Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Tencent Hunyuan Production Fundamentals
- Overview of serving scenarios for Tencent Hunyuan models.
- Production characteristics specific to large and MoE models.
- Identifying common bottlenecks related to latency, throughput, and cost.
- Establishing service-level objectives for inference workloads.
Deployment Architecture and Serving Flow
- Core components of a production inference stack.
- Evaluating containerized, on-premise, and cloud deployment models.
- Fundamentals of model loading, request routing, and GPU allocation.
- Designing for reliability and operational simplicity.
Practical Latency Optimization
- Utilizing optimized inference engines like TensorRT where applicable.
- Understanding KV-cache concepts and applying practical cache tuning.
- Minimizing startup, warmup, and response overhead.
- Measuring time to first token and token generation speed.
Throughput, Batching, and GPU Efficiency
- Strategies for continuous batching and request batching.
- Managing concurrency and queue behavior.
- Enhancing GPU utilization while maintaining user experience.
- Processing long-context and mixed-workload requests.
Quantization and Cost Management
- The importance of quantization in production serving.
- Practical trade-offs between FP16, INT8, and other precision options.
- Balancing model quality, latency, and infrastructure costs.
- Developing a basic checklist for cost optimization.
Operations, Monitoring, and Readiness Review
- Configuring autoscaling triggers for inference services.
- Monitoring key metrics including latency, throughput, cache usage, and GPU health.
- Essentials of logging, alerting, and incident response.
- Reviewing reference deployments and formulating improvement plans.
Requirements
- Fundamental understanding of large language model deployment and inference workflows.
- Experience with containerization, cloud or on-premise infrastructure, and API-based services.
- Proficiency in Python or system engineering tasks.
Target Audience
- ML engineers responsible for bringing LLMs into production.
- Platform engineers managing GPU-based inference services.
- Solution architects designing scalable AI serving platforms.