Get in Touch

Course Outline

Introduction to CANN Optimization Features

  • Mechanisms for handling inference performance within CANN
  • Key optimization targets for edge and embedded AI systems
  • Concepts of AI Core usage and memory distribution

Analysis via the Graph Engine

  • Fundamentals of the Graph Engine and its execution workflow
  • Visualization of operator graphs and runtime metrics
  • Adjusting computational graphs to enhance performance

Performance Metrics and Profiling Utilities

  • Utilizing the CANN Profiling Tool for workload assessment
  • Evaluating kernel execution times and identifying bottlenecks
  • Profiling memory access patterns and implementing tiling strategies

Developing Custom Operators with TIK

  • Overview of TIK and its operator programming framework
  • Building a custom operator using the TIK DSL
  • Validating and benchmarking operator efficiency

Advanced Operator Tuning with TVM

  • Foundations of integrating TVM with CANN
  • Strategies for auto-tuning computational graphs
  • Determining when to switch between TVM and TIK

Techniques for Memory Optimization

  • Controlling memory layout and buffer positioning
  • Methods to minimize on-chip memory usage
  • Best practices for asynchronous execution and data reuse

Practical Deployments and Case Studies

  • Case study: Tuning performance for smart city camera pipelines
  • Case study: Optimizing the inference stack for autonomous vehicles
  • Guidelines for continuous profiling and iterative improvement

Conclusion and Future Directions

Requirements

  • Solid grasp of deep learning model architectures and training pipelines
  • Hands-on experience deploying models via CANN, TensorFlow, or PyTorch
  • Proficiency in Linux CLI, shell scripting, and Python programming

Intended Audience

  • AI performance engineers
  • Specialists in inference optimization
  • Developers focused on edge AI or real-time systems
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories