Course Outline
Introduction
Grasping the Core Principles of Heterogeneous Computing Methodology
The Case for Parallel Computing: Understanding the Necessity
Multi-Core Processors: Exploring Architecture and Design
Thread Essentials: Intro to Threads, Basics, and Parallel Programming Concepts
Foundations of GPU Software Optimization Workflows
OpenMP: The Standard for Directive-Based Parallel Programming
Practical Session: Demonstrating Various Programs on Multicore Machines
Getting Started with GPU Computing
Leveraging GPUs for Parallel Workloads
The GPU Programming Model Explained
Practical Session: Demonstrating Various Programs on GPUs
Setting Up the SDK, Toolkit, and Development Environment for GPUs
Utilizing Diverse Library Ecosystems
Demonstrating GPU Tools and Sample Programs with OpenACC
Unpacking the CUDA Programming Model
Delving into the CUDA Architecture
Configuring and Exploring CUDA Development Environments
Interacting with the CUDA Runtime API
Comprehending the CUDA Memory Model
Investigating Extended CUDA API Capabilities
Optimizing Global Memory Access in CUDA: Efficiency Strategies
Enhancing Data Transfer Performance in CUDA via CUDA Streams
Leveraging Shared Memory Within CUDA
Mastering Atomic Operations and Instructions in CUDA
Case Study: Implementing Basic Digital Image Processing with CUDA
Advanced Strategies for Multi-GPU Programming
Performing Advanced Hardware Profiling and Sampling on NVIDIA / CUDA
Utilizing the CUDA Dynamic Parallelism API for Dynamic Kernel Launch
Key Takeaways and Final Wrap-up
Requirements
- C Programming
- Linux GCC
Testimonials (1)
Trainers energy and humor.