High-performance ML systems: GPU computing, distributed training, compression, and LLM serving
The course covers the systems engineering required to train and deploy industrial-scale models. Students study GPU architecture, CUDA kernel optimization, distributed training techniques (data/model/pipeline parallelism), model compression, and high-performance LLM inference systems. The course directly underpins team success in Practicum 3 and serves as a prerequisite for research tracks on efficient architectures.