Skip to main content

Syllabus

WeekDatesTopicReading/ContentHomework
Week 18/24Introduction to LLM[slides]
8/26GPU Programming Basics 1[slides]Chap 2,3,4 of Programming Massively Parallel Processors, 4th Ed
8/28Recitation 1: PSC Guidelines, Simple CUDA Demo, HW1 slides
Week 28/31GPU Acceleration[slides]Chap 5,6 of Programming Massively Parallel Processors, 4th Ed
9/2Deep Learning Frameworks and Auto Differentiation [slides]Tensorflow Auto Diff survey Differentiable ProgrammingHW1 Due
9/4Recitation 2: HW2, MiniTorch, More GPU
Week 39/7no class
9/9LLM Architecture and Pre-training [slides]Attention is all you need, Deepseek-V2, GPT3HW2 Due
9/11Recitation 3: Annotated TransformersAnnotated Transformer
Week 49/14Tokenization and Embedding [slides]BPE, Sentence-Piece, VOLT
9/16Generation and Speculative Decoding [slides]Speculative DecodingProject Team Due
9/18Recitation 4: Tokens and Decoding
Week 59/21Accelerating Transformer on GPU Part 1[slides]LightSeq
9/23Accelerating Transformer on GPU Part 2[slides]LightSeq2HW3 Due
9/25Recitation 5: Lightseq demo
Week 69/28Guest Lecture by Srinath Mandalapu (Google): TPU and JAX [slides]
9/30Guest Lecture by Srinath Mandalapu (Google): Pallas and Splash Attention [slides]Project Proposal Due
10/2Recitation 6: JAX and TPU
Week 710/5Distributed Data Parallel Training [slides]DDP
10/7Model Parallel Training [slides]GPipe, Megatron-LMHW4 Due
10/9Recitation 7 Distributed training
Week 810/12spring break
Week 910/19Systems for Mixture-of-Expert Models [slides]GShard, Switch Transformer, DeepSpeed-MOE, Deepseek-MoE
10/21Memory Optimization in Distributed Training [slides]ZeRO (DeepSpeed)
10/23Recitation 8 MoE
Week 1010/26Model Quantization [slides]NN Quantization, AdaQuant, LLM.int8()HW5 Due
10/28Model Quantization II [slides]GPTQ
10/30Recitation 9 Quantization
Week 1111/2Efficient fine-tuning for Large Models [slides]CIAT, LORA, QLoRA
11/4Optimizing Attention for Modern Hardware (Tri Dao) [slides]FlashAttention FlashAttention2 FlashAttention3 FlashAttention4Mid-term Report Due
11/6Recitation 10 LoRA and FlashAttention
Week 1211/9LLM serving with SGL [slides]ORCA SGLang
11/11Efficient LLM Inference with Paged Attention and vLLM (Woosuk Kwon) [slides]vLLM
11/13Recitation 11 SGLang and vLLMHW6 Due
Week 1311/16Acceleration on TPU1
11/18Acceleration on TPU2
11/20Recitation 12: TPU accelerationDistributed Training
Week 1411/23Efficient Reinforcement Learning System for LLMsReaLHF
11/25no class
Week 1511/30Serving with Disaggregated Prefill-Decoding [slides]DistServeHW7 Due
12/2Better KV Cache for LLM Serving (Junchen Jiang) [slides]CacheGen CacheBlend
12/4Final project presentation
12/6Final report due
DistServe: Disaggregated Prefill-Decoding (Hao Zhang) [slides]DistServe
App Stack and Model Serving[slides]Triton, LightLLM
Triton for Kernel OptimizationJAX
Retrieval-augmented Language ModelsRAG
Nearest Vector Search for EmbeddingsHNSW
Multimodal LLMsFlamingo
Efficient Streaming Language Models with Attention SinksAttention Sink
LLM Serving on Heterogeneous Hardware [slides]Mooncake, kTransformer