Past projects
53 student projects from 11-868 Large Language Model Systems, grouped by research topic.
Attention & FlashAttention
13 projectsAttentions Are All You Need: Extending FlashAttention and PagedAttention to MiniTorch
Benchmarking FlashAttention and Its Variants for ViT
High-Performance GPU Kernels for DeepSeek Sparse Attention on NVIDIA Blackwell (B200) Best Poster
Implementing Grouped-Query Attention in MiniTorch: From Naive CUDA to FlashAttention-2 on V100
LLM Training and Inference Acceleration with FlashAttention and PagedAttention
High-Performance Sparse Attention Kernels for DeepSeek V3.2 on NVIDIA Blackwell
Efficient Attention Backends for LLM Inference in MiniTorch: FlashAttention and Paged KV Cache Optimization
FlashAttention for Long-Context Transformer Scaling
Efficient GPU Kernel Design for Block-Sparse Attention in Long-Context LLMs: Final Report
FlashAttention2 in MiniTorch & FlashAttention-4 in PithTrain
Implementing FlashAttention in MiniTorch: IO-Aware Fused Attention Kernels for Decoder-Only Transformer Training Best Poster
Sparse Pair-Biased Attention
Breaking the Memory Wall: Implementing IO-Aware FlashAttention in MiniTorch
KV Cache & Serving
10 projectsH2O: Attention-Aware KV Cache Eviction
AdaKV: Memory-Aware Dynamic KV-Cache Management for Edge LLM Inference
PagedAttention for Memory-Efficient LLM Serving on Volta Architecture
MiniTorch Extension: KV Cache and Efficient Decoding
PagedAttention in MiniTorch: KV Cache Efficiency, Prefix Reuse, and Memory Sharing
Efficient KV Cache Eviction for SGLang
Optimizing KV Cache with Quantization
Workload-Aware KV Cache Eviction and Cache-Aware Scheduling Best Poster
RAGCache++: Cache-Aware Document Ordering for Low-Latency RAG Serving
Conf-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Inference
RL & Fine-Tuning
10 projectsScalable DPO Training in MiniTorch
Improving LLM RL Training via Quantization and CPU Offloading
Asynchronous LLM Reinforcement Learning Under Limited Data and Constrained Hardware Performance
Fast Rollouts, Stable Updates: Systems Trade-Offs in GRPO Post-Training with Partial Rollouts
Efficient Policy Optimization for LLM Reasoning: GRPO in MiniTorch
Optimizing Transfer Protocols for RLHF Dataflows: Local-Batch Pull, Asynchronous Overlap, and Lightweight Compression in veRL
An Empirical Implementation and Analysis of Direct Preference Optimization in MiniTorch
Asynchronous Agentic Reinforcement Learning for Tool-Using Language Models: System Design, Ablations, and Findings on Qwen2.5-1.5B-Instruct Best Poster
LoRAMino: A Framework for Efficient Multi-Job LoRA Fine-Tuning
High-Throughput Rollout Optimization for Long-Context Coding Agents in veRL
Decoding & Long Context
7 projectsScaling Long-Input, Long-Output LLM Inference for Clinical Speech Documentation
Accelerating Large Reasoning Models via Semantics-Aware Lookahead Decoding with Distillation and Lightweight Hybrid Verification
Efficient Recursive Language Model: System Optimizations for Long-Context Question Answering
Verifier-Guided Context Selection for Long-Context Speculative Decoding
Accelerating LLM Inference via Set Block Decoding
DynaState: Training-Free Long-Context Recall for Hybrid Linear-Attention LLMs via Selective Token Replay
One Cache, Two Models: Bridging the Dimensionality Gap for Memory-Efficient Speculative Decoding
Kernels & MoE
8 projectsEvolutionary Kernel Agents with Adaptive Skill Wisdom for FP8 Fused MoE Inference on NVIDIA Blackwell
Extending AdaSplash to JAX and CUDA
Triage: A Triton-Backed Megakernel Compiler for Autoregressive Transformer Inference
BeamAgent: Lineage-Aware Beam Search for Agent-Driven CUDA Kernel Optimization
Temporal Expert Caching for Accelerating Inference in Mixture-of-Experts Diffusion Language Models Best Poster
Optimizing Qwen3-30B-A3B Inference on AWS Trainium with Custom NKI Kernels
Optimizing MoE and Sparse Attention Layers on Blackwell GPU
Fused Mixture of Experts Kernel Generation
Architectures & System Extensions
5 projectsEfficient Systems Implementation of Manifold-Constrained Hyper-Connections
Heterogeneous Execution of Spiking Neural Networks Across CPU-GPU Systems
Exploring Efficient Training and Inference of Diffusion LLMs
Extending MiniTorch Scalable Components
Parallelism and Optimizations for Mamba-2 Models