Skip to main content

Past projects

53 student projects from 11-868 Large Language Model Systems, grouped by research topic.

Attention & FlashAttention​

13 projects
  • Attentions Are All You Need: Extending FlashAttention and PagedAttention to MiniTorch

  • Benchmarking FlashAttention and Its Variants for ViT

  • High-Performance GPU Kernels for DeepSeek Sparse Attention on NVIDIA Blackwell (B200) Best Poster

  • Implementing Grouped-Query Attention in MiniTorch: From Naive CUDA to FlashAttention-2 on V100

  • LLM Training and Inference Acceleration with FlashAttention and PagedAttention

  • High-Performance Sparse Attention Kernels for DeepSeek V3.2 on NVIDIA Blackwell

  • Efficient Attention Backends for LLM Inference in MiniTorch: FlashAttention and Paged KV Cache Optimization

  • FlashAttention for Long-Context Transformer Scaling

  • Efficient GPU Kernel Design for Block-Sparse Attention in Long-Context LLMs: Final Report

  • FlashAttention2 in MiniTorch & FlashAttention-4 in PithTrain

  • Implementing FlashAttention in MiniTorch: IO-Aware Fused Attention Kernels for Decoder-Only Transformer Training Best Poster

  • Sparse Pair-Biased Attention

  • Breaking the Memory Wall: Implementing IO-Aware FlashAttention in MiniTorch

KV Cache & Serving​

10 projects
  • H2O: Attention-Aware KV Cache Eviction

  • AdaKV: Memory-Aware Dynamic KV-Cache Management for Edge LLM Inference

  • PagedAttention for Memory-Efficient LLM Serving on Volta Architecture

  • MiniTorch Extension: KV Cache and Efficient Decoding

  • PagedAttention in MiniTorch: KV Cache Efficiency, Prefix Reuse, and Memory Sharing

  • Efficient KV Cache Eviction for SGLang

  • Optimizing KV Cache with Quantization

  • Workload-Aware KV Cache Eviction and Cache-Aware Scheduling Best Poster

  • RAGCache++: Cache-Aware Document Ordering for Low-Latency RAG Serving

  • Conf-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM Inference

RL & Fine-Tuning​

10 projects
  • Scalable DPO Training in MiniTorch

  • Improving LLM RL Training via Quantization and CPU Offloading

  • Asynchronous LLM Reinforcement Learning Under Limited Data and Constrained Hardware Performance

  • Fast Rollouts, Stable Updates: Systems Trade-Offs in GRPO Post-Training with Partial Rollouts

  • Efficient Policy Optimization for LLM Reasoning: GRPO in MiniTorch

  • Optimizing Transfer Protocols for RLHF Dataflows: Local-Batch Pull, Asynchronous Overlap, and Lightweight Compression in veRL

  • An Empirical Implementation and Analysis of Direct Preference Optimization in MiniTorch

  • Asynchronous Agentic Reinforcement Learning for Tool-Using Language Models: System Design, Ablations, and Findings on Qwen2.5-1.5B-Instruct Best Poster

  • LoRAMino: A Framework for Efficient Multi-Job LoRA Fine-Tuning

  • High-Throughput Rollout Optimization for Long-Context Coding Agents in veRL

Decoding & Long Context​

7 projects
  • Scaling Long-Input, Long-Output LLM Inference for Clinical Speech Documentation

  • Accelerating Large Reasoning Models via Semantics-Aware Lookahead Decoding with Distillation and Lightweight Hybrid Verification

  • Efficient Recursive Language Model: System Optimizations for Long-Context Question Answering

  • Verifier-Guided Context Selection for Long-Context Speculative Decoding

  • Accelerating LLM Inference via Set Block Decoding

  • DynaState: Training-Free Long-Context Recall for Hybrid Linear-Attention LLMs via Selective Token Replay

  • One Cache, Two Models: Bridging the Dimensionality Gap for Memory-Efficient Speculative Decoding

Kernels & MoE​

8 projects
  • Evolutionary Kernel Agents with Adaptive Skill Wisdom for FP8 Fused MoE Inference on NVIDIA Blackwell

  • Extending AdaSplash to JAX and CUDA

  • Triage: A Triton-Backed Megakernel Compiler for Autoregressive Transformer Inference

  • BeamAgent: Lineage-Aware Beam Search for Agent-Driven CUDA Kernel Optimization

  • Temporal Expert Caching for Accelerating Inference in Mixture-of-Experts Diffusion Language Models Best Poster

  • Optimizing Qwen3-30B-A3B Inference on AWS Trainium with Custom NKI Kernels

  • Optimizing MoE and Sparse Attention Layers on Blackwell GPU

  • Fused Mixture of Experts Kernel Generation

Architectures & System Extensions​

5 projects
  • Efficient Systems Implementation of Manifold-Constrained Hyper-Connections

  • Heterogeneous Execution of Spiking Neural Networks Across CPU-GPU Systems

  • Exploring Efficient Training and Inference of Diffusion LLMs

  • Extending MiniTorch Scalable Components

  • Parallelism and Optimizations for Mamba-2 Models