Full Stack Advanced LLM Notes, always updating.
Above All
学习 LLM 最核心的技术,始终 follow 时代的前沿,同时 top down 学习对应的基础知识。
Model Architecture
基座模型架构
LLM Infra
LLM Infra 源码解读
Advanced LLMs 1: Attentions
从 MHA、MQA/GQA、MLA 到稀疏注意力(LongFormer / StreamingLLM / NSA / DSA / MSA)、线性注意力(Mamba / Gated DeltaNet)与 Flash Attention。
Advanced LLMs 2: Normalizations, FFNs & RoPE
Pre-Norm / Post-Norm 与 RMSNorm 的推导、激活函数与 SwiGLU FFN,以及 RoPE 的旋转矩阵与长度外推(PI / NTK-aware / YaRN)。
Advanced LLMs 3: Decoding
Prefill 与 Decode 两个阶段的 FLOPs 与 Memory-Bound 分析,以及 Speculative Decoding、Multi-Token Prediction。
Advanced LLMs 4: Optimizers
预训练优化器与训练稳定性。
Advanced LLMs 5: Mixture of Experts
MoE 的路由、负载均衡与训练。
Slime 101
THUDM slime:连接 Megatron 与 SGLang、为 RL scaling 设计的 LLM post-training 框架。
Slime: 训练主流程
从 train.py 入口出发,梳理 Slime 如何用 Ray 封装 actor,并挂接 Megatron 在 GPU 上完成训练。