Sources¶
Toutes les sources effectivement consultées lors de la rédaction, classées par type. Les affirmations chiffrées importantes du document renvoient à l'une d'elles.
Sources primaires¶
Documentations officielles NVIDIA¶
- CUDA C++ Programming Guide
- CUDA C++ Best Practices Guide
- PTX ISA
- Hopper Tuning Guide 13.3
- Blackwell Tuning Guide 13.3
- Nsight Compute Profiling Guide
- Nsight Compute documentation 13.3
- Compute Sanitizer User Manual
- CUTLASS documentation — Python DSL Overview · Functionality · Efficient GEMM in CUDA
- cuBLAS · cuDNN · cuFFT · cuRAND · CCCL · NCCL · NVSHMEM
- CUDA Toolkit 13.3 Release Notes
- CUDA Toolkit 13.4 Developer Preview Release Notes
- CUDA Installation Guide
- CUDA GPUs — Compute Capability
Documentations officielles AMD¶
- Documentation ROCm
- ROCm 7.14 compatibility matrix
- AMD Instinct MI300/MI350 workload optimization
- Helion sur GPU AMD, ROCm AI Developer Hub
Autres documentations¶
- Triton documentation · Gluon Overview · Introduction to Gluon
- PyTorch documentation · Custom C++ and CUDA Extensions · Custom Operators
- Helion documentation et dépôt
- JAX Pallas · CuTe DSL et JAX
- Mojo — module
gpu - WebGPU Shading Language, W3C
- OCP Microscaling Formats (MX) Specification
- Metal, Apple · MLX · Kokkos · SYCL, Khronos
Papiers de recherche¶
Megakernels — le corpus central¶
| Référence | Ce qu'elle soutient dans ce document |
|---|---|
| Gupta, Stuart, Owens, A Study of Persistent Threads Style GPU Programming for GPGPU Workloads, InPar 2012 — eScholarship, GTC 2012 | les quatre cas d'usage des persistent threads ; « peut aussi entraîner une perte de performance » |
| Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs — arXiv:2512.22219, HTML | ttGraph, événements, workers/schedulers, 1,0-1,7×, 14,5→12,5 ms, pages de 32 Ko, 352 octets/tâche, fusion gather-GEMM à 11 % |
| Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel — arXiv:2604.13327 | le problème du dynamisme, ETC |
| Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs — arXiv:2604.15379 | chiplet-tasks, L2 12→54 %, −37 % HBM, 1,3-1,5× vs vLLM |
| Ada-MK: Adaptive MegaKernel Optimization via Automated DAG-based Search — arXiv:2605.11581, HTML | 14,6 % de surcoût de lancement, +23,6 % / +50,2 %, déploiement Baidu |
| AutoMegaKernel: A Statically-Checked Agent Harness — arXiv:2606.09682 | 7 160 ordonnancements validés, parité jeton à jeton, 1,25-1,72×, L4/L40S |
Microarchitecture¶
- Dissecting the NVIDIA Volta GPU Architecture via Microbenchmarking — arXiv:1804.06826
- Dissecting the NVidia Turing T4 GPU via Microbenchmarking — arXiv:1903.07486
- Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and
Multiple Level Analysis — arXiv:2501.12084
→
wgmmaà 95 % du pic contre 62,9 % pourmma - Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural
Analysis — arXiv:2512.02189
→ latences
tcgen05(11,0-11,4 cycles), TMEM (256 Ko, 16 To/s, 420 cycles), débits par précision, moteur de décompression
Modèles de performance¶
- Williams, Waterman, Patterson, Roofline: An Insightful Visual Performance Model for Multicore Architectures, CACM 52(4), 2009 — DOI
- Hierarchical Roofline Performance Analysis for Deep Learning Applications — arXiv:2009.05257
- Volkov, Better Performance at Lower Occupancy, GTC 2010 — PDF
- Optimization Techniques for GPU Programming, ACM Computing Surveys 55(11) — ACM DL
Attention et noyaux d'IA¶
- Vaswani et al., Attention Is All You Need — arXiv:1706.03762
- Dao et al., FlashAttention — arXiv:2205.14135
- Dao, FlashAttention-2 — arXiv:2307.08691
- Shah et al., FlashAttention-3 — arXiv:2407.08608
- FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling — arXiv:2603.05451 → 1 605 TFLOPS sur B200, 71 % d'utilisation, 1,3× cuDNN, 2,7× Triton, exponentielles sur FMA, ~10× moins de rescalings, écrit en CuTe DSL
- Ainslie et al., GQA — arXiv:2305.13245
- Kwon et al., PagedAttention / vLLM — arXiv:2309.06180
- Leviathan et al., Speculative Decoding — arXiv:2211.17192
- Cai et al., Medusa — arXiv:2401.10774
- Shoeybi et al., Megatron-LM — arXiv:1909.08053
- Kaplan et al., Scaling Laws for Neural Language Models — arXiv:2001.08361
- Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode — arXiv:2605.30571
Quantification¶
- Frantar et al., GPTQ — arXiv:2210.17323
- Lin et al., AWQ — arXiv:2306.00978
- Xiao et al., SmoothQuant — arXiv:2211.10438
Compilateurs et DSL¶
- Tillet, Kung, Cox, Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations, MAPL 2019 — ACM DL
- Liger Kernel: Efficient Triton Kernels for LLM Training — arXiv:2410.10989
- The Anatomy of a Triton Attention Kernel — arXiv:2511.11581
- Optimal Software Pipelining and Warp Specialization for Tensor Core GPUs — arXiv:2512.18134
- Mojo: MLIR-Based Performance-Portable HPC Science Kernels — arXiv:2509.21039
MoE et communication¶
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations — arXiv:2512.14080
- Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations — arXiv:2606.05951
- GPU-Initiated Networking for NCCL — arXiv:2511.15076
- TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference — arXiv:2505.11329
- Deep Kernel Fusion for Transformers — arXiv:2602.11808
- Tile-Level Activation Overlap for Efficient LLM Inference — arXiv:2607.02521
Génération automatique de noyaux¶
- KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels — arXiv:2605.04956
- FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems — arXiv:2601.00227
- EvoEngineer: Mastering Automated CUDA Kernel Code Evolution with LLMs — arXiv:2510.03760
- ARGUS: Agentic GPU Optimization Guided by Data-Flow Invariants — arXiv:2604.18616
- Optimizing CUDA like a Human: Micro-Profiling Tools as Expert Surrogates — arXiv:2606.26453
- KForge: Program Synthesis for Diverse AI Hardware Accelerators — arXiv:2511.13274
- Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU — arXiv:2605.26118
Calcul scientifique et graphique¶
- MFC 5.0: An exascale many-physics flow solver — arXiv:2503.07953
- Aurora: Architecting Argonne's First Exascale Supercomputer — arXiv:2509.08207
- GPU Performance Portability needs Autotuning — arXiv:2505.03780
- GPU Acceleration of CFD Simulations in OpenFOAM, MDPI Aerospace 10(9), 792 — MDPI
- GPU-accelerated simulations of turbulence, Phys. Rev. Fluids — APS
- Kerbl et al., 3D Gaussian Splatting for Real-Time Radiance Field Rendering, SIGGRAPH 2023 — page projet
- Cooperative Kernels: GPU Multitasking for Blocking Algorithms — arXiv:1707.01989
Sources primaires — blogs des équipes concernées¶
Ces billets sont des sources primaires : ils sont écrits par les auteurs des travaux et contiennent des chiffres qu'on ne trouve nulle part ailleurs.
| Source | Ce qu'elle soutient |
|---|---|
| Hazy Research, Look Ma, No Bubbles! | 1,3 µs par lancement avec CUDA Graphs ; 80 SM inactifs sur 512 blocs ; décompte des 600 µs sur B200 ; 13 pages de 16 Ko ; < 1 ms sur H100, ~680 µs sur B200 ; 78 % de bande passante ; 2,5× vLLM, 1,5× SGLang ; ~770 contre ~1 350 passes/s ; limites de PDL ; barrières à ~60 ns |
| Hazy Research, We Bought the Whole GPU | 23 468 contre 19 170 jetons/s (+22 %) ; recouvrement à trois niveaux (2-6 %, 14,2 %) ; transposition distribuée (trafic ÷8) ; neuf instructions ; > 90 % d'inactivité CPU ; débordements de registres |
| Kog, Single-kernel LLM inference on MI300X | ~4,5 µs par lancement ; 35 % du temps en synchronisation de grille ; sentinelle 7,59-7,88 → 0,80-0,93 µs ; duplication par XCD ; > 3 000 jetons/s ; DTP |
| Zhihao Jia, Compiling LLMs into a MegaKernel | présentation de MPK par son auteur |
| Catalyst (CMU), page MPK | contexte du projet |
| CMU 15-779, lecture 7 : Mega-Kernel | présentation pédagogique |
Sources secondaires — analyses et tutoriels¶
Tutoriels techniques¶
- Simon Boehm, How to Optimize a CUDA Matmul Kernel et le dépôt
- salykova, Advanced Matrix Multiplication Optimization on NVIDIA GPUs
- salykova, Matrix Core Programming on AMD CDNA3 and CDNA4
- Outperforming cuBLAS on H100: a Worklog
- 6 Step Optimization of GeMMs in CUDA
- Colfax Research — série CUTLASS : TMA · WGMMA · Pipelining · Clusters Blackwell
- Modern GPU Programming for MLSys, MLC
- tcgen05 for dummies, gau-nernst
- tcgen05 and TMEM, Blackwell GPU Wiki
Blogs NVIDIA et PyTorch¶
- NVIDIA Hopper Architecture In-Depth
- NVIDIA CUDA 13.1 Powers Next-Gen GPU Programming with CUDA Tile
- What's New and Important in CUDA Toolkit 13.0
- Using CUDA Warp-Level Primitives
- Cooperative Groups
- An Efficient Matrix Transpose in CUDA C/C++
- CUDA Pro Tip: Grid-Stride Loops
- Getting Started with CUDA Graphs
- Deep Dive on CUTLASS Ping-Pong GEMM Kernel, PyTorch
- Deep Dive on the Hopper TMA Unit for FP8 GEMMs, PyTorch
- Warp Specialization in Triton, PyTorch
- Helion: A High-Level DSL, PyTorch
- Flash-Decoding for long-context inference, PyTorch
- Structured Mojo Kernels, Modular
- Matrix Multiplication on NVIDIA's Blackwell, Modular
Références et glossaires¶
- Modal GPU Glossary — SM · Warp Scheduler · Warp divergence · Roofline · CuTe DSL
- Modal, We reverse-engineered Flash Attention 4
- SemiAnalysis, NVIDIA Tensor Core Evolution: From Volta To Blackwell
- SemiAnalysis, Dissecting Nvidia Blackwell
- Chips and Cheese, Nvidia's B200
Communauté et curriculums¶
- GPU MODE — lectures · resource-stream
- Notes GPU MODE, Christian Mills
- gpu-perf-engineering-resources, wafer-ai
- cuda-learning, rkinas
- Mojo GPU Puzzles et le dépôt
- GPU Puzzles, Sasha Rush
Dépôts de code¶
- NVIDIA/cutlass
- HazyResearch/ThunderKittens · HazyResearch/Megakernels
- mirage-project/mirage
- triton-lang/triton
- Dao-AILab/flash-attention
- flashinfer-ai/flashinfer
- vllm-project/vllm · sgl-project/sglang
- deepseek-ai/DeepEP · deepseek-ai/DeepGEMM
- linkedin/Liger-Kernel
- pytorch/helion
- IST-DASLab/marlin
- ScalingIntelligence/KernelBench
- Issue vLLM #38022 — alignement de Marlin MoE
Sources de contexte¶
- Programming Massively Parallel Processors, 5e éd., Hwu, Kirk, El Hajj, Elsevier, février 2026 — ISBN 978-0-443-43900-1
- Stephen Jones, How GPU Computing Works, GTC
- Stephen Jones, How to Write a CUDA Program, GTC 2024
- Mark Harris, Optimizing Parallel Reduction in CUDA
- Merrill & Garland, Single-pass Parallel Prefix Scan with Decoupled Look-back
- NVIDIA Launches Rubin AI Compute Platform at CES 2026, ServeTheHome
- How to build unmaintainable kernels, Ian Barber
- Helion and the evolving GPU programming model, Ian Barber
- SYCL vs OpenCL vs Vulkan Compute, Till Code — l'écart de 10 à 30 % du code portable
- Choosing Vulkan, OpenCL, SYCL or CUDA for GPU Compute, TechnoLynx
- WebGPU 2026: Bringing GPU Compute and AI Inference to the Browser
- What is FP8 Quantization?, Spheron
- Deploy DeepEP and DeepGEMM, Spheron
- GPU Acceleration of Molecular Modeling Applications, UIUC
Méthodologie¶
Ce qui a été fait et ce qui ne l'a pas été
Fait :
- consultation directe de chaque source listée ;
- vérification croisée des chiffres importants entre au moins deux sources quand c'était possible ;
- attribution de chaque affirmation chiffrée à sa source dans le corps du document ;
- calculs arithmétiques (intensités, planchers de temps, tailles de cache) refaits et vérifiables au papier.
Pas fait :
- aucune mesure n'a été reproduite. L'environnement de rédaction ne dispose ni de GPU ni de compilateur CUDA ;
- aucun code n'a été compilé ni exécuté. Les extraits suivent les documentations citées mais n'ont pas été validés ;
- les papiers longs ont été lus par leur résumé, leurs figures et leurs sections de résultats, pas intégralement.
Les affirmations chiffrées sont donc des rapports de sources, pas des vérifications indépendantes. Elles sont attribuées comme telles dans tout le document.
Document créé le 22 août 2026.
Retour : Annexes · Index du cours