Agentic Kernel Engineering
KernelWiki 负责找路,ncu-report-skill 负责验路。
Kernel Design Agents 的顶层仓库很小,真正有信息密度的是两个 submodule:一个把 Blackwell/Hopper kernel 优化经验整理成可检索知识库,一个把 Nsight Compute profiling 固化成证据链。
先看顶层仓库
顶层的 kernel-design-agents 更像一本操作手册,而不是 kernel 源码库。它只保留通用流程:
task contract、starter prompt、agent flow、以及安装两个 skills 的方式。真正实现任务时,要另建
task workspace,把代码、数据、benchmark、profiling artifact 和 candidate 记录放在那里。
Skill 1
KernelWiki 里面有什么
KernelWiki 是 Blackwell-first 的 GPU kernel 优化知识库,打包成 Claude Code skill。它的知识截止日期是
2026-04-27,覆盖 SM100/B200 和一部分带 Blackwell 迁移价值的 SM90/Hopper 内容。
PR、官方文档、blog、competition 页面,保留原始来源摘要。
硬件、技术、kernel case、问题模式、语言、迁移指南。
按 problem、technique、hardware、repo、kernel type、language 自动生成索引。
主目录怎么读
| 目录 | 作用 | 典型内容 |
|---|---|---|
wiki/hardware/ |
SM100 硬件特性入口 | tcgen05、TMEM、TMA、CLC、2-SM cooperative、NVFP4、mbarrier |
wiki/techniques/ |
优化手法 | warp specialization、persistent kernels、swizzling、epilogue fusion、pipeline stages |
wiki/kernels/ |
kernel case study | FlashAttention-4、DeepGEMM、FlashMLA、NVFP4 GEMM/GEMV、Fused MoE、Gated Delta Net |
wiki/patterns/ |
从症状到方案 | low SM utilization、memory-bound、register pressure、tail effect、pipeline stalls |
wiki/languages/ |
语言和 DSL 指南 | CuTe DSL、CUDA C++、PTX SM100、Triton Blackwell |
artifacts/ |
可追溯代码材料 | PR diff、关键 kernel 文件、derived teaching skeleton,带 PROVENANCE.yaml |
它真正有用的点
把问题归类
看到 “tail effect” 或 “memory-bound”,先去 queries/by-problem.md,再跳到 pattern 页面和候选技术。
把方案追到证据
wiki 页有 sources: 字段,可追到具体 PR、doc、blog。回答时要看 confidence:verified、source-reported、inferred、experimental。
把硬件名词落到代码
例如 hw-tcgen05-mma 会说明 tcgen05 和 wgmma 的差异,hw-tmem 会讲 TMEM 分配、读写和 fence。
把 case study 当路线图
FlashAttention-4、NVFP4 GEMV、CuTe DSL tutorial 这类页面给出了 technique 组合和性能 claim。
常用查询命令
python3 scripts/query.py "flash attention blackwell" --compact --limit 5
python3 scripts/query.py --tag nvfp4 --type kernel --compact
python3 scripts/get_page.py kernel-flash-attention-4 --follow-sources
python3 scripts/get_page.py hw-tcgen05-mma --body-only
python3 scripts/grep_wiki.py "tcgen05\\.mma" --only sources
Skill 2
ncu-report-skill 里面有什么
ncu-report-skill 是一套 B200/sm_100 优先的 Nsight Compute 工作流。它不直接给你“猜优化方向”,而是强迫 agent
先建 harness、跑 profile、抽 metric、按六个维度诊断,最后写一个可复现的 REPORT.md。
核心参考文档
| 文件 | 重点 |
|---|---|
reference/00-directory-layout.md |
规定 profiling artifact 的目录结构:一个 run 一个目录,避免混结果。 |
reference/03-collection.md |
标准 ncu 命令:--set full 加 PM sampling,以及 --set source 加 SourceCounters。 |
reference/04-python-api.md |
用 ncu_report Python API 读 report,不靠肉眼读 CLI 表格。 |
reference/05-analysis-dimensions.md |
六个诊断维度:occupancy、tail effect、stall、tensor core、timeline、memory pattern。 |
reference/06-diagnosis-playbook.md |
把 NCU signal 映射到 cause 和 fix:small grid、tail、uncoalesced、long scoreboard、register spill 等。 |
reference/08-b200-metric-names.md |
B200 上实际可用 metric 名称。很多旧 GPU 名称在 sm_100 上会失效。 |
helper 脚本
harness_template.cu
独立 profiling harness 模板,用来隔离 kernel 并编译 -lineinfo。
analyze_reports.py
抽取 B200 key metrics,保存 full JSON 和可读 txt,多 report 时生成对比表。
extract_stall_hotspots.py
把 source-level per-PC stall sample 聚合到源文件行号,定位最热 stall 行。
plot_timeline.py
把 PM sampling 时间序列画成 ASCII 图,直接看 tail effect 和 pipeline bubbles。
list_flashinfer_workloads.py
浏览 flashinfer-trace workload,选真实 shape,而不是凭空造输入。
safetensors_loader.h
header-only safetensors reader,方便 harness 直接读真实 tensor。
标准采集命令
export PROFILE_RUN_DIR=profile/my_kernel_v1_baseline
mkdir -p "$PROFILE_RUN_DIR"/{harness,reports,analysis}
ncu --set full \
--section PmSampling \
--section PmSampling_WarpStates \
-k "regex:KERNEL_REGEX" -c 1 \
-o "$PROFILE_RUN_DIR/reports/full_main" \
"$PROFILE_RUN_DIR/harness/my_harness" [args]
ncu --set source --section SourceCounters \
-k "regex:KERNEL_REGEX" -c 1 \
-o "$PROFILE_RUN_DIR/reports/source_main" \
"$PROFILE_RUN_DIR/harness/my_harness" [args]
两者怎么配合 KDA
定 contract
把目标、正确性、评测命令、promotion criteria 写清楚。没有这个,agent 很容易乱优化。
用 KernelWiki 找候选方向
根据症状或 kernel 类型找 hardware feature、technique、case study,再追 sources 看可信度。
实现一个 candidate
每次只改一个有明确机制的方向,例如 tile scheduling、vectorized loads、epilogue fusion 或 CLC。
用 ncu-report-skill 验证
跑 full/source profiles,抽 key metrics、stall hotspots、timeline,写入 REPORT.md。
promote 或 reject
只有 correctness 通过且 metric 有证据支持时才保留;失败也记录原因,避免重复踩坑。
我的使用建议
把 KernelWiki 当“设计检索器”,不要当权威答案。
它的优势是索引、source chain 和术语归一化;遇到 performance claim,要看 source_id、shape、dtype、GPU 和 confidence。
把 ncu-report-skill 当“证据纪律”。 优化前先确认瓶颈,不要看到 Blackwell 就直接上 tcgen05、TMEM 或 CLC。NCU 的 rule engine、stall hotspot 和 PM timeline 往往能先告诉你该不该改。
两者之间缺的那一层,是 candidate ledger。
KDA 顶层已经建议记录 benchmark.csv 和 candidates.jsonl。实际做任务时,建议每个 candidate 写清 parent、改动机制、验证结果、profiling 证据和最终状态。
来源
- mit-han-lab/kernel-design-agents
- DongyunZou/KernelWiki
- DongyunZou/ncu-report-skill
- 本页整理时间:2026-06-01。KernelWiki 自身标注的知识截止日期:2026-04-27。