[CPU] 为 csrc/cpu/ 添加 batch-invariant greedy/top-k/top-p 采样内核
Issue #27433 · [Feature]: Batch Invariant Feature and Performance Optimization ↗
已指派 bwasti, PaulZhang12, yewentao256, sfeng33已有 PR #399 条评论good first issuefeature request更新 2026-10-04
用到的专长:使用候选人在算子融合、内核实现、分布式训练确定性方面的专长,直接在 C++ 层实现 batch-invariant 采样。
目标:在 csrc/cpu/sampling_kernels.cpp 中新增 batch-invariant 的 greedy、top-k、top-p 采样实现,使 VLLM_BATCH_INVARIANT=1 在 CPU 后端也能端到端通过 tests/v1/determinism/ 的相关测试。
为什么值得长期做:Issue #27433 是 vLLM 当前最活跃的 batch-invariance tracker(99 条评论、good first issue、4 位 assignee),但 CPU 侧采样仍是空白。csrc/cpu/sampling_kernels.cpp 已有 fused_gumbel_argmax_kernel 与 greedy_argmax_kernel,但无 batch-invariant 版本。这是无 GPU 环境下能直接发挥 kernel 专长、且维护者明确想要的贡献,通向 CPU 后端 ownership 的最短路径。
第一个 PR 的边界:第一个 PR 仅新增内核 + 测试,不修改现有 CUDA 路径或 scheduler 逻辑;边界在 csrc/cpu/sampling_kernels.cpp 与 tests/kernels/test_cpu_sampling_batch_invariant.py(新增)。
认领留言(英文,可直接贴到 Issue)
Hi all, I'd like to pick up the CPU-side batch-invariant sampling kernels (greedy/top-k/top-p) under this tracker. I've read csrc/cpu/sampling_kernels.cpp and csrc/core/batch_invariant.hpp — the existing fused_gumbel_argmax_kernel gives a clear template. My plan: add batch_invariant_greedy/top_k/top_p kernels in C++, register them in the CPU backend, and add tests/kernels/test_cpu_sampling_batch_invariant.py covering bitwise consistency under VLLM_BATCH_INVARIANT=1. Could you confirm (1) if this sub-task is still unclaimed, (2) whether you prefer pure C++ kernels or a Python wrapper, and (3) if temperature scaling should be included? I'll have a draft PR ready in about 5–7 days.
大致实施方案
- 阅读 csrc/cpu/sampling_kernels.cpp 与 csrc/core/batch_invariant.hpp,理解现有 fused_gumbel_argmax_kernel / greedy_argmax_kernel 的接口与 batch invariant 抽象
- 阅读 tests/v1/determinism/test_batch_invariance.py,理解 batch-invariant 测试的断言逻辑与调用路径
- 在 csrc/cpu/sampling_kernels.cpp 中新增 batch_invariant_greedy / top_k / top_p 内核,遵循 batch_invariant.hpp 的接口约定
- 在 csrc/cpu/ 的 CMakeLists.txt 或对应构建文件中注册新内核,确保 -mcpu=apple-m1 / NEON 路径可编译
- 在 tests/kernels/ 新增 test_cpu_sampling_batch_invariant.py(新增),覆盖 greedy/top-k/top-p 的 bitwise 一致性
- 在 MacBook 上编译并跑通测试,记录基准数据
可能涉及的目录或文件
- csrc/cpu/sampling_kernels.cpp → 新增内核实现
- csrc/cpu/ 构建注册 → CMakeLists.txt 或 torch_bindings
- tests/kernels/test_cpu_sampling_batch_invariant.py → 新增测试文件
- 需先定位 batch invariant 抽象在 CPU 侧的现有 hook 点
验收方式
- 在 MacBook 上编译 csrc/cpu/ 并运行新增的 test_cpu_sampling_batch_invariant.py,断言 bitwise 一致
- 运行 tests/v1/determinism/test_batch_invariance.py 中 CPU 相关用例(若有),确认不回归
- 在 Colab T4 上编译并跑同一测试,确认 CPU 路径与 CUDA 路径行为一致
开工前问题与风险
向维护者确认
- CPU 侧 batch-invariant 采样是否已有设计文档或接口约定?
- 维护者偏好纯 C++ kernel 还是 Python wrapper 调用现有算子?
- 是否需要同时覆盖 sampling post-processing(如 temperature scaling)?
风险
- batch invariant 抽象可能仍在快速迭代,接口可能变动
- CPU 采样路径可能依赖 scheduler 侧改动,需确认边界
- Apple Silicon 上编译 C++ 内核可能遇到 NEON / DNNL 兼容性问题