一、项目定位 Molt 是 NVIDIA NeMo 实验室维护的一个 agentic-first 强化学习训练框架,定位是研究导向、PyTorch / HuggingFace 原生、面向 1T 级 MoE 规模的全异步 agentic RL。它把整个栈压到三个薄层:Ray 负责 placement 与 rollout↔trainer 之间的异步队列,vLLM 负责 rollout 生成,NVIDIA AutoModel + FSDP2 负责纯 PyTorch 训练。README 自称「~9.2K LOC of RL code」,并强调同一份脚本可以从 8B 一路跑到 DeepSeek-V3 级别(--fsdp.ep_size 256),且支持多轮、工具调用、VLM、LLM-as-judge 等任意 Python reward。
项目目前处于 快速迭代但尚未稳定 的阶段:最近三周连发 v0.1.8 / v0.1.9 / v0.1.10(2026-09-14 → 0.9-28),提交热点集中在 tests/unit、.github/workflows、README.md、molt/trainer、examples/python,说明团队在同步补测试、CI 与文档;最近 10 个 commit 里一半是 fix(agents 上的 event-loop 阻塞、rollout router 转换、grading 逻辑、loss 聚合),一半是 docs/ci 打磨,属于「功能基本可用、工程收尾中」的状态。README 没有给自己贴 Experimental / Beta 标签,但 PyPI 包明确写「not a supported install path right now」,且默认依赖 git-pinned 的 AutoModel 特定 commit(8f73178ca),因此实际成熟度介于「可用」与「研究预览」之间。
与同类(OpenRLHF / verl / slime)相比,Molt 的差异点是:训练栈坚持 PyTorch/FSDP2 原生而非 Megatron-LM,runtime 坚持 Ray + vLLM 而非自研,且把 agent 接口抬到一等公民(Gymnasium 对齐的 Env.step() / ChatAgent.run()),让环境与 reward 用纯 Python 写、trainer 不动。它同时维护了一个 automodel-slim 分支,把上游 AutoModel 从 758 文件 / 282k 行裁到 252 文件 / ~97k 行,只保留 molt 实际用到的模型族与并行栈,作为可选轻量后端。
解决什么问题、给谁用 Molt 解决的是 大规模 agentic RL 训练的基础设施碎片化 问题:现有 RL 训练栈要么绑定 Megatron-LM 这类重型后端、要么 rollout 与训练紧耦合、要么对多轮 / 工具 / VLM 等 agentic 场景支持薄弱。Molt 把栈收敛成 Ray + vLLM + AutoModel/FSDP2 三件套,用 token-first 的 contract(token ids / logprobs / action ranges / rewards / 多模态张量全程对齐)统一多轮、多模态、工具调用轨迹;同时通过 fully-async 队列、partial rollout、weight sync 重叠,让 DeepSeek-V3 级别的 MoE actor 在 vLLM rollout 下不被饿死。它面向的是需要 在纯 PyTorch 里 hack 模型和环境、又要跑到百卡级 MoE 的研究团队。
目标用户是做 LLM / 多模态 agentic RL 研究 的算法研究员与系统工程师,典型场景包括:
- 用 GRPO / RLOO / REINFORCE-baseline / FlashREINFORCE / on-policy distillation 等算法训对话、数学、代码、工具-use agent;
- 在 Env.step() 或 ChatAgent.run() 里写自定义 grader、多轮工具、VLM 环境、LLM-as-judge,reward 用任意 Python;
- 从 8B Dense 一路扩展到 1T-class MoE(DeepSeek-V3、Qwen3.8 等),不改脚本只改 --fsdp.ep_size / TP / CP;
- 用 vLLM 做 rollout 并需要 partial rollout、async queue、rollout dump/replay、weight-update 覆盖检查等调试能力;
- 用 molt.cli.train_sft 做与 RL 同模型加载路径的 SFT 预训练 / 蒸馏。
对「资深 GPU kernel / 训练性能工程师」而言,切入点主要在 rollout-trainer 重叠、MoE dispatch(DeepEP/HybridEP)、FSDP2 并行策略、以及 vLLM 权重同步这些系统层。
同类项目与差别 OpenRLHF :基于 Ray + vLLM,但训练后端走 DeepSpeed / Megatron,模型并行栈更重;Molt 坚持 PyTorch/FSDP2 + NVIDIA AutoModel 原生,且 agent 接口是一等公民。verl :字节跳动出品,同样 Ray + vLLM,但训练栈绑定 Megatron-LM / Megatron-LM 风格的并行,生态更偏工业规模;Molt 走纯 PyTorch/FSDP2,自称 ~9.2K LOC 更易 fork。slime :也是 vLLM-based RL,但更偏 Megatron 路线;Molt 强调 AutoModel-first、token-first contract、以及 agentic-first 的 Gymnasium 对齐 API。TRL (HuggingFace) :HuggingFace 官方 RL 库,生态最广、模型支持最全,但默认不是 fully-async、不原生面向 1T-class MoE + vLLM rollout 规模;Molt 在 agentic + 大规模 MoE 上更激进。
核心能力
阶段: 快速迭代的研究级框架:v0.1.8–v0.1.10 三周三连发,功能基本可用但工程仍在收尾(测试/CI/文档提交密集),PyPI 包暂不支持,默认依赖 git-pinned AutoModel commit。
技术栈: Python 3.10+;PyTorch 2.13+cu130(CUDA 13 硬依赖);Ray(runtime / placement / async queue);vLLM(rollout);NVIDIA AutoModel(nemo-automodel,git-pinned 或 automodel-slim 分支);FSDP2 + TransformerEngine attention(THD packing);DeepEP / HybridEP(MoE dispatch);flash-attn、mamba、TransformerEngine、Dion/Muon 优化器;setuptools + pyproject.toml 构建;pytest(unit/integration/system/acceptance 四级 marker)+ ruff + black + isort + pre-commit;GitHub Actions CI + Jenkins(部分 job 在 Jenkins);Docker(多架构 amd64+arm64,CUDA forward-compat)
规模: 自称 ~9.2K LOC RL 代码;仓库目录树显示 molt/ 55 入口、tests/ 38(含 e2e 与 unit)、examples/ 23、.github 19;最近提交热点 tests/unit(7)、.github/workflows(5)、README.md(5)、molt/trainer(3)、examples/python(3);Release v0.1.8 / v0.1.9 / v0.1.10 于 2026-09-14 / 09-17 / 09-28 三连发,最近提交 2026-10-04;AutoModel 上游 758 文件 / 282k 行,slim 分支裁到 252 / ~97k;Docker 镜像 hijkzzz/molt:latest。
三、本地跑起来(没有 GPU 的 Mac) 安装 # 1. 克隆仓库(默认分支 main,当前最新 tag v0.1.10)git clone https://github.com/NVIDIA-NeMo/labs-molt.gitcd labs-molt# 2. 创建 Python 3.10+ 虚拟环境(Apple Silicon 推荐 conda / venv)python3.11 -m venv .venv && source .venv/bin/activate# 3. 安装 PyTorch CPU/MPS 版本(Mac 上没有 CUDA,必须覆盖 setup.py 里的 cu130 pin)# 证据:README 要求 torch==2.13.0+cu130,但 PyPI 现在提供 2.13.0 mac 版;# 这里先装非 CUDA 版本,避免 pip 去拉 cu130 wheel 失败。pip install torch==2.13.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cpu# 如果希望走 MPS 后端(部分算子仍会回退 CPU):# pip install torch==2.13.0 torchvision torchaudio# 4. 安装 molt 本身(跳过 vllm 与 AutoModel 的 GPU 依赖)# 证据:setup.py 中 AUTOMODEL 默认 git-pinned 上游 commit,requirements.txt 含 vllm;# vLLM 0.21.0 在 macOS 上无 wheel 且依赖 CUDA,必须跳过。# 方法 A:只装 molt 核心(不装 [vllm] extra)pip install -e .# 方法 B:如果 requirements.txt 里 vllm 是硬依赖导致失败,则手动裁剪:# grep -v '^vllm' requirements.txt > requirements-nogpu.txt# pip install -r requirements-nogpu.txt# pip install -e .# 5. 安装 AutoModel(可选,仅当你需要跑含训练/rollout 的 e2e 时)# 证据:setup.py AUTOMODEL 字典默认 'upstream' 指向 git commit 8f73178ca;# AutoModel 本身依赖 TransformerEngine / flash-attn 等 GPU 库,Mac 上大概率装不上。# 纯读代码 / 跑 CPU-only 单元测试可跳过此步。# pip install nemo-automodel @ git+https://github.com/NVIDIA-NeMo/Automodel.git@8f73178ca51d4c1e55ccf05df5da6540a9e24f7e# 6. 安装 pre-commit(CONTRIBUTING.md 要求,PR 会被卡住)pip install pre-commitpre-commit install# 7. 验证基础 import 链(不触发 CUDA 初始化)python -c "import molt; from molt.agents import Env, ChatAgent; print('import ok')"
哪些路径能真跑 ['✅ 可在 CPU / MPS 上真正执行(无需 GPU,不触发 CUDA):', ' - molt/agents/base.py — Env / ChatAgent / Result / Trajectory 抽象基类,纯 Python + torch 张量', ' - molt/agents/__init__.py — 公开接口导出', ' - molt/datasets/prompts_dataset.py / sft_dataset.py / utils.py — 数据集加载与 tokenize(用 HF tokenizer,CPU 可用)', ' - molt/models/loss.py — policy loss / value loss 计算(torch 算子,CPU 可跑)', ' - molt/models/actor.py / critic.py / base.py / utils.py — 模型定义(forward 可在 CPU 跑小模型)', ' - molt/trainer/algorithm/ — advantage / kl_controller / experience / replay_buffer(纯 torch 张量)', ' - molt/trainer/fsdp/packing.py — THD packing 逻辑(test_cp_thd_packing.py 在 CI 里 CPU 跑)', ' - molt/trainer/rollout/router.py、samples_generator.py、experience_maker.py — 路由与经验构造(不依赖 vLLM 实例化)', ' - molt/utils/ — config / logging / distributed_sampler / seqlen_balancing / vlm_utils', ' - molt/cli/common_args.py — argparse 定义', ' - examples/python/agents/ — math.py / geo3k.py / chat_minimal.py / chat_geo3k.py(agent 逻辑本身可 import 看)', ' - examples/python/utils/ — math_grader.py / prepare_dapo.py / prepare_geo3k.py', ' - tests/unit/ 大部分测试 — 见 smoke_test 与 test_suite', '', '⚠️ 只能读代码 / 静态分析,无法在 Mac 上真正执行(依赖 CUDA / vLLM / NCCL):', ' - molt/cli/train_rl_ray.py — 入口会 ray.init() + create_vllm_engines(),需要 GPU placement group', ' - molt/cli/train_sft.py — torchrun + FSDP2,需要 CUDA', ' - molt/trainer/rl_trainer.py — 主训练循环,依赖 Ray actor + vLLM + FSDP2', ' - molt/trainer/sft_trainer.py — 同上', ' - molt/trainer/placement.py — Ray placement group → GPU bundle 映射', ' - molt/trainer/vllm/vllm_engine.py + vllm_worker_wrap.py — vLLM 引擎封装,硬依赖 CUDA(_MIN_VLLM_VERSION=0.21.0)', ' - molt/trainer/workers/actor_group.py / policy_actor.py / critic_actor.py — Ray actor,内部 CUDA 操作', ' - molt/trainer/fsdp/checkpoint.py / refit.py / strategy.py / optimizer_offload.py / muon.py — FSDP2 + TransformerEngine + CPU offload,需 GPU', ' - molt/trainer/rollout/router.py 中与 vLLM 交互的运行时路径(import 可看,运行需引擎)', ' - tests/e2e/ — 端到端测试,需多 GPU', '', '🔶 需 Colab 免费 T4 验证(单卡几十分钟):', ' - 任何含 --vllm.num_engines 1 的 RL quick start', ' - SFT torchrun --nproc_per_node=1 的小模型实验', ' - tests/unit/test_vllm_engine.py(需 vLLM 实例化)']
最小可运行 # 冒烟测试 1:验证公开 agent 接口可 import 且抽象方法签名正确(CPU,<10s)python -c "from molt.agents import Env, ChatAgent, ChatAgentRunner, ChatContext, Result, StepEnvRunner, Trajectory; print('agents import ok')"# 冒烟测试 2:跑 agents 基类单元测试(CPU,纯 Python 逻辑,~30s)python -m pytest tests/unit/test_agents_base.py -v -m unit# 冒烟测试 3:跑 chat agent 单元测试(CPU,~1min)python -m pytest tests/unit/test_chat_agent.py -v -m unit# 冒烟测试 4:跑 rollout router 单元测试(CPU,不依赖 vLLM 实例化,~1min)python -m pytest tests/unit/test_router.py tests/unit/test_routing_replay.py -v -m unit# 冒烟测试 5:跑算法层单元测试(advantage / KL / experience / loss,CPU,~2min)python -m pytest tests/unit/test_gae_value_loss.py tests/unit/test_group_advantage.py tests/unit/test_kl_controller.py tests/unit/test_experience.py tests/unit/test_policy_loss.py tests/unit/test_global_token_loss.py -v -m unit# 冒烟测试 6:跑 FSDP packing 单元测试(CPU,THD packing 逻辑,~30s)python -m pytest tests/unit/test_cp_thd_packing.py tests/unit/test_fsdp_packing.py -v -m unit
测试 ['# 测试框架:pytest(pyproject.toml [tool.pytest.ini_options] 配置)', '# addopts = "--verbose --pyargs --durations=0 --strict-markers"', '# testpaths = ["./tests"]', '# markers: unit / integration / system / acceptance / docs / skipduringci / pleasefixme', '#', '# 目录结构:', '# tests/unit/ — 38 个文件,CPU 可跑(主要目标)', '# tests/e2e/ — 仅 summarize_metrics.py,需 GPU', '#', '# 跑全量 CPU 单元测试(Mac 上推荐,约 5-15min,取决于模型加载):', 'python -m pytest tests/unit -m unit -v --timeout=120', '#', '# 只跑纯 CPU、不加载 HF 模型的子集(最快,<2min):', 'python -m pytest tests/unit/test_agents_base.py tests/unit/test_router.py tests/unit/test_routing_replay.py tests/unit/test_cp_thd_packing.py tests/unit/test_fsdp_packing.py tests/unit/test_muon_param_classify.py tests/unit/test_logging_utils.py tests/unit/test_prompts_dataset.py tests/unit/test_sft_dataset.py -m unit -v', '#', '# 跳过已知 broken 测试(marker pleasefixme):', 'python -m pytest tests/unit -m "unit and not pleasefixme" -v', '#', '# 跑单个测试文件示例:', 'python -m pytest tests/unit/test_chat_agent.py -v -m unit', '#', '# 覆盖率(需 pip install pytest-cov):', 'python -m pytest tests/unit -m unit --cov=molt --cov-report=term-missing']
调试 # 1. 日志开关:molt 使用 molt/utils/logging_utils.py 里的 init_logger # 设置环境变量控制 verbosity: export LOG_LEVEL=DEBUG # 默认 INFO # 或在代码中:from molt.utils.logging_utils import init_logger; logger = init_logger(__name__, level='DEBUG') # 2. 断点入口推荐(按层): # - Agent 层:molt/agents/base.py 的 StepEnvRunner.run() / ChatAgentRunner.run() # - 调度层:molt/trainer/rl_trainer.py 的 train() 主循环 # - Rollout 层:molt/trainer/rollout/router.py 的 route() / samples_generator.py 的 generate() # - 训练层:molt/trainer/workers/policy_actor.py 的 update_policy() # - 算法层:molt/trainer/algorithm/advantage.py 的 compute_advantage() # 3. 关键环境变量(README + train_rl_ray.py _ray_runtime_env_vars()): export NCCL_DEBUG=WARN # NCCL 日志级别 export TOKENIZERS_PARALLELISM=true export RAY_ENABLE_ZERO_COPY_TORCH_TORCH_TENSORS=1 export VLLM_WORKER_MULTIPROC_METHOD=spawn # 多进程方式 export TORCH_COMPILE_DISABLE=1 # 禁用 torch.compile 便于调试 export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:base_size:4096 # CUDA 内存分配 # 4. Profiling 入口: # - molt/trainer/rl_trainer.py 主循环内可插入 torch.profiler # - molt/trainer/vllm/vllm_engine.py 的 generate() 可加 vLLM stats 打印 # - 推荐在 Colab T4 上用 torch.profiler.record_function() 标记关键段 # 5. 配置打印:strategy.print(args) 在 train_rl_ray.py 里会 dump 全部参数, # 调试时可在 ray.init() 前加 breakpoint() 看解析后的 args 命名空间。 # 6. Ray 本地调试(Mac 上只能单节点,无 GPU): # ray.init(num_cpus=4, resources={'mock': 1}) # 模拟资源,不跑真实 actor CI ['# CI 系统:GitHub Actions(.github/workflows/,commit_hotspots 显示 5 次提交在此)', '#', '# 主要 workflow(根据目录推断,需验证具体文件名):', '# - python-package.yml 或类似:PyPI 构建(MOLT_PYPI_BUILD=1)', '# - cicd.yml 或类似:PR 检查(copyright + unit tests)', '# - docs.yml 或类似:文档站构建', '#', '# PR 会被以下检查卡住(证据:CONTRIBUTING.md + 最近 commit):', '# 1. pre-commit hooks:black / ruff / isort(line_length=119)', '# 2. Copyright header check(最近 commit 74573e45 修复 docs-only push 不失败)', '# 3. DCO sign-off:git commit -s 必须有 Signed-off-by 行', '# 4. Unit tests:至少 tests/unit 子集在 CI 容器里跑(CUDA 环境)', '# 5. Marker 检查:--strict-markers 要求所有 @pytest.mark.X 必须在 pyproject.toml 注册', '#', '# 本地预检(提交前必跑):', 'python -m compileall -q molt examples/python tests # AGENTS.md 要求', 'pre-commit run --all-files # 格式 + copyright', 'python -m pytest tests/unit -m "unit and not pleasefixme" -q # 快速回归']
坑 # 坑 1:PyPI 包不是官方支持路径 # 证据:README 明确写 'molt-rl on PyPI ... is not a supported install path right now'; # setup.py 中 MOLT_PYPI_BUILD=1 会放松 AutoModel pin 到 >=0.5.0,但默认 git pin 是 8f73178ca。 # 不要 pip install molt-rl,用 pip install -e . 或容器。 # 坑 2:vLLM 是硬依赖但 macOS 无 wheel # 证据:vllm_engine.py 要求 vLLM>=0.21.0,而 vLLM 0.21.0 无 macOS wheel。 # 如果 pip install -e '.[vllm]' 失败,改用 pip install -e . 并手动装 requirements.txt 中除 vllm 外的依赖。 # 坑 3:AutoModel 上游依赖 TransformerEngine / flash-attn / DeepEP # 证据:setup.py AUTOMODEL 字典指向 git commit,这些库需要 CUDA 编译。 # 在 Mac 上不要尝试 pip install nemo-automodel,除非走 MOLT_AUTOMODEL=slim 分支(仍需验证)。 # 坑 4:torch==2.13.0+cu130 在 macOS 不存在 # 证据:README 要求 CUDA 13 + torch 2.13 cu130。 # macOS 必须覆盖为 torch 2.13.0 CPU 或 MPS 版本,否则 pip 解析失败。 # 坑 5:tests/unit 里部分测试会尝试加载 HF 模型 # 证据:test_actor_forward_output.py、test_critic_trainer.py 等文件名暗示需要模型权重。 # 在 Mac 上跑前先检查是否需 --model.model_name_or_path,可用小模型(如 Qwen2.5-0.5B)或 mock。 # 坑 6:Ray 在 macOS 上行为差异 # 证据:train_rl_ray.py 用 ray.util.placement_group + GPU bundle。 # Mac 上 ray.init() 可运行,但 create_vllm_engines() 会失败(无 GPU placement group)。 # 本地调试建议 mock Ray actor 或只跑非 rollout 路径。 # 坑 7:FSDP2 相关测试在 CPU 上可能数值不一致 # 证据:test_fsdp_packing.py 与 test_cp_thd_packing.py 测试 packing 逻辑,但 FSDP2 的 # distribute_mesh / apply 需要 CUDA 上下文。纯 packing 算法可跑,涉及通信的跳过。 # 坑 8:pre-commit 的 copyright header 检查会卡 PR # 证据:最近 commit 74573e45 'copyright-check: a docs-only push is not a failure' 说明 # 该检查存在且会失败。新文件必须加 SPDX-FileCopyrightText 头(仓库统一 Apache-2.0)。 # 坑 9:tests/unit 中 pleasefixme marker 的测试会失败 # 证据:pyproject.toml markers 列表含 'pleasefixme: marks tests that are broken and need fixing'。 # 跑全量测试时请加 -m "unit and not pleasefixme" 避免误判。 # 坑 10:MOLT_AUTOMODEL=slim 分支的 CI 独立 # 证据:README 提到 'the branch runs its own seven-minute CI on molt's image and runners'。 # 如果你基于 automodel-slim 分支开发,PR 需打 cicd label,且 review 用 /review 而非 /claude review。 七、任务卡 任务 1 优先 · medium · Mac · 2–3 个晚上
test(agents): CPU-only conftest fixtures + mark GPU-only agent tests Issue #51 · Are community contributions welcome? ↗
无人认领 3 条评论 更新 2026-07-30
用到的专长: 测试基础设施与 Python 调度逻辑,CPU 可验证。
目标: 在 tests/unit/conftest.py 新增 cpu-only / gpu-only fixture 与 marker,把现有 agent 测试中依赖 CUDA/vLLM 的用 skipif 隔离,确保 Mac 上 pytest tests/unit -m cpu_only 全绿。
为什么值得长期做: agents 是 Molt agentic-first 定位的门面,commit 热点 #185/#186/#191 都在修 event-loop 阻塞。CPU/GPU 测试分离是所有后续 agent 测试的基础设施。
怎么介入: Issue #51 是社区贡献意愿确认,无 assignee、无 open PR;借其「start small (docs, tests)」语境切入测试基建是得体的。第一个 PR 的边界: 仅限 tests/unit/conftest.py(新增)与现有 agent 测试加 marker,不动 molt/ 源码。
第一步: 调查 tests/unit/conftest.py 是否存在、现有 agent 测试文件(test_agents_base.py、test_chat_agent.py、test_chat_server.py)里哪些 import 或 fixture 触发 CUDA。
本机怎么复现 / 验证: 在 Mac 上 pip install -e ".[vllm]" 后 pytest tests/unit -m cpu_only --collect-only 看 marker 是否生效;pytest tests/unit/test_agents_base.py 验证 agent 测试可跑。
认领留言(英文,可直接贴到 Issue) Hi — I'd like to pick up the CPU-first test infra from the 'start small (docs, tests)' spirit in this issue. Plan: add tests/unit/conftest.py with cpu_only/gpu_only markers, mark CUDA/vLLM-dependent agent tests with skipif, and add a smoke test that imports molt.agents on CPU. No changes to molt/ source. Does cpu_only/gpu_only marker naming look good, or do you prefer requires_cuda? I'll have a PR up within a couple of evenings. 复制留言
大致实施方案 在 tests/unit/conftest.py(新增)里定义 cpu_only / gpu_only marker 与 is_cuda_available 检查 在现有 agent 测试中给依赖 vLLM/CUDA 的测试函数加 @pytest.mark.gpu_only + skipif(not is_cuda_available) 在 tests/unit 下新增 test_import_agent_cpu.py 验证 molt.agents.base / molt.agents.chat_agent 可在 CPU import 本地跑 pytest tests/unit -m cpu_only --collect-only 与 pytest tests/unit -m cpu_only 确认 提交 PR,scope 仅限 tests/unit 与 conftest,不动 molt/ 源码 可能涉及的目录或文件 tests/unit/conftest.py(新建) tests/unit/test_agents_base.py tests/unit/test_chat_agent.py tests/unit/test_chat_server.py
验收方式 pytest tests/unit -m cpu_only 在 Mac 上全绿pytest tests/unit -m gpu_only 在 Mac 上全 skip、不报错CI 中现有 GPU 测试不受影响(marker 不改变其行为) 开工前问题与风险 向维护者确认 是否偏好 cpu_only / gpu_only 命名,还是沿用 requires_cuda 之类? conftest.py 是否应放在 tests/unit/ 还是根 tests/? 风险 若 误 把 非 G P U 测 试 标 为 g p u _ o n l y 会 降 低 覆 盖 率 ; 需 逐 文 件 确 认 i m p o r t 链 。 任务 2 可选 · medium · Mac · 2–3 个晚上
test(vlm): unit tests for vlm_utils image placeholder expansion Issue #104 · [feature] OrcaRouter provider support for Molt ↗
无人认领 0 条评论 更新 2026-09-07
用到的专长: 多模态 token 对齐是训练性能工程师熟悉领域,纯函数测试 CPU 可验证。
目标: 新建 tests/unit/test_vlm_utils.py(新增),用 PIL.Image.new 或 torch.zeros 构造假图片,覆盖 process_prompt_with_images 与 estimate_vllm_input_expansion_delta 的纯函数逻辑。
为什么值得长期做: VLM 是 README 主打卖点之一(multi-turn/VLM/tool-call 共享格式),vlm_utils.py 是 token 对齐关键函数,但 tests/unit/ 下无对应测试。
怎么介入: Issue #104 本身与 OrcaRouter 相关,但 vlm_utils 测试是其「OpenAI-compatible provider」无关的独立缺口;在 Issue 下简短评论说明认领的是测试子任务即可。第一个 PR 的边界: 仅新增 tests/unit/test_vlm_utils.py,不动 molt/ 源码。
第一步: 阅读 molt/utils/vlm_utils.py 与 molt/agents/base.py 中 _tokenize_observation 的调用方式,确定函数签名与期望输出。
本机怎么复现 / 验证: 在 Mac 上 pytest tests/unit/test_vlm_utils.py -v 运行新增测试;若无该文件则先确认 molt/utils/vlm_utils.py 在 CPU 上可 import。
认领留言(英文,可直接贴到 Issue) I'd like to add unit tests for molt/utils/vlm_utils.py — it's currently untested and is on the VLM token-alignment path the README highlights. Plan: create tests/unit/test_vlm_utils.py covering process_prompt_with_images and estimate_vllm_input_expansion_delta with fake images (PIL/torch.zeros), all CPU-runnable. No changes to molt/ source. Any implicit assumptions on image shape/format I should encode? PR ready in a couple of evenings. 复制留言
大致实施方案 在 tests/unit/test_vlm_utils.py(新增)中构造 fake image tensor / PIL image 测试 process_prompt_with_images 在不同 placeholder 模式下的 token 展开 测试 estimate_vllm_input_expansion_delta 在空/非空 image list 下的返回值 本地跑 pytest tests/unit/test_vlm_utils.py -v 确认全绿 提交 PR,scope 仅限新增测试文件 可能涉及的目录或文件 tests/unit/test_vlm_utils.py(新建) molt/utils/vlm_utils.py(只读)
验收方式 pytest tests/unit/test_vlm_utils.py 在 Mac 上全绿覆盖 placeholder 展开与 delta 估算两个函数的主要分支 开工前问题与风险 向维护者确认 process_prompt_with_images 是否对 image 尺寸/格式有隐含假设需要在测试里显式约束? 风险 若 函 数 内 部 隐 式 依 赖 C U D A t e n s o r , 需 确 认 其 在 C P U 上 可 跑 ; 否 则 调 整 测 试 输 入 类 型 。 任务 3 可选 · medium · Mac · 2–3 个晚上
test(agents): DistillationEnv / chat_minimal env 单元测试 Issue #95 · [feature] Quickstart scripts missing prepare_math ↗
已有 PR #96 0 条评论 更新 2026-08-24
用到的专长: agent contract 测试是纯 Python 逻辑,CPU 可验证。
目标: 新建 tests/unit/test_chat_minimal.py(新增)与 tests/unit/test_distill_agent.py(新增),覆盖 Result 形状、terminated 默认值、reward=0.0 placeholder 等纯 Python 逻辑。
为什么值得长期做: examples/python/agents/ 下 chat_minimal.py / distill_agent.py 无对应 unit test,是 agent 层测试覆盖缺口;与 #51 的「start small (tests)」一致。
怎么介入: Issue #95 已有 open PR #96 处理 prepare_math 脚本,但测试覆盖是独立缺口;在 #95 下简短评论说明认领的是 agent 测试子任务即可。第一个 PR 的边界: 仅新增 tests/unit/test_chat_minimal.py 与 tests/unit/test_distill_agent.py,不动 molt/ 或 examples/ 源码。
第一步: 阅读 examples/python/agents/chat_minimal.py、molt/agents/distill_agent.py,确定 Env 接口与 Result 字段。
本机怎么复现 / 验证: 在 Mac 上 pytest tests/unit/test_chat_minimal.py tests/unit/test_distill_agent.py -v 运行新增测试;若文件不存在则先确认 examples/python/agents/chat_minimal.py 在 CPU 上可 import。
认领留言(英文,可直接贴到 Issue) I'd like to add unit tests for the example agents that currently lack coverage — chat_minimal.py and DistillationEnv. Plan: create tests/unit/test_chat_minimal.py and tests/unit/test_distill_agent.py covering Result shape, terminated defaults, and reward placeholders. All CPU-runnable, no changes to molt/ source. Do these envs have any hidden external LLM dependencies I should mock? PR ready in a couple of evenings. 复制留言
大致实施方案 在 tests/unit/test_chat_minimal.py(新增)中实例化 env、调用 step、断言 Result 字段类型与 terminated 默认值 在 tests/unit/test_distill_agent.py(新增)中覆盖 DistillationEnv 的基本轨迹生成 本地跑 pytest tests/unit/test_chat_minimal.py tests/unit/test_distill_agent.py -v 提交 PR,scope 仅限新增测试文件 可能涉及的目录或文件 tests/unit/test_chat_minimal.py(新建) tests/unit/test_distill_agent.py(新建) examples/python/agents/chat_minimal.py(只读) molt/agents/distill_agent.py(只读)
验收方式 pytest tests/unit/test_chat_minimal.py tests/unit/test_distill_agent.py 在 Mac 上全绿覆盖 Result 形状、terminated 默认、reward placeholder 开工前问题与风险 向维护者确认 chat_minimal env 是否需要 mock 外部 LLM 调用?若是,确认是否用 unittest.mock 即可在 CPU 上跑。 风险 若 e n v 内 部 隐 式 依 赖 v L L M 实 例 化 , 需 确 认 测 试 路 径 可 跳 过 该 依 赖 。 任务 4 可选 · medium · Mac · 3–4 个晚上
test(rollout): rollout dump→replay token-exact parity 测试 Issue #10 · [feature] Roadmap for DeepSeek-V4 RL training support (DSA/MLA kernels, mixed precision, GPU footprint, and the agentic environment path) ↗
已有 PR #132 2 条评论 更新 2026-07-31
用到的专长: rollout 数据通路是训练性能工程师熟悉领域,纯逻辑测试 CPU 可验证。
目标: 新建 tests/unit/test_dump_replay_parity.py(新增),用 mock 或 CPU tensor 验证 dump→replay 的 token ids / logprobs / rewards round-trip 一致性。
为什么值得长期做: README 已暴露 --train.rollout_dump_dir / --rollout_replay_dir 标志,但无端到端验证脚本;router/experience_maker 数据通路是 CPU 可测的纯逻辑。
怎么介入: Issue #10 有 open PR #132 处理 Qwen-vLLM parity,但 dump/replay 测试是独立缺口;在 #10 下简短评论说明认领的是测试子任务即可。第一个 PR 的边界: 仅新增 tests/unit/test_dump_replay_parity.py,不动 molt/ 源码。
第一步: 阅读 molt/trainer/rollout/router.py、experience_maker.py 与 tests/unit/test_routing_replay.py,确定 dump/replay 数据结构与序列化路径。
本机怎么复现 / 验证: 在 Mac 上 pytest tests/unit/test_dump_replay_parity.py -v 运行新增测试;若文件不存在则先确认 molt/trainer/rollout/experience_maker.py 在 CPU 上可 import。
认领留言(英文,可直接贴到 Issue) I'd like to add a dump→replay parity test for the rollout path — README exposes --train.rollout_dump_dir / --rollout_replay_dir but there's no end-to-end verification. Plan: create tests/unit/test_dump_replay_parity.py that constructs fake Experience (token ids/logprobs/rewards), runs dump then replay, and asserts bit-exact round-trip. All CPU-runnable with mock tensors. Does the dump path assume GPU tensors or a specific filesystem I should accommodate? PR ready in a couple of evenings. 复制留言
大致实施方案 在 tests/unit/test_dump_replay_parity.py(新增)中构造 fake Experience(token ids / logprobs / rewards) 调用 dump 路径序列化到临时目录,再走 replay 路径反序列化 断言 round-trip 后 token ids / logprobs / rewards 完全一致 本地跑 pytest tests/unit/test_dump_replay_parity.py -v 提交 PR,scope 仅限新增测试文件 可能涉及的目录或文件 tests/unit/test_dump_replay_parity.py(新建) molt/trainer/rollout/router.py(只读) molt/trainer/rollout/experience_maker.py(只读)
验收方式 pytest tests/unit/test_dump_replay_parity.py 在 Mac 上全绿round-trip 后 token ids / logprobs / rewards 完全一致 开工前问题与风险 向维护者确认 dump/replay 是否依赖特定文件系统或 GPU tensor?若是,确认是否可用 CPU tensor + tempfile 替代。 风险 若 d u m p 路 径 隐 式 依 赖 C U D A t e n s o r , 需 调 整 测 试 输 入 为 C P U t e n s o r 。 任务 5 可选 · easy · Mac · 1–2 个晚上
docs: CPU-first 开发指南 + CONTRIBUTING 更新 Issue #51 · Are community contributions welcome? ↗
无人认领 3 条评论 更新 2026-07-30
用到的专长: 文档与贡献指南是低风险高可见度贡献。
目标: 新建 docs/dev-cpu.md(新增)并在 CONTRIBUTING.md 加「CPU-only 贡献」指引,说明哪些 marker 是 CPU-safe、如何在 Mac 上跑测试。
为什么值得长期做: README 安装章节只给 container 与 pip install -e '.[vllm]' 两条 GPU 路径,无 Apple Silicon / CPU-only 文档;CONTRIBUTING.md 只提 sign-off + pre-commit,没说无 GPU 怎么跑测试。
怎么介入: Issue #51 是社区贡献意愿确认,无 assignee、无 open PR;借其「start small (docs, tests)」语境切入文档是得体的。第一个 PR 的边界: 仅新增 docs/dev-cpu.md 与 CONTRIBUTING.md 追加,不动代码。
第一步: 确认 docs/ 目录结构与 CONTRIBUTING.md 当前内容,结合前几张测试卡的经验整理 CPU-first 流程。
本机怎么复现 / 验证: 在 Mac 上 pre-commit run --all-files 确认格式;cat docs/dev-cpu.md 检查内容。
认领留言(英文,可直接贴到 Issue) Following the 'start small (docs, tests)' spirit in this issue, I'd like to add a CPU-first development guide. Plan: create docs/dev-cpu.md documenting how to run tests on Mac (cpu_only marker, install steps), and add a 'CPU-only contributions' section to CONTRIBUTING.md referencing the marker convention. Pure docs, no code changes. Any preferred doc format (md vs rst)? PR ready in an evening or two. 复制留言
大致实施方案 在 docs/dev-cpu.md(新增)中写明 Mac 上 pip install -e ".[vllm]" 后的测试命令、cpu_only marker 用法 在 CONTRIBUTING.md 加一节「CPU-only contributions」,引用前几张测试卡建立的 marker 约定 本地 pre-commit run --all-files 确认格式合规 提交 PR,scope 仅限 docs/ 与 CONTRIBUTING.md 可能涉及的目录或文件 docs/dev-cpu.md(新建) CONTRIBUTING.md(追加)
验收方式 pre-commit run --all-files 通过文档中命令在 Mac 上可复现 开工前问题与风险 向维护者确认 docs/ 目录是否偏好特定格式(如 markdown 还是 rst)? 风险