一、项目定位 🤗 Diffusers 是 Hugging Face 维护的 PyTorch 扩散模型推理与训练库,定位为 image / video / audio / 3D 分子生成的「模块化工具箱」。README 明确把哲学写在最前面:usability over performance、simple over easy、customizability over abstraction,因此它不是追求极致 kernel 性能的项目,而是追求可组合、易扩展、易贡献。核心抽象三层:DiffusionPipeline(端到端推理入口)、Scheduler(噪声调度与步进)、Model(UNet / DiT / VAE 等可替换积木),外加可插拔的 LoRA / ControlNet / Adapter 机制。
从证据看,这是一个成熟的中年期项目:根目录 1047 个 src 文件、615 个测试文件、509 个文档文件、95 个转换脚本、464 个 examples;最近 release 是 0.40.0(2026-08-20),再往前 0.39.0、0.38.0,约 6–8 周一个 minor,节奏稳定;最近提交热点集中在 src/diffusers(109 次)、examples/community(12)、docs(5)、tests/models(4)、workflows(3),说明仍在活跃地加新模型/新 pipeline,同时补测试和文档。
同类项目里,它是「生态最大、模型最多、社区最活跃」的那一个,但性能不是第一位——这正是 GPU kernel / 训练性能工程师可以切入的空间。README 明确写了 Apple Silicon / MPS 支持,意味着 CPU / Apple Silicon 上可以跑通很多流程,只是速度慢;真正要验证性能仍需 CUDA 机器(可用 Colab T4 做短时确认)。
解决什么问题、给谁用 解决的具体问题:把学术界和工业界不断扩散(pun intended)的扩散模型变体统一到一个可组合、可训练、可推理的框架里,让用户用几行代码跑 SOTA,同时让研究者能快速把新论文落地成 pipeline / model / scheduler。
目标用户和典型场景:
1. 应用开发者:用 DiffusionPipeline.from_pretrained 几行代码做 text-to-image / video / audio 推理,部署到 Hub、ONNX、TensorRT;
2. 研究者 / 训练工程师:用 examples/ 和 training scripts 做 DreamBooth、LoRA、ControlNet、一致性蒸馏、RL 训练;
3. 模型作者:把自家 checkpoint 通过 scripts/convert_*.py 转成 diffusers 格式,或新增 pipeline / model / scheduler 合入核心库;
4. 优化工程师(你的切入点):在 optimization/ 和 benchmarks/ 下做 FP16 / FP8、tensor parallelism、kernel 融合、推理加速——这些正是当前提交热点里反复出现的关键词(FP8、tensor-parallel、single-file、MPS)。
同类项目与差别 CompVis / StableDiffusion(原始 SD 代码) :只覆盖 SD 一代,没有统一 pipeline / scheduler / model 抽象,不维护多模型生态;diffusers 是它的「超集 + 工程化继任者」Stability-AI / generative-models(SDXL / SVD 官方仓库) :聚焦自家模型,训练和耦合较紧;diffusers 把同一类模型抽象成可插拔积木,模型数量远超单一厂商仓库NVIDIA / Megatron-DiT / 官方扩散实现 :偏研究与单模型复现,工程通用性弱;diffusers 胜在 Hub 生态、社区贡献、统一 APIvllm / SGLang(LLM 推理加速栈) :同属 HuggingFace 生态但面向 LLM;diffusers 是扩散模型侧的对应物,但当前性能优化深度不如 vllm 对 LLM 的极致 kernel 优化——这正是机会HuggingFace / transformers :兄弟项目,共享 Hub / 训练 / 推理基础设施;diffusers 把同一套哲学搬到扩散模型领域,但模型形态(UNet/DiT + VAE + Scheduler)完全不同
核心能力 能力 在哪 成熟度 端到端扩散推理 Pipeline(text-to-image / video / audio / inpainting / ControlNet / LoRA 等) src/diffusers/pipelines/ 成熟 可插拔 Scheduler(DDPM / DDIM / Euler / FlowMatch / 等几十种) src/diffusers/schedulers/ 成熟 可替换 Model 积木(UNet2D、UNet3D、DiT、Vae、Transformer2D 等) src/diffusers/models/ 成熟 LoRA / ControlNet / Adapter / IP-Adapter 等可插拔注入机制 src/diffusers/(loaders / peft / controlnet 等子目录) 成熟 训练脚本(DreamBooth、LoRA、ControlNet、一致性蒸馏、RL、advanced training) examples/ 下 dreambooth / text_to_image / controlnet / consistency_distillation / reinforcement_learning / advanced_diffusion_training 成熟 Modular Diffusers(可组合的 pipeline 组件化新范式) src/diffusers/modular_pipelines/、docs/source/en/modular_diffusers 实验 Tensor Parallelism(FP8 / 分布式推理与加载) src/diffusers/(tensor parallel 相关)、最近提交含 FP8 / tensor-parallel 修复 实验 Single-File 支持(单 checkpoint 直接加载,无需多文件目录) src/diffusers/(single_file 相关)、最近提交多次 Add single file support for ... 实验 Apple Silicon / MPS 推理支持 docs/source/en/optimization/mps、README 明确提及 成熟 ONNX / Hub 导出与 CLI 工具 docker/(onnxruntime 镜像)、diffusers-cli(AGENTS.md / CLAUDE.md 提及) 成熟 模型格式转换脚本(数百种 convert_*.py) scripts/(95 个文件,大量 convert_*.py) 成熟 Benchmarking 与 profiling 工具 benchmarks/(benchmarking_flux / sdxl / ltx / wan 等)、examples/profiling 成熟
阶段: 成熟活跃期(maintenance + 持续扩张)。证据:0.38–0.40 三个 minor release 在 4 个月内完成;提交热点仍在加新模型(Krea 2、Minimax H3、Cosmos3、QwenImage 2.1)和新能力(FP8 tensor parallelism、single-file support、modular pipelines);同时在做清理(Remove deprecated code paths)和补测试(tests/models、tests/modular_pipelines)。README 没有标 Experimental,整体是生产可用状态;但 modular pipelines、tensor parallelism、部分新 pipeline 仍在快速迭代,属于「核心稳、前沿动」。
技术栈: Python;PyTorch(核心依赖);HuggingFace Hub / huggingface_hub(模型加载);transformers(文本编码器等);safetensors(权重格式);PEFT(LoRA 训练);accelerate(分布式训练);ruff(lint/format,pyproject.toml 配置);make(style / fix-copies / pre-release);GitHub Actions CI(.github/workflows,46 个文件);uv(推荐虚拟环境,AGENTS.md / CLAUDE.md)
规模: 根目录一级条目 27 个;src/ 1047 个文件;tests/ 615 个文件;docs/ 509 个文件;examples/ 464 个文件;scripts/ 95 个文件;utils/ 31 个文件;.github 46 个文件。最近提交热点:src/diffusers 109 次、examples/community 12 次、docs/source 5 次、tests/models 4 次、workflows 3 次。Release:0.40.0(2026-08-20)、0.39.0(2026-07-03)、0.38.0(2026-05-01),约 6–8 周一个 minor。pushed_at 2026-10-05,仍在活跃推进。star 数未在证据中提供,需验证。
三、本地跑起来(没有 GPU 的 Mac) 安装 # 1. 用 uv 创建虚拟环境(AGENTS.md / CLAUDE.md 推荐方式)uv venv --python 3.10 && source .venv/bin/activate# 2. 先装 PyTorch macOS 版(CPU + MPS),不拉 CUDA 变体pip install torch torchvision pillow # 让 pip 自动选 arm64 wheel# 3. 以 editable 模式装 diffusers,跳过 CUDA-only 的可选依赖pip install -e "." # 不装 [torch] extras,避免被拉进 torch 的 cuda 子包# 4. 显式装 CPU/MPS 推理真正需要的依赖pip install transformers accelerate safetensors huggingface_hub scipy# 5. 验证安装与设备python -c "import diffusers, torch; print(diffusers.__version__); print('MPS', torch.backends.mps.is_available())"# 6. 可选:装测试冒烟用的轻量依赖(不装 xformers / triton,它们需要 CUDA)pip install pytest pytest-xdist filelock numpy"
哪些路径能真跑 ['可在 CPU / MPS 上真正执行(慢但能跑):', '- tests/schedulers/:纯 tensor 调度逻辑,无 CUDA kernel 依赖', '- tests/others/:工具函数、配置、dummy 测试', '- tests/quantization/ 中 CPU 量化路径(需验证具体子目录)', "- examples/inference/ 中纯推理脚本(如 image_to_image.py、inpainting.py),配合 torch.float32 和 device='cpu' 或 'mps'", '- src/diffusers/schedulers/ 全部:DDIM、DDPM、FlowMatchEuler 等可在 CPU 上 step', '- src/diffusers/models/ 中模型 forward:UNet、DiT、VAE 用 PyTorch 原生 op,CPU/MPS 可执行,但大模型在 16GB 上会 OOM', '', '只能读代码或必须上 Colab T4 验证:', "- 任何 pipeline.to('cuda') 调用路径", '- AttnProcessor2_0、xformers、FlashAttention 相关分支(src/diffusers/models/attention_processor.py 中 CUDA-only 路径)', '- torch.compile 的 CUDA graph 回退', '- examples/ 下所有 train_*.py(训练脚本)', '- benchmarks/ 全部(需要真实 GPU 测吞吐)', '- docker/ 下所有 cuda 镜像对应路径', "- scripts/convert_*.py 中硬编码 device = 'cuda' 的脚本(如 convert_amused.py)"]
最小可运行 # 1. 最小 CPU 推理:DDPM 无条件生成(tests 自带 fixture,不下载大模型)python -c "from diffusers import DDPMScheduler, UNet2DModelimport torchmodel = UNet2DModel.from_pretrained('google/ddpm-cat-256').to('cpu')scheduler = DDPMScheduler.from_pretrained('google/ddpm-cat-256')scheduler.set_timesteps(10)noise = torch.randn(1, 3, model.config.sample_size, model.config.sample_size)input = noisefor t in scheduler.timesteps: with torch.no_grad(): noisy_residual = model(input, t).sample prev = scheduler.step(noisy_residual, t, input).prev_sample input = prevprint('DDPM CPU OK, shape:', input.shape)"# 2. MPS 推理(若 Apple Silicon):同模型切到 mpspython -c "from diffusers import DDPMScheduler, UNet2DModelimport torchdevice = 'mps' if torch.backends.mps.is_available() else 'cpu'model = UNet2DModel.from_pretrained('google/ddpm-cat-256').to(device)scheduler = DDPMScheduler.from_pretrained('google/ddpm-cat-256')scheduler.set_timesteps(5)noise = torch.randn(1, 3, model.config.sample_size, model.config.sample_size, device=device)input = noisefor t in scheduler.timesteps: with torch.no_grad(): noisy_residual = model(input, t).sample prev = scheduler.step(noisy_residual, t, input).prev_sample input = prevprint('MPS/CPU OK:', input.device, input.shape)"# 3. Scheduler 纯逻辑测试(不下载模型权重)python -c "from diffusers import DDIMScheduler, FlowMatchEulerDiscreteSchedulerimport torchsched = DDIMScheduler(num_train_timesteps=1000)sched.set_timesteps(20)print('DDIM timesteps:', sched.timesteps)flow = FlowMatchEulerDiscreteScheduler(num_train_timesteps=1000)print('FlowMatch sigmas:', flow.sigmas[:5])"# 4. 跑 tests/schedulers 子集(纯 CPU,无 GPU 依赖)pytest tests/schedulers/test_scheduling_ddim.py -x -q --no-header -rN# 5. 跑 tests/others 子集(配置、工具类测试)pytest tests/others/ -x -q --no-header -rN# 6. 验证 Hub 加载与 pipeline 构造(不推理,只验 from_pretrained 路径)python -c "from diffusers import DiffusionPipelinepipe = DiffusionPipeline.from_pretrained('google/ddpm-cat-256', device='cpu')print('Pipeline loaded:', type(pipe).__name__)"
测试 框架:pytest + pytest-xdist(见 pyproject.toml 与 tests/conftest.py)。目录:tests/ 下 615 个文件,分 models/、pipelines/、schedulers/、lora/、quantization/、modular_pipelines/、others/、single_file/、hooks/、remote/。只跑 CPU 子集:pytest tests/schedulers/ tests/others/ -x -q --no-header -rN,预计 5–15 分钟(MacBook Air 16GB)。跳过 GPU 测试:CI 用 @pytest.mark.skipIf(not torch.cuda.is_available()) 装饰器,本地不满足会自动 skip。大模型测试会下载权重,建议用 HF_HUB_OFFLINE=1 或 TRANSFORMERS_OFFLINE=1 限制网络。完整套件需 CUDA 机器,不在本地跑。
调试 断点入口:src/diffusers/pipelines/stable_diffusion/pipeline_stable_diffusion.py 的 __call__ 方法;src/diffusers/models/unets/unet_2d_condition.py 的 forward;src/diffusers/schedulers/scheduling_ddim.py 的 step。 日志开关:diffusers.utils.logging.set_verbosity_debug() 启用内部日志;DIFFUSERS_DISABLE_IMPORT_ERRORS=1 跳过可选依赖报错。 环境变量:HF_HUB_OFFLINE=1 强制离线;TRANSFORMERS_OFFLINE=1 同上;DIFFUSERS_CACHE 控制缓存目录;CUDA_VISIBLE_DEVICES='' 在 CPU 脚本里显式禁 CUDA。 Profiling 入口:examples/profiling/profiling_pipelines.py 与 profiling_utils.py(需 GPU 才能跑完整);CPU 上可用 torch.profiler 手动包 pipeline.__call__ 或 model.forward。 MPS 调试:PYTORCH_MPS_HIGH_WATERMARK_RATIO=0.0 调整 MPS 内存高水位;MPS_FALLBACK=1 让不支持的 op 回退 CPU。 常见断点模式:在 attention_processor.py 的 AttnProcessor2_0.__call__ 设断点,观察注意力路径是否走 CUDA 分支。 CI CI 系统:GitHub Actions(见 .github/workflows/)。跑哪些:make style(ruff lint + format 检查,见 pyproject.toml [tool.ruff] 配置)、make fix-copies(一致性检查)、pytest 分模型/调度器/pipeline 多矩阵跑、文档构建、自测试(utils/check_repo.py 等)。PR 会被卡住:1) make style 失败(行尾、import 排序、未用变量);2) make fix-copies 失败(src/diffusers 与 utils/copy_utils.py 生成的代码不一致);3) 缺少测试(utils/check_test_missing.py 会扫描新增模型/ pipeline 是否带测试);4) 提交信息格式(utils/log_reports.py 检查);5) 文档 TOC 不一致(utils/check_doc_toc.py)。PR 前本地必跑:make style && make fix-copies。
坑 1. 直接 pip install diffusers[torch] 会拉 CUDA 变体 torch,在 macOS 上可能装错版本;应先单独装 arm64 torch 再装 diffusers。 2. 16GB 内存跑 UNet2DConditionModel 或 Transformer2DModel 大模型会 OOM;冒烟测试用 google/ddpm-cat-256(64x64 小模型)。 3. AttnProcessor2_0 在 CPU/MPS 上会回退到慢路径,不要误判为 bug;这是预期行为。 4. torch.compile 在 macOS 上后端不稳定,遇到奇怪报错先禁掉 pipeline.enable_model_cpu_offload() 或 torch.compile 调用。 5. tests/ 下很多测试用 @pytest.mark.skipIf(not torch.cuda.is_available()) 装饰,本地不显示失败但会 skip,勿误以为通过。 6. scripts/convert_*.py 中部分脚本硬编码 device = 'cuda'(如 convert_amused.py),直接跑会报 CUDA 不可用;需读代码改 device='cpu' 或上 Colab。 7. make fix-copies 会重写 src/diffusers 下大量 # Copied from 注释块,提交前必须跑,否则 CI 红。 8. examples/ 下训练脚本全部需要多 GPU,本地只能读代码理解流程,不能直接跑。 9. HuggingFace Hub 下载大模型时默认走 ~/.cache/huggingface/,16GB 磁盘紧张时设 HF_HUB_CACHE 到外置盘。 10. MPS 后端对部分 op(如某些 scatter、index_put)支持不全,遇到 MPS device not supported 错误时用 device='cpu' 回退验证。 七、任务卡 任务 1 优先 · medium · CPU · 2 个晚上
修复 UniPCMultistepScheduler 在 sigma=0 末步返回 NaN/inf 问题(solver_type=bh1、lower_order_final=False、predict_x0=False) Issue #14887 · UniPCMultistepScheduler returns nan with solver_type="bh1", and inf or nan with lower_order_final=False or predict_x0=False, on the final step to sigma = 0 ↗
无人认领 0 条评论 更新 2026-09-27
用到的专长: GPU 性能工程师对数值精度与算子边界条件敏感;本任务核心是 scheduler 的 dtype / 边界步逻辑,与 CUDA 无关。
目标: 让 UniPCMultistepScheduler 在 final_sigmas_type="zero"(默认配置)下,所有 solver_type / lower_order_final / predict_x0 组合在末步都返回有限数,行为与 DPMSolverMultistepScheduler 一致。
为什么值得长期做: Scheduler 是 Diffusers 三大核心抽象之一,数值正确性直接影响所有使用 UniPC 的 pipeline。修复 sigma=0 末步的 NaN/inf 属于纯逻辑层,不依赖 CUDA,且能建立你在调度器方向的可信度——这是 ownership_target 的核心区域。
怎么介入: Issue 14887 无 assignee、无 open PR、0 条评论,2026-09-27 创建,状态新鲜。可直接认领。第一个 PR 的边界: 第一个 PR 只改 src/diffusers/schedulers/scheduling_unipc_multistep.py 的 step、multistep_uni_p_bh_update、__init__,并新增 tests/schedulers/test_scheduler_unipc.py 中的 test_final_step_finiteness。不动其他 scheduler。
第一步: 在本地 MacBook 上用 Issue 提供的复现脚本跑通(CPU,float32),确认 bh1 末步 NaN、lower_order_final=False 末步 inf、predict_x0=False 末步 NaN 三种现象。
本机怎么复现 / 验证: python -c "import torch; from diffusers import UniPCMultistepScheduler; s=UniPCMultistepScheduler(); s.set_timesteps(10); x=torch.rand(1,3,8,8); [s.step(x*t/(t+1),t,x) for t in s.timesteps]" # CPU 上复现 NaN/inf;然后 pytest tests/schedulers/test_scheduler_unipc.py -x 验证修复。
认领留言(英文,可直接贴到 Issue) I'd like to take this. The root cause is that UniPC's final step to sigma=0 runs at full order with inf lambda, and only the default (bh2/lower_order_final=True/predict_x0=True) path survives. I'll align with DPMSolverMultistepScheduler: force this_order=1 on the last step when final_sigmas_type=='zero', guard the B_h correction in multistep_uni_p_bh_update to D1s is not None, and raise ValueError for predict_x0=False + final_sigmas_type='zero' in __init__. I'll add a finiteness test covering all solver_type/lower_order_final/predict_x0 combinations. Reproduction and tests run on CPU. PR ready in ~2 evenings. Does this scope look right? 复制留言
大致实施方案 阅读 src/diffusers/schedulers/scheduling_unipc_multistep.py 的 step、multistep_uni_p_bh_update、multistep_uni_c_bh_update 三个函数。 在 step 中加入末步强制一阶逻辑:当 final_sigmas_type == "zero" 且当前步是末步时,this_order = 1(与 DPMSolverMultistepScheduler.step 对齐)。 在 multistep_uni_p_bh_update 中,仅当 D1s is not None 时才减去 B_h * pred_res 项,避免 -inf * 0 = NaN。 在 __init__ 中加入 ValueError:predict_x0=False 与 final_sigmas_type="zero" 组合非法(与 DPMSolverMultistepScheduler 一致)。 在 tests/schedulers/test_scheduler_unipc.py 新增 test_final_step_finiteness,覆盖 bh1/bh2 × lower_order_final True/False × predict_x0 True/False × 多步数,断言 prev_sample.isfinite().all()。 本地跑 pytest tests/schedulers/test_scheduler_unipc.py -x,确认全部通过。 可能涉及的目录或文件 src/diffusers/schedulers/scheduling_unipc_multistep.py(step、multistep_uni_p_bh_update、multistep_uni_c_bh_update、__init__) tests/schedulers/test_scheduler_unipc.py(新增测试)
验收方式 本地 CPU 复现脚本(Issue 正文)不再产出 NaN/inf。 pytest tests/schedulers/test_scheduler_unipc.py -x 全部通过(含新增 test_final_step_finiteness)。 ruff check src/diffusers/schedulers/scheduling_unipc_multistep.py 无新增 warning。 开工前问题与风险 向维护者确认 是否同意与 DPMSolverMultistepScheduler 对齐的修复方案(末步强制一阶 + __init__ 拒绝 predict_x0=False + final_sigmas_type="zero")? 是否需要同时处理 final_sigmas_type="sigma_min" 路径(该路径末步 h=0 但一阶更新安全,Issue 未要求)? 风险 改动 step 主路径可能影响现有 pipeline 的数值回归,需确保只在末步触发新逻辑。 __init__ 加 ValueError 是破坏性变更,需确认没有现有配置依赖该组合。 任务 2 优先 · easy · CPU · 1 个晚上
修复 UniPCMultistepScheduler 在 float64 默认 dtype 下 torch.linalg.solve 的 dtype 不匹配崩溃 Issue #14888 · UniPCMultistepScheduler fails in torch.linalg.solve under a float64 default dtype: the unit entry of rks takes the default dtype ↗
无人认领 0 条评论 更新 2026-09-27
用到的专长: dtype 传播与算子边界是 GPU 性能工程师日常;本任务只需 CPU 验证。
目标: 让 UniPCMultistepScheduler 在 torch.set_default_dtype(torch.float64) 下正常运行,rks 单位项继承 sigmas 的 dtype。
为什么值得长期做: 与 #14887 同属 UniPCScheduler 数值鲁棒性方向,修复极小(两处 torch.ones(..., device=device) → torch.ones_like(h)),但影响所有在 float64 环境下使用 UniPC 的用户。维护者对这类 dtype 边界 bug 接受度高。
怎么介入: Issue 14888 无 assignee、无 open PR(linked PR #14920 和 #1 都已 closed,未合并),0 条评论,2026-09-27 创建。可重新认领。第一个 PR 的边界: 第一个 PR 只改两处 torch.ones 调用并新增一个测试函数。不动其他逻辑。
第一步: 在本地 MacBook 跑 Issue 复现脚本(CPU,float64 默认 dtype),确认 RuntimeError。
本机怎么复现 / 验证: python -c "import torch; torch.set_default_dtype(torch.float64); from diffusers import UniPCMultistepScheduler; s=UniPCMultistepScheduler(); s.set_timesteps(10); x=torch.rand(1,3,8,8); [s.step(x*t/(t+1),t,x) for t in s.timesteps]" # CPU 复现 RuntimeError;然后 pytest tests/schedulers/test_scheduler_unipc.py -x 验证。
认领留言(英文,可直接贴到 Issue) I'll take this. The fix is two one-liners: replace torch.ones((), device=device) with torch.ones_like(h) in both multistep_uni_p_bh_update and multistep_uni_c_bh_update so the unit entry inherits the sigmas' dtype. I'll add a test_default_dtype_float64 test. CPU-only reproduction and verification. PR ready in ~1 evening. Should I combine this with #14887 in one PR or keep separate? 复制留言
大致实施方案 定位 src/diffusers/schedulers/scheduling_unipc_multistep.py 中 multistep_uni_p_bh_update 和 multistep_uni_c_bh_update 两处 torch.ones((), device=device)。 改为 torch.ones_like(h),使单位项与 sigmas 同 dtype。 在 tests/schedulers/test_scheduler_unipc.py 新增 test_default_dtype_float64,用 torch.set_default_dtype(torch.float64) 跑完整 step 循环。 本地跑 pytest tests/schedulers/test_scheduler_unipc.py -x 确认通过。 可能涉及的目录或文件 src/diffusers/schedulers/scheduling_unipc_multistep.py(multistep_uni_p_bh_update、multistep_uni_c_bh_update) tests/schedulers/test_scheduler_unipc.py(新增测试)
验收方式 Issue 复现脚本在 float64 默认 dtype 下不再抛 RuntimeError。 pytest tests/schedulers/test_scheduler_unipc.py -x 全部通过。 ruff check 无新增 warning。 开工前问题与风险 向维护者确认 是否同意 torch.ones_like(h) 方案?Issue 作者已验证 540 配置 bit-identical。 风险 几 乎 无 风 险 : t o r c h . o n e s _ l i k e ( h ) 在 f l o a t 3 2 默 认 下 与 原 来 b i t - i d e n t i c a l 。 任务 3 优先 · medium · CPU · 2 个晚上
修复 DPMSolver 与 UniPC 在 final_sigmas_type=sigma_min + 转换 sigma 时末步 h=0 返回 NaN Issue #14890 · Converted sigma schedules with final_sigmas_type="sigma_min" repeat the last sigma, and the third-order and heun updates of DPMSolverMultistepScheduler and UniPCMultistepScheduler return NaN on that zero-width final step ↗
无人认领 0 条评论 更新 2026-09-28
用到的专长: scheduler 数值边界与 dtype 分析是性能工程师强项;纯 CPU 可验证。
目标: 让 DPMSolverMultistepScheduler(solver_order=3 或 heun)和 UniPCMultistepScheduler(order=3, lower_order_final=False)在使用 converted sigmas(Karras/exp/beta/lu_lambdas)+ final_sigmas_type="sigma_min" 时,末步 h=0 不返回 NaN。
为什么值得长期做: 与 #14887 / #14888 构成 scheduler 数值正确性三连击。本任务涉及 DPMSolverMultistepScheduler 和 UniPCMultistepScheduler 两个最常用 scheduler 的 converted sigma 路径,影响面广。Issue 作者已给出完整根因分析,只需落地修复。
怎么介入: Issue 14890 无 assignee、无 open PR、0 条评论,2026-09-28 创建,状态新鲜。可直接认领。第一个 PR 的边界: 第一个 PR 只改两个 scheduler 的 step 函数末步检测逻辑,并新增 finiteness 测试。不动 set_timesteps。
第一步: 在本地 MacBook 用 Issue 的 karras 3-steps 配置复现 NaN(CPU,float32)。
本机怎么复现 / 验证: python -c "import torch; from diffusers import DPMSolverMultistepScheduler; s=DPMSolverMultistepScheduler(use_karras_sigmas=True); s.set_timesteps(3); s.config.final_sigmas_type='sigma_min'; x=torch.rand(1,3,8,8); [s.step(x,t,x) for t in s.timesteps]" # CPU 复现 NaN;然后 pytest 验证。
认领留言(英文,可直接贴到 Issue) I'd like to take this. The root cause is that converted sigmas already end at sigma_min, so appending sigma_min again makes the last step h=0, and the third-order/heun updates divide by h. I'll detect h=0 in DPMSolverMultistepScheduler.step and UniPCMultistepScheduler.step and force this_order=1 there (mirroring the existing final_sigmas_type=='zero' guard). I'll add finiteness tests for Karras/exp/beta/lu_lambdas + sigma_min. CPU-only reproduction. PR in ~2 evenings. Should I also cover DPMSolverSinglestep/DEIS/SASolver/Euler in the same PR? 复制留言
大致实施方案 阅读 DPMSolverMultistepScheduler.step 中已有的 final_sigmas_type=="zero" 强制一阶逻辑,理解为何 "sigma_min" 路径未覆盖。 在 DPMSolverMultistepScheduler.step 中检测末步 h=0(或 sigmas[-1]==sigmas[-2])时强制 this_order=1。 在 UniPCMultistepScheduler.step 中同样检测末步 h=0 并强制 this_order=1。 在 tests/schedulers/test_scheduler_dpm_multi.py 和 test_scheduler_unipc.py 新增 converted-sigma + sigma_min 末步 finiteness 测试。 本地跑 pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py -x。 可能涉及的目录或文件 src/diffusers/schedulers/scheduling_dpm_solver_multistep.py(step) src/diffusers/schedulers/scheduling_unipc_multistep.py(step) tests/schedulers/test_scheduler_dpm_multi.py、test_scheduler_unipc.py(新增测试)
验收方式 karras 3-steps + final_sigmas_type=sigma_min 末步不再 NaN。 pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py -x 全部通过。 开工前问题与风险 向维护者确认 是否同意在 step 中检测 h=0 并强制一阶,而不是在 set_timesteps 中避免重复 sigma? 是否需要同时处理 DPMSolverSinglestep / DEIS / SASolver / EulerDiscrete(Issue 提到它们也有重复 sigma)? 风险 强 制 一 阶 可 能 改 变 末 步 数 值 精 度 , 但 I s s u e 已 分 析 一 阶 更 新 在 h = 0 时 是 精 确 的 ( s a m p l e u n c h a n g e d ) 。 任务 4 可选 · easy · Mac · 1 个晚上
修复 QwenImage21 FlexAttention mask 在 MPS 上因 ~ 操作符编译失败 Issue #14889 · QwenImage21 FlexAttention mask fails to compile on MPS because of boolean ~ ↗
无人认领 0 条评论 更新 2026-09-28
用到的专长: MPS 后端适配与编译期问题属于端侧部署范畴;候选人熟悉 TensorRT/量化,MPS 是同类技能。
目标: 让 build_qwenimage21_block_causal_mask 在 MPS + torch.compile(fullgraph=True) 下正常编译,不抛 NotImplementedError。
为什么值得长期做: README 明确写 Apple Silicon 支持,MPS 是 Diffusers 一等公民。本任务修复 MPS 后端 FlexAttention 的编译失败,直接提升 Mac 用户体验,且改动极小(~ → torch.logical_not),风险低。
怎么介入: Issue 14889 无 assignee、无 open PR、0 条评论,2026-09-28 创建。可直接认领。第一个 PR 的边界: 第一个 PR 只改 transformer_qwenimage21.py 一行,并新增一个 MPS-only 测试。不动其他模型。
第一步: 在本地 MacBook(MPS 可用)跑 Issue 复现脚本,确认 NotImplementedError。
本机怎么复现 / 验证: python -c "import torch; from torch.nn.attention.flex_attention import flex_attention; from diffusers.models.transformers.transformer_qwenimage21 import build_qwenimage21_block_causal_mask; device='mps'; ids=torch.full((97,),-1,dtype=torch.long,device=device); valid=torch.ones(2,97,dtype=torch.bool,device=device); mask=build_qwenimage21_block_causal_mask(ids,valid,2,device); q=torch.randn(2,2,mask.shape[-1],128,device=device); flex_attention(q,q,q,block_mask=mask)" # MPS 上复现 NotImplementedError;修复后通过。
认领留言(英文,可直接贴到 Issue) I'll take this on my Apple Silicon Mac. The fix is a one-liner: replace ~is_padding with torch.logical_not(is_padding) in build_qwenimage21_block_causal_mask. The MPS FlexAttention lowerer rejects aten.bitwise_not but accepts torch.logical_not for boolean tensors. I'll add a MPS-only regression test using torch.compile(fullgraph=True). PR ready in ~1 evening. Any preference on where the test should live? 复制留言
大致实施方案 定位 src/diffusers/models/transformers/transformer_qwenimage21.py 中 build_qwenimage21_block_causal_mask 的 return allowed & ~is_padding。 改为 return allowed & torch.logical_not(is_padding)。 在 tests/models/ 新增 test_qwenimage21_mps_flex_attention(或扩展现有测试),在 MPS 可用时跑 torch.compile(fullgraph=True) 验证。 本地在 MPS 上跑该测试确认通过。 可能涉及的目录或文件 src/diffusers/models/transformers/transformer_qwenimage21.py(build_qwenimage21_block_causal_mask) tests/models/test_transformer_qwenimage21.py 或新增测试文件(新增)
验收方式 本地 MPS 复现脚本不再抛 NotImplementedError。 新增测试在 MPS 上通过,在 CPU 上 skip。 开工前问题与风险 向维护者确认 是否同意 torch.logical_not 方案?Issue 作者已验证 MPS 上可行。 是否需要同时检查其他模型中类似的 ~boolean 模式? 风险 几 乎 无 风 险 : t o r c h . l o g i c a l _ n o t 与 ~ 在 语 义 上 对 b o o l e a n t e n s o r 完 全 等 价 。 任务 5 可选 · easy · CPU · 1 个晚上
修复 numpy_to_pil 对灰度图 singleton 空间维度的错误 squeeze Issue #14901 · numpy_to_pil removes singleton spatial dimensions from grayscale images ↗
无人认领 1 条评论 更新 2026-09-29
用到的专长: 图像预处理与 tensor 形状分析是部署工程师日常;纯 CPU 可验证。
目标: 让 numpy_to_pil 在 HWC 灰度分支只 squeeze channel 维度,保留 singleton 的 H 或 W。
为什么值得长期做: 图像预处理是 pipeline 最末一环,numpy_to_pil 的 squeeze 错误会导致 H×1 或 1×W 灰度图被压成 1×1。Issue 作者已给出完整复现与修复方案(squeeze(-1)),维护者尚未回应。纯工具函数修复,风险极低。
怎么介入: Issue 14901 无 assignee、无 open PR、1 条评论(作者自己),2026-09-29 创建。可直接认领。第一个 PR 的边界: 第一个 PR 只改 pil_utils.py 和 image_processor.py 的 squeeze 调用,并新增测试。examples 修复视维护者反馈决定是否同 PR。
第一步: 在本地 MacBook 跑 Issue 复现脚本(CPU),确认 (1,7,1) 被压成 (1,1)。
本机怎么复现 / 验证: python -c "import numpy as np; from diffusers.utils.pil_utils import numpy_to_pil; x=np.zeros((1,7,1),dtype=np.float32); y=numpy_to_pil(x)[0]; print(y.size)" # 期望 (7,1),实际 (1,7);修复后通过。
认领留言(英文,可直接贴到 Issue) I'll take this. The fix is to replace image.squeeze() with image.squeeze(-1) in the grayscale branch of numpy_to_pil, and the same pattern in image_processor.py (2 occurrences). I'll also review the unrestricted squeeze in the two Flux2 DreamBooth examples. I'll add a regression test covering (1,1,1), (1,7,1), (7,1,1), and ordinary grayscale/RGB/RGBA cases. CPU-only reproduction. PR in ~1 evening. Should I fix the examples in the same PR or file a separate issue for them? 复制留言
大致实施方案 定位 src/diffusers/utils/pil_utils.py 中 numpy_to_pil 的 grayscale 分支 Image.fromarray(image.squeeze(), mode="L")。 改为 Image.fromarray(image.squeeze(-1), mode="L")。 检查 src/diffusers/image_processor.py 中两处相同模式并同步修复。 检查 examples/dreambooth/train_dreambooth_lora_flux2_img2img.py 和 train_dreambooth_lora_flux2_klein_img2img.py 中 CHW 灰度 squeeze,改为 channel-axis-specific。 在 tests/ 新增 test_numpy_to_pil_singleton_dims,覆盖 (1,1,1)、(1,7,1)、(7,1,1)、(1,7,3) 等 case。 本地跑 pytest tests/ -k numpy_to_pil 或对应测试文件。 可能涉及的目录或文件 src/diffusers/utils/pil_utils.py(numpy_to_pil) src/diffusers/image_processor.py(两处) examples/dreambooth/train_dreambooth_lora_flux2_img2img.py、train_dreambooth_lora_flux2_klein_img2img.py(需 review) tests/utils/test_pil_utils.py 或新增测试文件(新增)
验收方式 Issue 复现脚本输出 (7,1) 而非 (1,7)。 新增测试全部通过。 ruff check 无新增 warning。 开工前问题与风险 向维护者确认 是否同意 squeeze(-1) 方案?Issue 作者已验证 9 cases。 examples 中的 CHW squeeze 是否一并修,还是留给示例作者? 风险 极 低 : s q u e e z e ( - 1 ) 只 在 c h a n n e l 维 度 为 1 时 生 效 , 对 普 通 R G B 图 无 影 响 。 任务 6 可选 · easy · Mac · 1 个晚上
新增 DDIM 与 FlowMatchEuler 调度器的 CPU/MPS 跨后端一致性测试 Issue #12760 · StateManager, Hook, State: missing context setting (cond or uncond) in pipelines ↗
无人认领 6 条评论 stale contributions-welcome 更新 2026-10-05
用到的专长: 跨后端数值一致性测试是性能工程师写 micro-benchmark 的日常;Mac 上可完整验证。
目标: 在 tests/schedulers/ 新增 test_cross_backend_consistency.py,覆盖 DDIMScheduler 和 FlowMatchEulerDiscreteScheduler 在 CPU vs MPS 上的 step 输出一致性。
为什么值得长期做: gaps 明确指出 tests/schedulers/ 缺少跨后端数值一致性测试。本任务在 CPU/MPS 上固定 seed 与 timestep,断言 prev_sample 的 max abs error < 1e-5,填补 README 声称的 Apple Silicon 支持与实际测试覆盖之间的鸿沟。与 ownership_target 的「调度与性能基线」方向完全对齐。
怎么介入: Issue 12760 是 contributions-welcome 标签的 stale issue,但正文讨论的是 cache_context 缺失,与本任务不直接相关。本任务作为 proposal 卡,source_url 指向仓库 commits,engagement 写明先开 Issue 提案。第一个 PR 的边界: 第一个 PR 只新增 tests/schedulers/test_cross_backend_consistency.py,不改任何源文件。
第一步: 在本地 MacBook 确认 MPS 后端可用(torch.backends.mps.is_available()),阅读 tests/schedulers/ 现有测试结构。
本机怎么复现 / 验证: pytest tests/schedulers/test_cross_backend_consistency.py -x -v # 在 Mac 上直接运行新增测试,MPS 可用则跑 CPU vs MPS 对比,否则 skip。
认领留言(英文,可直接贴到 Issue) [Proposal] Add CPU/MPS cross-backend consistency tests for schedulers
I'd like to add tests/schedulers/test_cross_backend_consistency.py that fixes seed + timestep and asserts max abs error < 1e-5 between CPU and MPS for DDIMScheduler and FlowMatchEulerDiscreteScheduler. This fills the gap noted in the repo's Apple Silicon support claim vs actual test coverage. All tests run on Mac (MPS) with CPU fallback. Should I open a dedicated issue first, or is a PR with the tests sufficient? 复制留言
大致实施方案 新建 tests/schedulers/test_cross_backend_consistency.py。 对每个 scheduler(DDIM、FlowMatchEuler),固定 seed、timestep、sample,分别在 CPU 和 MPS 上跑 step,断言 max abs error < 1e-5。 MPS 不可用时自动 skip(pytest.mark.skipif)。 本地在 MPS 上跑 pytest tests/schedulers/test_cross_backend_consistency.py -x。 可能涉及的目录或文件 tests/schedulers/test_cross_backend_consistency.py(新增)
验收方式 本地 MPS 上测试通过。 CPU 上测试通过(MPS skip)。 pytest tests/schedulers/ -x 全部通过。 开工前问题与风险 向维护者确认 是否同意 max abs error < 1e-5 的阈值?是否需要覆盖更多 scheduler(DDPM、LCMScheduler)? 风险 极 低 : 纯 新 增 测 试 , 不 改 源 逻 辑 。 若 M P S 数 值 偏 差 超 过 阈 值 , 会 暴 露 现 有 问 题 而 非 引 入 新 b u g 。