← 所有项目

Contribution Tasks

huggingface/diffusers

Diffusers 是 PyTorch 扩散模型库,usability 优先、性能其次;核心抽象(Scheduler / UNet / DiT / VAE / Pipeline)几乎全是 PyTorch 原生算子,CPU/MPS 可执行。候选人的 GPU kernel / 训练性能专长能发挥在:scheduler 数值正确性与跨后端一致性、CPU/MPS 性能基线与 profiling 示例、算子 dtype 鲁棒性(bf16/fp32 边界)、转换脚本与量化路径的纯逻辑测试。仓库 README 明确写 Apple Silicon 支持,16GB MacBook Air 足以跑通调度器、小模型 forward、转换脚本与测试套件;真正需要 CUDA 的只是最终吞吐验证,可用 Colab T4 做短时确认。

当前方向:近 6–8 周一个 minor release(最新 0.40.0),热点集中在 src/diffusers(109 次)、examples/community(12 次)、docs(5 次)、tests/models(4 次)。方向是:持续加新模型/新 pipeline(Flux2、QwenImage2.1、Cosmos3、Krea 2、Minimax H3)、推广 modular pipelines、补测试与文档。性能与 CPU/MPS 正确性不是主线,但维护者对 scheduler 数值 bug、跨设备一致性、死代码清理反应积极——这正是候选人能切入的空白。

★ 34653Fork 73766 个候选任务Gemini:LongCat-2.0

更新于 2026-10-08T07:37:23+00:00 · 打开仓库 ↗

一、项目定位

🤗 Diffusers 是 Hugging Face 维护的 PyTorch 扩散模型推理与训练库,定位为 image / video / audio / 3D 分子生成的「模块化工具箱」。README 明确把哲学写在最前面:usability over performance、simple over easy、customizability over abstraction,因此它不是追求极致 kernel 性能的项目,而是追求可组合、易扩展、易贡献。核心抽象三层:DiffusionPipeline(端到端推理入口)、Scheduler(噪声调度与步进)、Model(UNet / DiT / VAE 等可替换积木),外加可插拔的 LoRA / ControlNet / Adapter 机制。 从证据看,这是一个成熟的中年期项目:根目录 1047 个 src 文件、615 个测试文件、509 个文档文件、95 个转换脚本、464 个 examples;最近 release 是 0.40.0(2026-08-20),再往前 0.39.0、0.38.0,约 6–8 周一个 minor,节奏稳定;最近提交热点集中在 src/diffusers(109 次)、examples/community(12)、docs(5)、tests/models(4)、workflows(3),说明仍在活跃地加新模型/新 pipeline,同时补测试和文档。 同类项目里,它是「生态最大、模型最多、社区最活跃」的那一个,但性能不是第一位——这正是 GPU kernel / 训练性能工程师可以切入的空间。README 明确写了 Apple Silicon / MPS 支持,意味着 CPU / Apple Silicon 上可以跑通很多流程,只是速度慢;真正要验证性能仍需 CUDA 机器(可用 Colab T4 做短时确认)。

解决什么问题、给谁用

解决的具体问题:把学术界和工业界不断扩散(pun intended)的扩散模型变体统一到一个可组合、可训练、可推理的框架里,让用户用几行代码跑 SOTA,同时让研究者能快速把新论文落地成 pipeline / model / scheduler。 目标用户和典型场景: 1. 应用开发者:用 DiffusionPipeline.from_pretrained 几行代码做 text-to-image / video / audio 推理,部署到 Hub、ONNX、TensorRT; 2. 研究者 / 训练工程师:用 examples/ 和 training scripts 做 DreamBooth、LoRA、ControlNet、一致性蒸馏、RL 训练; 3. 模型作者:把自家 checkpoint 通过 scripts/convert_*.py 转成 diffusers 格式,或新增 pipeline / model / scheduler 合入核心库; 4. 优化工程师(你的切入点):在 optimization/ 和 benchmarks/ 下做 FP16 / FP8、tensor parallelism、kernel 融合、推理加速——这些正是当前提交热点里反复出现的关键词(FP8、tensor-parallel、single-file、MPS)。

同类项目与差别

核心能力

能力在哪成熟度
端到端扩散推理 Pipeline(text-to-image / video / audio / inpainting / ControlNet / LoRA 等)src/diffusers/pipelines/成熟
可插拔 Scheduler(DDPM / DDIM / Euler / FlowMatch / 等几十种)src/diffusers/schedulers/成熟
可替换 Model 积木(UNet2D、UNet3D、DiT、Vae、Transformer2D 等)src/diffusers/models/成熟
LoRA / ControlNet / Adapter / IP-Adapter 等可插拔注入机制src/diffusers/(loaders / peft / controlnet 等子目录)成熟
训练脚本(DreamBooth、LoRA、ControlNet、一致性蒸馏、RL、advanced training)examples/ 下 dreambooth / text_to_image / controlnet / consistency_distillation / reinforcement_learning / advanced_diffusion_training成熟
Modular Diffusers(可组合的 pipeline 组件化新范式)src/diffusers/modular_pipelines/、docs/source/en/modular_diffusers实验
Tensor Parallelism(FP8 / 分布式推理与加载)src/diffusers/(tensor parallel 相关)、最近提交含 FP8 / tensor-parallel 修复实验
Single-File 支持(单 checkpoint 直接加载,无需多文件目录)src/diffusers/(single_file 相关)、最近提交多次 Add single file support for ...实验
Apple Silicon / MPS 推理支持docs/source/en/optimization/mps、README 明确提及成熟
ONNX / Hub 导出与 CLI 工具docker/(onnxruntime 镜像)、diffusers-cli(AGENTS.md / CLAUDE.md 提及)成熟
模型格式转换脚本(数百种 convert_*.py)scripts/(95 个文件,大量 convert_*.py)成熟
Benchmarking 与 profiling 工具benchmarks/(benchmarking_flux / sdxl / ltx / wan 等)、examples/profiling成熟

阶段:成熟活跃期(maintenance + 持续扩张)。证据:0.38–0.40 三个 minor release 在 4 个月内完成;提交热点仍在加新模型(Krea 2、Minimax H3、Cosmos3、QwenImage 2.1)和新能力(FP8 tensor parallelism、single-file support、modular pipelines);同时在做清理(Remove deprecated code paths)和补测试(tests/models、tests/modular_pipelines)。README 没有标 Experimental,整体是生产可用状态;但 modular pipelines、tensor parallelism、部分新 pipeline 仍在快速迭代,属于「核心稳、前沿动」。

技术栈:Python;PyTorch(核心依赖);HuggingFace Hub / huggingface_hub(模型加载);transformers(文本编码器等);safetensors(权重格式);PEFT(LoRA 训练);accelerate(分布式训练);ruff(lint/format,pyproject.toml 配置);make(style / fix-copies / pre-release);GitHub Actions CI(.github/workflows,46 个文件);uv(推荐虚拟环境,AGENTS.md / CLAUDE.md)

规模:根目录一级条目 27 个;src/ 1047 个文件;tests/ 615 个文件;docs/ 509 个文件;examples/ 464 个文件;scripts/ 95 个文件;utils/ 31 个文件;.github 46 个文件。最近提交热点:src/diffusers 109 次、examples/community 12 次、docs/source 5 次、tests/models 4 次、workflows 3 次。Release:0.40.0(2026-08-20)、0.39.0(2026-07-03)、0.38.0(2026-05-01),约 6–8 周一个 minor。pushed_at 2026-10-05,仍在活跃推进。star 数未在证据中提供,需验证。

二、架构与代码地图

Diffusers 的代码组织可以清晰地划分为五层,从上到下依次是: 1. 接入层(Entry Layer):面向用户的统一入口。核心是 DiffusionPipeline 及其数百个具体子类(如 StableDiffusionXLPipeline、FluxPipeline、CogVideoXPipeline 等),负责把 from_pretrained()、to(device)、__call__() 这些用户接口翻译成内部调度。ModularDiffusionPipeline 是近几个版本新增的模块化组合入口,允许用声明式方式拼装 pipeline。这一层还包含 Hub / CLI / ONNX / TensorRT 等外部对接。 2. 调度层(Scheduling Layer):决定「每一步去多少噪」。src/diffusers/schedulers/ 下聚集了 DDIM、DDPM、FlowMatchEuler、LCMScheduler 等几十种 scheduler,每个 scheduler 实现 set_timesteps() 和 step(),输出 prev_sample 与附加信息。调度层是性能优化的关键杠杆之一——一致性蒸馏、少步生成、视频生成的时序调度都发生在这里。 3. 执行层(Execution Layer):真正的「去噪」计算发生在这里。src/diffusers/models/ 下是 UNet、DiT(Transformer2DModel 及各类 DiT 变体)、VAE(AutoencoderKL 及其派生)、ControlNetModel、MotionAdapter 等可组合积木。注意力机制被抽象成 AttnProcessor / AttnProcessor2_0 等 processor,是 xformers / SDPA / FlashAttention 的挂载点。 4. 硬件与内核适配层(Hardware/Kernel Adaptation Layer):这一层并不直接提供 Python kernel,而是通过 PyTorch 后端、torch.compile、xformers、FlashAttention、CuTe/CUTLASS(间接通过 PyTorch)、torch.backends.cuda.matmul.allow_tf32、enable_mem_efficient_sdp 等开关来影响底层 kernel 选择。src/diffusers/pipelines/_optional_dependencies.py 和设备分发逻辑(src/diffusers/utils/ 下的 torch_utils.py 等)负责探测可用后端。分布式训练和推理则依赖 Accelerate、FSDP、tensor parallelism(最近提交热点之一)等。 5. 工具与生态层(Tooling & Ecosystem Layer):examples/(训练脚本、社区 pipeline、研究原型)、scripts/(95 个 checkpoint 转换脚本)、benchmarks/、docker/、utils/(仓库一致性检查、文档生成、发布流程)、tests/(615 个测试文件)。这一层让库保持可维护、可发布、可复现。 这五层之间是严格单向依赖的:接入层调用调度层和执行层,调度层和执行层调用硬件适配层,工具层横切所有层。

接入层调度层执行层硬件适配层工具层DiffusionPipeline → Scheduler:配置配置DiffusionPipeline → UNet2DConditionModel:潜在变量潜在变量DiffusionPipeline → Transformer2DModel:潜在变量潜在变量DiffusionPipeline → AutoencoderKL:潜在变量潜在变量DiffusionPipeline → ControlNetModel:控制信号控制信号DiffusionPipeline → HubMixin:模型加载模型加载Scheduler → UNet2DConditionModel:timesteptimestepScheduler → Transformer2DModel:timesteptimestepUNet2DConditionModel → AttnProcessor:注意力注意力Transformer2DModel → AttnProcessor:注意力DiffusionPipeline → LoRA:权重融合权重融合DiffusionPipeline → OnnxRuntime:导出导出DiffusionPipeline → TensorRT:导出导出TrainingScripts → DiffusionPipeline:训练训练ConversionScripts → DiffusionPipeline:转换转换Tests → DiffusionPipeline:测试测试DiffusionPipelineDiffusionPipelineModularDiffusionPipelineModularDiffusionPipelineHubMixinHubMixinSchedulerSchedulerAutoencoderKLAutoencoderKLControlNetModelControlNetModelLoRALoRAUNet2DConditionModelUNet2DConditionModelTransformer2DModelTransformer2DModelAttnProcessorAttnProcessorOnnxRuntimeOnnxRuntimeTensorRTTensorRTTrainingScriptsTrainingScriptsConversionScriptsConversionScriptsTestsTestsUtilsUtils
Diffusers 架构图:五层结构,自上而下为接入层、调度层、执行层、硬件适配层、工具层
模块 / 路径职责 · 入口 · 依赖
DiffusionPipeline
src/diffusers/pipelines/diffusion_pipeline.py
1 个核心文件 + 数百个具体 pipeline 子类
所有 pipeline 的基类,提供 from_pretrained、to、save_pretrained、__call__ 等统一接口,是用户入口
入口:DiffusionPipeline.from_pretrained, DiffusionPipeline.__call__, DiffusionPipeline.save_pretrained, DiffusionPipeline.to
依赖:Scheduler, Model, Optional dependencies
具体 pipeline 在 src/diffusers/pipelines/ 下按模型名分目录,如 stable_diffusion_xl/、flux/、cogvideo/ 等
Scheduler
src/diffusers/schedulers/
几十种 scheduler 实现
噪声调度算法集合,决定每步去噪的步长、噪声比例,是生成质量与速度的核心
入口:Scheduler.set_timesteps, Scheduler.step, Scheduler.scale_model_input
依赖:需验证具体依赖,通常依赖 torch 和配置类
包括 DDIMScheduler、DDPM Scheduler、FlowMatchEulerDiscreteScheduler、LCMScheduler、ConsistancyScheduler 等
UNet2DConditionModel
src/diffusers/models/unets/
需验证具体文件数
条件 UNet 模型,是 Stable Diffusion 系列的核心去噪网络
入口:UNet2DConditionModel.forward
依赖:AttentionProcessor, Resnet, Down/Up blocks
SD 1.x/2.x/XL 等使用,正被 DiT 逐步取代
Transformer2DModel
src/diffusers/models/transformers/
需验证具体文件数
DiT(Diffusion Transformer)系列模型,Flux、SD3、HunyuanDiT 等使用
入口:Transformer2DModel.forward
依赖:Attention, PatchEmbed, PositionalEmbedding
近年新增模型的主流架构
AutoencoderKL
src/diffusers/models/autoencoders/
需验证具体文件数
VAE 模型,负责图像与潜在空间之间的编解码
入口:AutoencoderKL.encode, AutoencoderKL.decode
依赖:Encoder, Decoder, Quantizer
包括 AutoencoderKLQwenImage 等变体
ControlNetModel
src/diffusers/models/controlnets/
需验证具体文件数
ControlNet 条件控制模型,为 pipeline 提供额外控制信号
入口:ControlNetModel.forward
依赖:需验证,通常依赖 UNet 结构
支持多种 ControlNet 变体
AttnProcessor
src/diffusers/models/attention_processor.py
1 个核心文件
注意力处理器抽象层,挂载 xformers / SDPA / FlashAttention 等后端
入口:AttnProcessor.__call__, AttnProcessor2_0.__call__
依赖:torch.nn.functional.scaled_dot_product_attention
性能优化的关键扩展点
LoRA
src/diffusers/loaders/lora.py
需验证具体文件数
LoRA 加载与融合机制,支持低秩适配训练和推理
入口:LoraLoaderMixin.load_lora_weights, LoraLoaderMixin.fuse_lora
依赖:DiffusionPipeline, PEFT
训练和推理都广泛使用
ModularDiffusionPipeline
src/diffusers/modular_pipelines/
需验证具体文件数
模块化 pipeline 组合机制,允许声明式拼装
入口:ModularDiffusionPipeline.from_pretrained, ModularDiffusionPipeline.__call__
依赖:DiffusionPipeline, Scheduler, Model
近几个版本新增,提交热点之一
HubMixin
src/diffusers/hub_mixin.py
1 个核心文件
HuggingFace Hub 集成,提供 from_pretrained / save_pretrained / push_to_hub
入口:HubMixin.from_pretrained, HubMixin.save_pretrained
依赖:huggingface_hub
所有可保存模型的基类
OnnxRuntime
src/diffusers/pipelines/onnx_utils.py
需验证具体文件数
ONNX 导出与推理支持
入口:需验证
依赖:onnxruntime
端侧部署路径之一
TensorRT
examples/community/
需验证具体文件数
TensorRT 推理支持,通过社区 pipeline 和转换脚本
入口:需验证
依赖:tensorrt, polygraphy
见 examples/community/stable_diffusion_tensorrt_*.py
TrainingScripts
examples/
464 个文件
训练脚本集合,包括 DreamBooth、LoRA、ControlNet、一致性蒸馏等
入口:train_dreambooth.py, train_lora.py, train_controlnet.py
依赖:DiffusionPipeline, Accelerate, PEFT
examples/dreambooth/、examples/text_to_image/ 等
ConversionScripts
scripts/
95 个文件
checkpoint 转换脚本,把原始权重转为 diffusers 格式
入口:convert_*.py
依赖:diffusers, torch, transformers
如 convert_flux_to_diffusers.py、convert_sd3_to_diffusers.py 等
Tests
tests/
615 个文件
测试套件,覆盖 models、pipelines、schedulers、lora 等
入口:pytest
依赖:pytest, torch, diffusers
tests/models/、tests/pipelines/、tests/schedulers/ 等
Utils
utils/
31 个文件
仓库维护工具,包括一致性检查、文档生成、发布流程
入口:check_*.py, release.py
依赖:需验证
check_copies.py、check_repo.py、release.py 等
目录树(按文件数)
  • src/ 1047 个文件
    diffusers
  • tests/ 615 个文件
    __init__.py, conftest.py, fixtures, hooks, lora, models, modular_pipelines, others, pipelines, quantization, remote, schedulers, single_file, testing_utils.py
  • docs/ 509 个文件
    README.md, TRANSLATING.md, source
  • examples/ 464 个文件
    README.md, advanced_diffusion_training, amused, cogvideo, cogview4-control, community, conftest.py, consistency_distillation, controlnet, cosmos, cosmos3, custom_diffusion, discrete_diffusion, dreambooth
  • scripts/ 95 个文件
    __init__.py, change_naming_configs_and_checkpoints.py, conversion_ldm_uncond.py, convert_ace_step_to_diffusers.py, convert_amused.py, convert_anima_to_diffusers.py, convert_animatediff_motion_lora_to_diffusers.py, convert_animatediff_motion_module_to_diffusers.py, convert_animatediff_sparsectrl_to_diffusers.py, convert_anyflow_to_diffusers.py, convert_asymmetric_vqgan_to_diffusers.py, convert_aura_flow_to_diffusers.py, convert_blipdiffusion_to_diffusers.py, convert_cogvideox_to_diffusers.py
  • .github/ 46 个文件
    ISSUE_TEMPLATE, PULL_REQUEST_TEMPLATE.md, actions, dependabot.yml, labeler.yml, workflows
  • utils/ 31 个文件
    check_ai.py, check_config_docstrings.py, check_copies.py, check_doc_toc.py, check_dummies.py, check_forward_call_docstrings.py, check_inits.py, check_repo.py, check_return_annotations.py, check_support_list.py, check_table.py, check_test_missing.py, consolidated_test_report.py, custom_init_isort.py
  • .ai/ 17 个文件
    .claude-plugin, .codex-plugin, AGENTS.md, plugin.json, references, skills
  • benchmarks/ 10 个文件
    README.md, __init__.py, benchmarking_flux.py, benchmarking_ltx.py, benchmarking_sdxl.py, benchmarking_utils.py, benchmarking_wan.py, push_results.py, requirements.txt, run_all.py
  • docker/ 8 个文件
    diffusers-cli-cuda, diffusers-doc-builder, diffusers-onnxruntime-cpu, diffusers-onnxruntime-cuda, diffusers-pytorch-cpu, diffusers-pytorch-cuda, diffusers-pytorch-minimum-cuda, diffusers-pytorch-xformers-cuda
  • .claude-plugin/ 1 个文件
    marketplace.json
  • .gitignore/ 1 个文件
  • AGENTS.md/ 1 个文件
  • CITATION.cff/ 1 个文件
  • CLAUDE.md/ 1 个文件
  • CODE_OF_CONDUCT.md/ 1 个文件

一次调用怎么流过这些模块

一次典型的推理请求(以 pipeline("A cat").images[0] 为例)的数据流如下: 1. 入口解析:DiffusionPipeline.__call__() 接收 prompt 和参数,进行输入验证和默认值填充。 2. 文本编码:调用 self.text_encoder(CLIP 或 T5 等)将 prompt 编码为 text_embeddings,可能经过 tokenizer 预处理。 3. 潜在初始化:生成随机噪声 latents,形状为 (batch_size, channels, height//8, width//8)。 4. 调度器配置:调用 self.scheduler.set_timesteps(num_inference_steps) 设置去噪步数。 5. 去噪循环:对每个 timestep: a. 模型输入准备:scheduler.scale_model_input(latents, timestep) 缩放潜在变量。 b. 噪声预测:调用 self.unet(latents, timestep, text_embeddings) 或 self.transformer(...) 预测噪声,得到 noise_pred。这是计算最密集的部分,涉及注意力计算、残差连接等。 c. 调度步进:调用 scheduler.step(noise_pred, timestep, latents) 更新潜在变量,得到 prev_sample。 d. 循环:更新 latents = prev_sample,进入下一步。 6. 解码:循环结束后,调用 self.vae.decode(latents) 将潜在变量解码为图像像素。 7. 后处理:转换为 PIL Image 或 numpy 数组,返回给用户。 训练流程类似,但包含: - 数据加载和预处理 - 前向扩散(加噪) - 噪声预测 - 损失计算(MSE 或 flow matching loss) - 反向传播和参数更新 - 可能涉及梯度累积、混合精度、分布式训练 关键数据结构: - latents:潜在空间表示,4D Tensor - text_embeddings:文本条件,2D 或 3D Tensor - timestep:当前步数,标量或 1D Tensor - noise_pred:模型预测的噪声,与 latents 同形状 - SchedulerOutput:调度器输出,包含 prev_sample 和附加信息 调度点: - scheduler.set_timesteps():设置总步数 - scheduler.step():每步更新 - model.forward():噪声预测 - vae.decode():潜在解码

1入口解析DiffusionPipeline · src/diffusers/pipelines/diffusion_pipeline.py2文本编码CLIPTextModel · 需验证3潜在初始化torch.randn · 需验证4调度配置Scheduler · src/diffusers/schedulers/5噪声预测UNet2DConditionModel · src/diffusers/models/unets/6调度步进Scheduler · src/diffusers/schedulers/7潜在更新Scheduler · src/diffusers/schedulers/8解码AutoencoderKL · src/diffusers/models/autoencoders/9后处理DiffusionPipeline · src/diffusers/pipelines/diffusion_pipeline.py

关键类型与函数

名称路径用途
DiffusionPipelinesrc/diffusers/pipelines/diffusion_pipeline.py
Schedulersrc/diffusers/schedulers/scheduling_utils.py
UNet2DConditionModelsrc/diffusers/models/unets/unet_2d_condition.py
Transformer2DModelsrc/diffusers/models/transformers/transformer_2d.py
AutoencoderKLsrc/diffusers/models/autoencoders/vae.py
ControlNetModelsrc/diffusers/models/controlnets/controlnet.py
AttnProcessorsrc/diffusers/models/attention_processor.py
LoRAsrc/diffusers/loaders/lora.py
HubMixinsrc/diffusers/hub_mixin.py
ModularDiffusionPipelinesrc/diffusers/modular_pipelines/modular_pipeline.py
OnnxRuntimesrc/diffusers/pipelines/onnx_utils.py
TensorRTexamples/community/
TrainingScriptsexamples/
ConversionScriptsscripts/

扩展点

最近在动的地方

建议阅读顺序

  1. README.md - 了解项目定位和哲学
  2. src/diffusers/pipelines/diffusion_pipeline.py - 理解核心入口
  3. src/diffusers/schedulers/scheduling_utils.py - 理解调度器接口
  4. src/diffusers/models/unets/unet_2d_condition.py - 理解 UNet 实现
  5. src/diffusers/models/transformers/transformer_2d.py - 理解 DiT 实现
  6. src/diffusers/models/autoencoders/vae.py - 理解 VAE
  7. src/diffusers/models/attention_processor.py - 理解注意力机制
  8. src/diffusers/loaders/lora.py - 理解 LoRA
  9. examples/text_to_image/train_text_to_image.py - 理解训练流程
  10. tests/ - 理解测试覆盖和预期行为

三、本地跑起来(没有 GPU 的 Mac)

安装

  1. # 1. 用 uv 创建虚拟环境(AGENTS.md / CLAUDE.md 推荐方式)
  2. uv venv --python 3.10 && source .venv/bin/activate
  3. # 2. 先装 PyTorch macOS 版(CPU + MPS),不拉 CUDA 变体
  4. pip install torch torchvision pillow # 让 pip 自动选 arm64 wheel
  5. # 3. 以 editable 模式装 diffusers,跳过 CUDA-only 的可选依赖
  6. pip install -e "." # 不装 [torch] extras,避免被拉进 torch 的 cuda 子包
  7. # 4. 显式装 CPU/MPS 推理真正需要的依赖
  8. pip install transformers accelerate safetensors huggingface_hub scipy
  9. # 5. 验证安装与设备
  10. python -c "import diffusers, torch; print(diffusers.__version__); print('MPS', torch.backends.mps.is_available())"
  11. # 6. 可选:装测试冒烟用的轻量依赖(不装 xformers / triton,它们需要 CUDA)
  12. pip install pytest pytest-xdist filelock numpy"

哪些路径能真跑

['可在 CPU / MPS 上真正执行(慢但能跑):', '- tests/schedulers/:纯 tensor 调度逻辑,无 CUDA kernel 依赖', '- tests/others/:工具函数、配置、dummy 测试', '- tests/quantization/ 中 CPU 量化路径(需验证具体子目录)', "- examples/inference/ 中纯推理脚本(如 image_to_image.py、inpainting.py),配合 torch.float32 和 device='cpu' 或 'mps'", '- src/diffusers/schedulers/ 全部:DDIM、DDPM、FlowMatchEuler 等可在 CPU 上 step', '- src/diffusers/models/ 中模型 forward:UNet、DiT、VAE 用 PyTorch 原生 op,CPU/MPS 可执行,但大模型在 16GB 上会 OOM', '', '只能读代码或必须上 Colab T4 验证:', "- 任何 pipeline.to('cuda') 调用路径", '- AttnProcessor2_0、xformers、FlashAttention 相关分支(src/diffusers/models/attention_processor.py 中 CUDA-only 路径)', '- torch.compile 的 CUDA graph 回退', '- examples/ 下所有 train_*.py(训练脚本)', '- benchmarks/ 全部(需要真实 GPU 测吞吐)', '- docker/ 下所有 cuda 镜像对应路径', "- scripts/convert_*.py 中硬编码 device = 'cuda' 的脚本(如 convert_amused.py)"]

最小可运行

  1. # 1. 最小 CPU 推理:DDPM 无条件生成(tests 自带 fixture,不下载大模型)
  2. python -c "
  3. from diffusers import DDPMScheduler, UNet2DModel
  4. import torch
  5. model = UNet2DModel.from_pretrained('google/ddpm-cat-256').to('cpu')
  6. scheduler = DDPMScheduler.from_pretrained('google/ddpm-cat-256')
  7. scheduler.set_timesteps(10)
  8. noise = torch.randn(1, 3, model.config.sample_size, model.config.sample_size)
  9. input = noise
  10. for t in scheduler.timesteps:
  11. with torch.no_grad():
  12. noisy_residual = model(input, t).sample
  13. prev = scheduler.step(noisy_residual, t, input).prev_sample
  14. input = prev
  15. print('DDPM CPU OK, shape:', input.shape)"
  16. # 2. MPS 推理(若 Apple Silicon):同模型切到 mps
  17. python -c "
  18. from diffusers import DDPMScheduler, UNet2DModel
  19. import torch
  20. device = 'mps' if torch.backends.mps.is_available() else 'cpu'
  21. model = UNet2DModel.from_pretrained('google/ddpm-cat-256').to(device)
  22. scheduler = DDPMScheduler.from_pretrained('google/ddpm-cat-256')
  23. scheduler.set_timesteps(5)
  24. noise = torch.randn(1, 3, model.config.sample_size, model.config.sample_size, device=device)
  25. input = noise
  26. for t in scheduler.timesteps:
  27. with torch.no_grad():
  28. noisy_residual = model(input, t).sample
  29. prev = scheduler.step(noisy_residual, t, input).prev_sample
  30. input = prev
  31. print('MPS/CPU OK:', input.device, input.shape)"
  32. # 3. Scheduler 纯逻辑测试(不下载模型权重)
  33. python -c "
  34. from diffusers import DDIMScheduler, FlowMatchEulerDiscreteScheduler
  35. import torch
  36. sched = DDIMScheduler(num_train_timesteps=1000)
  37. sched.set_timesteps(20)
  38. print('DDIM timesteps:', sched.timesteps)
  39. flow = FlowMatchEulerDiscreteScheduler(num_train_timesteps=1000)
  40. print('FlowMatch sigmas:', flow.sigmas[:5])"
  41. # 4. 跑 tests/schedulers 子集(纯 CPU,无 GPU 依赖)
  42. pytest tests/schedulers/test_scheduling_ddim.py -x -q --no-header -rN
  43. # 5. 跑 tests/others 子集(配置、工具类测试)
  44. pytest tests/others/ -x -q --no-header -rN
  45. # 6. 验证 Hub 加载与 pipeline 构造(不推理,只验 from_pretrained 路径)
  46. python -c "
  47. from diffusers import DiffusionPipeline
  48. pipe = DiffusionPipeline.from_pretrained('google/ddpm-cat-256', device='cpu')
  49. print('Pipeline loaded:', type(pipe).__name__)"

测试

框架:pytest + pytest-xdist(见 pyproject.toml 与 tests/conftest.py)。目录:tests/ 下 615 个文件,分 models/、pipelines/、schedulers/、lora/、quantization/、modular_pipelines/、others/、single_file/、hooks/、remote/。只跑 CPU 子集:pytest tests/schedulers/ tests/others/ -x -q --no-header -rN,预计 5–15 分钟(MacBook Air 16GB)。跳过 GPU 测试:CI 用 @pytest.mark.skipIf(not torch.cuda.is_available()) 装饰器,本地不满足会自动 skip。大模型测试会下载权重,建议用 HF_HUB_OFFLINE=1 或 TRANSFORMERS_OFFLINE=1 限制网络。完整套件需 CUDA 机器,不在本地跑。

调试

CI

CI 系统:GitHub Actions(见 .github/workflows/)。跑哪些:make style(ruff lint + format 检查,见 pyproject.toml [tool.ruff] 配置)、make fix-copies(一致性检查)、pytest 分模型/调度器/pipeline 多矩阵跑、文档构建、自测试(utils/check_repo.py 等)。PR 会被卡住:1) make style 失败(行尾、import 排序、未用变量);2) make fix-copies 失败(src/diffusers 与 utils/copy_utils.py 生成的代码不一致);3) 缺少测试(utils/check_test_missing.py 会扫描新增模型/ pipeline 是否带测试);4) 提交信息格式(utils/log_reports.py 检查);5) 文档 TOC 不一致(utils/check_doc_toc.py)。PR 前本地必跑:make style && make fix-copies。

坑

四、维护者与社区

约 6–8 周一个 minor release:最近三个 release 分别是 0.40.0(2026-08-20)、0.39.0(2026-07-03)、0.38.0(2026-05-01),间隔约 6–7 周。提交热点集中在 src/diffusers(109 次)、examples/community(12)、docs/source(5)、tests/models(4)、.github/workflows(3),说明仍在活跃地加新模型/新 pipeline,同时补测试、文档和 CI。

谁角色依据
kashif核心维护者,负责模型加载与 Hub 集成(从 Issue #13279 assignee 推断)Issue #13279 的 assignees 字段为 ["kashif"]
stevhliu文档维护者,负责 docs 与 guidesIssue #14924 的 assignees 字段为 ["stevhliu"],标题为 Add a Hugging Face Jobs guide to the docs
EigenAx2Pi社区贡献者 / 被 assign 的量化相关维护Issue #14691 的 assignees 字段为 ["EigenAx2Pi"],涉及 GGUF 量化 bug
PrakshaaleJain社区贡献者 / 被 assign 的量化后端维护Issue #14705 的 assignees 字段为 ["PrakshaaleJain"],涉及 Comfy Quants 后端
Gelercatty社区贡献者,modular pipelines 相关PR #14941 标题含 Echo team joy future academy jd echo by @Gelercatty
HuggingFace 团队(未在证据中具名的 reviewer)PR 合并与 CI 维护PR #14944、#14789 等 CI/workflow 相关 PR 持续在 .github/workflows 提交(commit_hotspots 中 .github/workflows 3 次)

流程与 Review 风格

CONTRIBUTING.md 明确:先读 CODE_OF_CONDUCT.md 与 PHILOSOPHY.md;鼓励先在 Discord 打招呼;贡献路径从易到难:1) 论坛/Discord 问答 → 2) 开 Issue/Discussion → 3) 回答 Issue → 4) Good first issue → 5) 文档 → 6) Community Pipeline → 7) examples → 8) Good second issue → 9) 新 pipeline/model/scheduler。PR 前必须跑 make style 和 make fix-copies(AGENTS.md / CLAUDE.md 也强调)。pyproject.toml 配置 ruff 做 lint + isort,line-length 119。没有证据显示需要 CLA 或 DCO(LICENSE 是 Apache 2.0,但无 DCO 签名流程描述)。AI 辅助贡献有专门章节(AI-assisted and agentic contributions),要求 agent 先读 .ai/ 下的 skill 与 reference guide。

从近期 PR 看,review 较快且注重 CI 与格式:PR #14944(harden GitHub Actions workflows)和 #14789(Bump actions)显示对 CI 安全的重视;PR #14943、#14942 等 bugfix/modular 类 PR 在几天内就有响应。Issue 标签体系细(bug / good first issue / Good second issue / help wanted / contributions-welcome / roadmap / performance / needs-env-info / stale),说明维护者会主动分类和路由。未具名核心 reviewer 似乎对代码风格(ruff)、文档一致性(fix-copies)和测试覆盖有硬性要求。

渠道

这里的规矩

维护者现在最想要的帮助

五、切入方案

建议长期负责:src/diffusers/schedulers/ + tests/schedulers/ + benchmarks/ + examples/profiling/ 构成的「调度与性能基线」方向
scheduler 是扩散模型推理与训练的步进核心,直接影响速度、数值稳定性与跨后端一致性;该方向天然可在 CPU/MPS 上开发与验证;benchmarks/ 与 examples/profiling/ 是性能工程师的主场;从 scheduler 切入可逐步扩展到 models/ 的注意力处理器与 pipeline 执行层,路径清晰通向核心维护者。

它现在缺什么(你无 GPU 也能补)

缺口依据为什么是你
benchmarks/ 仅有 5 个模型 benchmark 脚本,缺少 CPU/MPS 可执行的轻量基准套件与自动化对比报告benchmarks/ 目录仅含 benchmarking_flux.py、benchmarking_ltx.py、benchmarking_sdxl.py、benchmarking_wan.py、benchmarking_utils.py、run_all.py、push_results.py 共 10 个文件;run_all.py 默认依赖 CUDA 环境,无 CPU/MPS 路径;无 pytest-benchmark 集成,无 CI 回归对比。GPU kernel 性能工程师擅长写 micro-benchmark 与回归对照;可在纯 CPU/MPS 上建立调度、模型 forward、pipeline step 的耗时基线,无需 CUDA 即可产出可复现数据。
tests/ 中缺少 scheduler 数值稳定性与跨后端(CPU/MPS/CUDA)一致性测试tests/schedulers/ 存在但覆盖以功能正确性为主;commit_hotspots 中 tests/models 仅 4 次、tests/schedulers 未进热点,说明近期无大规模测试补强;README 明确支持 Apple Silicon/MPS,但无证据表明有 MPS 数值一致性 CI。你擅长算子数值精度与后端对照;可补 scheduler step 的跨设备一致性测试、边界 timestep 测试,全部可在 CPU/MPS 上跑。
examples/profiling/ 仅有 profiling_pipelines.py 与 profiling_utils.py,缺少面向 CPU/MPS 的内存-时间剖析指南与可复用脚本examples/profiling/ 仅 2 个 Python 文件;README 强调 Apple Silicon 支持,但无对应 profiling 示例;commit_hotspots 中 examples/profiling 未出现,说明近期无人维护。性能工程师熟悉 torch.profiler、memory snapshot、peak memory 测量;可补 MPS/CPU 推理的 profiling 示例与最佳实践文档,无需 GPU。
docs/source 中缺少系统性的「CPU / Apple Silicon 性能优化」指南,仅有零散 optimization 页面docs/ 共 509 文件,README 提及 optimization/fp16 与 mps 指南,但无证据存在系统性 CPU/MPS 性能调优章节;commit_hotspots 中 docs/source 仅 5 次,且无相关标题提交。可基于自身经验撰写 CPU/MPS 推理优化、dtype 选择、注意力后端选择指南,纯文档工作无需 GPU。
scripts/ 下 95 个转换脚本缺少单元测试与输入输出 schema 校验scripts/ 目录 95 个文件,tree 中无对应 tests/scripts/ 子目录;check_repo.py 等工具脚本存在,但无证据表明转换脚本有自动化测试覆盖。可补转换脚本的轻量测试(mock 小模型、校验 state_dict key 映射),全部在 CPU 上完成,安全且不依赖 CUDA。
tests/quantization/ 中 CPU 量化路径覆盖不足,缺少与 CUDA 量化的对照基准tests/quantization/ 存在,但 commit_hotspots 未进前 10;EigenAx2Pi 与 PrakshaaleJain 被 assign 到量化相关 Issue(#14691、#14705),说明量化是活跃但缺人方向;CPU 量化路径可在 MacBook 上跑。你熟悉量化后端与部署;可补 CPU 量化测试、精度对照、文档,无需 GPU 即可贡献。
ModularDiffusionPipeline 测试覆盖与文档示例不足tests/modular_pipelines/ 仅 2 次提交热点;README 新增 modular_diffusers 入口;PR #14941 标题含 modular pipelines,说明是新功能但尚在早期;modular.md 参考指南存在但示例少。可补 modular pipeline 的使用示例、测试用例、迁移指南,全部 CPU 可跑,且能快速建立在新机制中的话语权。
CI 中缺少 CPU/MPS 上的快速冒烟测试矩阵,导致非 CUDA 贡献无法在 PR 阶段得到验证.github/workflows 仅 3 次提交热点;docker/ 含 diffusers-pytorch-cpu 镜像,但无证据表明 CI 在 PR 阶段跑 CPU 冒烟;fetch_torch_cuda_pipeline_test_matrix.py 命名暗示 CUDA 为主。可贡献 CPU 冒烟 workflow 与测试矩阵配置,提升项目对你的 PR 的验证能力,同时不依赖 GPU。

第 1–30 天:看懂并露面

第 31–60 天:稳定产出

第 61–90 天:接管一块

第一批 PR

题目范围为什么安全
Add CPU/MPS cross-backend consistency tests for DDIM and FlowMatchEuler schedulerstests/schedulers/ 新增测试文件,固定 seed 与 timestep,断言 CPU 与 MPS 上 prev_sample 的 max abs error < 1e-5纯测试代码,不修改源逻辑;可在 MacBook 上跑通;失败仅暴露现有数值偏差,不引入新 bug
Add benchmarking_scheduler.py with CPU/MPS/CUDA supportbenchmarks/ 新增脚本,复用 benchmarking_utils.py,测量 scheduler.step 的 wall-clock 与 peak memory,输出 JSON新增文件不改动现有代码;设备可选,默认 CPU;无副作用
Add profile_pipeline_mps.py example with torch.proprofiler traceexamples/profiling/ 新增脚本,跑一个轻量 pipeline(如 ddpm-cat-256)在 MPS 上,生成 trace.json 与 memory snapshot新增示例文件,用户主动运行;不修改库代码;文档说明清晰即可
Add missing return type annotations to schedulers/*.pysrc/diffusers/schedulers/ 中部分函数缺少 -> 返回类型注解,按 ruff 规则补全纯类型注解,不影响运行时;commit 历史中已有类似 PR([fix] Add return types #14874),说明社区接受此类贡献
Add docstring example to FlowMatchEulerDiscreteScheduler.stepsrc/diffusers/schedulers/scheduling_flow_match_euler_discrete.py 的 step 方法 docstring 补一个最小可运行示例纯文档字符串,不改变逻辑;示例可在 CPU 上验证

怎么知道自己站住了

风险与对策
  • MacBook 16GB 内存不足以跑大模型 forward,导致无法验证真实场景:使用 google/ddpm-cat-256 等小模型做功能验证;大模型验证留到 Colab T4 短时确认;在 PR 描述中明确标注测试设备与模型
  • MPS 与 CUDA 数值差异导致跨后端一致性测试失败,被误认为 bug:在测试文档中明确 MPS 与 CUDA 的误差来源(fp32 精度、算子实现差异);设置合理 tolerance;必要时将测试标记为 MPS-only 或 CUDA-only
  • scheduler 是核心模块,维护者审查严格,PR 易被要求大改:先从小范围测试/文档 PR 建立信任;修改 scheduler 逻辑前先在 issue 中提出方案并获得口头同意;附 benchmark 数据证明性能收益
  • CI 中缺少 CPU 资源,导致你的 CPU 冒烟测试无法在 PR 阶段运行:
  • benchmark 结果受 macOS 后台进程干扰,数据不稳定:多次运行取中位数;在脚本中记录系统负载与内存压力;在报告中注明运行条件
  • 贡献方向与核心维护者当前优先级不匹配,导致 PR 被搁置:在 Discord 或 issue 中先确认方向符合当前 release 目标;优先绑定已有 issue(如 #14691、#14705 量化相关)而非自创需求

六、怎么介入这个项目

社区入口:discord.gg/G7tWnz98XR · img.shields.io/discord/82381315959200153

建议顺序

成为长期维护者的路径
  • 首个 PR 落在 tests/schedulers/ 或 benchmarks/:纯新增文件、不改源逻辑,维护者审查门槛最低(参考已合并的 [fix] Add return types #14874、Consolidate torch device backend dispatch #14792)。
  • 第二个 PR 进入 src/diffusers/schedulers/ 修 NaN:附 CPU 复现脚本 + 回归测试,引用 Issue 编号,维护者对 scheduler 数值 bug 响应快(#14887 / #14888 / #14890 均 0 评论、无 assignee、无 open PR)。
  • 第三个 PR 进入 tests/ 补跨后端一致性:与 scheduler 修复形成组合拳,建立「调度与性能基线」方向的可见度。
  • 后续可申请维护 tests/schedulers/ + benchmarks/ + examples/profiling/ 方向,或牵头 CPU/MPS 冒烟 CI 改进。

七、任务卡

任务 1优先 · medium · CPU · 2 个晚上

修复 UniPCMultistepScheduler 在 sigma=0 末步返回 NaN/inf 问题(solver_type=bh1、lower_order_final=False、predict_x0=False)

无人认领0 条评论更新 2026-09-27

用到的专长:GPU 性能工程师对数值精度与算子边界条件敏感;本任务核心是 scheduler 的 dtype / 边界步逻辑,与 CUDA 无关。

目标:让 UniPCMultistepScheduler 在 final_sigmas_type="zero"(默认配置)下,所有 solver_type / lower_order_final / predict_x0 组合在末步都返回有限数,行为与 DPMSolverMultistepScheduler 一致。

为什么值得长期做:Scheduler 是 Diffusers 三大核心抽象之一,数值正确性直接影响所有使用 UniPC 的 pipeline。修复 sigma=0 末步的 NaN/inf 属于纯逻辑层,不依赖 CUDA,且能建立你在调度器方向的可信度——这是 ownership_target 的核心区域。

怎么介入:Issue 14887 无 assignee、无 open PR、0 条评论,2026-09-27 创建,状态新鲜。可直接认领。
第一个 PR 的边界:第一个 PR 只改 src/diffusers/schedulers/scheduling_unipc_multistep.py 的 step、multistep_uni_p_bh_update、__init__,并新增 tests/schedulers/test_scheduler_unipc.py 中的 test_final_step_finiteness。不动其他 scheduler。
第一步:在本地 MacBook 上用 Issue 提供的复现脚本跑通(CPU,float32),确认 bh1 末步 NaN、lower_order_final=False 末步 inf、predict_x0=False 末步 NaN 三种现象。
本机怎么复现 / 验证:python -c "import torch; from diffusers import UniPCMultistepScheduler; s=UniPCMultistepScheduler(); s.set_timesteps(10); x=torch.rand(1,3,8,8); [s.step(x*t/(t+1),t,x) for t in s.timesteps]" # CPU 上复现 NaN/inf;然后 pytest tests/schedulers/test_scheduler_unipc.py -x 验证修复。
认领留言(英文,可直接贴到 Issue)
I'd like to take this. The root cause is that UniPC's final step to sigma=0 runs at full order with inf lambda, and only the default (bh2/lower_order_final=True/predict_x0=True) path survives. I'll align with DPMSolverMultistepScheduler: force this_order=1 on the last step when final_sigmas_type=='zero', guard the B_h correction in multistep_uni_p_bh_update to D1s is not None, and raise ValueError for predict_x0=False + final_sigmas_type='zero' in __init__. I'll add a finiteness test covering all solver_type/lower_order_final/predict_x0 combinations. Reproduction and tests run on CPU. PR ready in ~2 evenings. Does this scope look right?
大致实施方案
  • 阅读 src/diffusers/schedulers/scheduling_unipc_multistep.py 的 step、multistep_uni_p_bh_update、multistep_uni_c_bh_update 三个函数。
  • 在 step 中加入末步强制一阶逻辑:当 final_sigmas_type == "zero" 且当前步是末步时,this_order = 1(与 DPMSolverMultistepScheduler.step 对齐)。
  • 在 multistep_uni_p_bh_update 中,仅当 D1s is not None 时才减去 B_h * pred_res 项,避免 -inf * 0 = NaN。
  • 在 __init__ 中加入 ValueError:predict_x0=False 与 final_sigmas_type="zero" 组合非法(与 DPMSolverMultistepScheduler 一致)。
  • 在 tests/schedulers/test_scheduler_unipc.py 新增 test_final_step_finiteness,覆盖 bh1/bh2 × lower_order_final True/False × predict_x0 True/False × 多步数,断言 prev_sample.isfinite().all()。
  • 本地跑 pytest tests/schedulers/test_scheduler_unipc.py -x,确认全部通过。
可能涉及的目录或文件
  • src/diffusers/schedulers/scheduling_unipc_multistep.py(step、multistep_uni_p_bh_update、multistep_uni_c_bh_update、__init__)
  • tests/schedulers/test_scheduler_unipc.py(新增测试)
验收方式
  • 本地 CPU 复现脚本(Issue 正文)不再产出 NaN/inf。
  • pytest tests/schedulers/test_scheduler_unipc.py -x 全部通过(含新增 test_final_step_finiteness)。
  • ruff check src/diffusers/schedulers/scheduling_unipc_multistep.py 无新增 warning。
开工前问题与风险

向维护者确认

  • 是否同意与 DPMSolverMultistepScheduler 对齐的修复方案(末步强制一阶 + __init__ 拒绝 predict_x0=False + final_sigmas_type="zero")?
  • 是否需要同时处理 final_sigmas_type="sigma_min" 路径(该路径末步 h=0 但一阶更新安全,Issue 未要求)?

风险

  • 改动 step 主路径可能影响现有 pipeline 的数值回归,需确保只在末步触发新逻辑。
  • __init__ 加 ValueError 是破坏性变更,需确认没有现有配置依赖该组合。
任务 2优先 · easy · CPU · 1 个晚上

修复 UniPCMultistepScheduler 在 float64 默认 dtype 下 torch.linalg.solve 的 dtype 不匹配崩溃

无人认领0 条评论更新 2026-09-27

用到的专长:dtype 传播与算子边界是 GPU 性能工程师日常;本任务只需 CPU 验证。

目标:让 UniPCMultistepScheduler 在 torch.set_default_dtype(torch.float64) 下正常运行,rks 单位项继承 sigmas 的 dtype。

为什么值得长期做:与 #14887 同属 UniPCScheduler 数值鲁棒性方向,修复极小(两处 torch.ones(..., device=device) → torch.ones_like(h)),但影响所有在 float64 环境下使用 UniPC 的用户。维护者对这类 dtype 边界 bug 接受度高。

怎么介入:Issue 14888 无 assignee、无 open PR(linked PR #14920 和 #1 都已 closed,未合并),0 条评论,2026-09-27 创建。可重新认领。
第一个 PR 的边界:第一个 PR 只改两处 torch.ones 调用并新增一个测试函数。不动其他逻辑。
第一步:在本地 MacBook 跑 Issue 复现脚本(CPU,float64 默认 dtype),确认 RuntimeError。
本机怎么复现 / 验证:python -c "import torch; torch.set_default_dtype(torch.float64); from diffusers import UniPCMultistepScheduler; s=UniPCMultistepScheduler(); s.set_timesteps(10); x=torch.rand(1,3,8,8); [s.step(x*t/(t+1),t,x) for t in s.timesteps]" # CPU 复现 RuntimeError;然后 pytest tests/schedulers/test_scheduler_unipc.py -x 验证。
认领留言(英文,可直接贴到 Issue)
I'll take this. The fix is two one-liners: replace torch.ones((), device=device) with torch.ones_like(h) in both multistep_uni_p_bh_update and multistep_uni_c_bh_update so the unit entry inherits the sigmas' dtype. I'll add a test_default_dtype_float64 test. CPU-only reproduction and verification. PR ready in ~1 evening. Should I combine this with #14887 in one PR or keep separate?
大致实施方案
  • 定位 src/diffusers/schedulers/scheduling_unipc_multistep.py 中 multistep_uni_p_bh_update 和 multistep_uni_c_bh_update 两处 torch.ones((), device=device)。
  • 改为 torch.ones_like(h),使单位项与 sigmas 同 dtype。
  • 在 tests/schedulers/test_scheduler_unipc.py 新增 test_default_dtype_float64,用 torch.set_default_dtype(torch.float64) 跑完整 step 循环。
  • 本地跑 pytest tests/schedulers/test_scheduler_unipc.py -x 确认通过。
可能涉及的目录或文件
  • src/diffusers/schedulers/scheduling_unipc_multistep.py(multistep_uni_p_bh_update、multistep_uni_c_bh_update)
  • tests/schedulers/test_scheduler_unipc.py(新增测试)
验收方式
  • Issue 复现脚本在 float64 默认 dtype 下不再抛 RuntimeError。
  • pytest tests/schedulers/test_scheduler_unipc.py -x 全部通过。
  • ruff check 无新增 warning。
开工前问题与风险

向维护者确认

  • 是否同意 torch.ones_like(h) 方案?Issue 作者已验证 540 配置 bit-identical。

风险

  • 几
  • 乎
  • 无
  • 风
  • 险
  • :
  • t
  • o
  • r
  • c
  • h
  • .
  • o
  • n
  • e
  • s
  • _
  • l
  • i
  • k
  • e
  • (
  • h
  • )
  • 在
  • f
  • l
  • o
  • a
  • t
  • 3
  • 2
  • 默
  • 认
  • 下
  • 与
  • 原
  • 来
  • b
  • i
  • t
  • -
  • i
  • d
  • e
  • n
  • t
  • i
  • c
  • a
  • l
  • 。
任务 3优先 · medium · CPU · 2 个晚上

修复 DPMSolver 与 UniPC 在 final_sigmas_type=sigma_min + 转换 sigma 时末步 h=0 返回 NaN

无人认领0 条评论更新 2026-09-28

用到的专长:scheduler 数值边界与 dtype 分析是性能工程师强项;纯 CPU 可验证。

目标:让 DPMSolverMultistepScheduler(solver_order=3 或 heun)和 UniPCMultistepScheduler(order=3, lower_order_final=False)在使用 converted sigmas(Karras/exp/beta/lu_lambdas)+ final_sigmas_type="sigma_min" 时,末步 h=0 不返回 NaN。

为什么值得长期做:与 #14887 / #14888 构成 scheduler 数值正确性三连击。本任务涉及 DPMSolverMultistepScheduler 和 UniPCMultistepScheduler 两个最常用 scheduler 的 converted sigma 路径,影响面广。Issue 作者已给出完整根因分析,只需落地修复。

怎么介入:Issue 14890 无 assignee、无 open PR、0 条评论,2026-09-28 创建,状态新鲜。可直接认领。
第一个 PR 的边界:第一个 PR 只改两个 scheduler 的 step 函数末步检测逻辑,并新增 finiteness 测试。不动 set_timesteps。
第一步:在本地 MacBook 用 Issue 的 karras 3-steps 配置复现 NaN(CPU,float32)。
本机怎么复现 / 验证:python -c "import torch; from diffusers import DPMSolverMultistepScheduler; s=DPMSolverMultistepScheduler(use_karras_sigmas=True); s.set_timesteps(3); s.config.final_sigmas_type='sigma_min'; x=torch.rand(1,3,8,8); [s.step(x,t,x) for t in s.timesteps]" # CPU 复现 NaN;然后 pytest 验证。
认领留言(英文,可直接贴到 Issue)
I'd like to take this. The root cause is that converted sigmas already end at sigma_min, so appending sigma_min again makes the last step h=0, and the third-order/heun updates divide by h. I'll detect h=0 in DPMSolverMultistepScheduler.step and UniPCMultistepScheduler.step and force this_order=1 there (mirroring the existing final_sigmas_type=='zero' guard). I'll add finiteness tests for Karras/exp/beta/lu_lambdas + sigma_min. CPU-only reproduction. PR in ~2 evenings. Should I also cover DPMSolverSinglestep/DEIS/SASolver/Euler in the same PR?
大致实施方案
  • 阅读 DPMSolverMultistepScheduler.step 中已有的 final_sigmas_type=="zero" 强制一阶逻辑,理解为何 "sigma_min" 路径未覆盖。
  • 在 DPMSolverMultistepScheduler.step 中检测末步 h=0(或 sigmas[-1]==sigmas[-2])时强制 this_order=1。
  • 在 UniPCMultistepScheduler.step 中同样检测末步 h=0 并强制 this_order=1。
  • 在 tests/schedulers/test_scheduler_dpm_multi.py 和 test_scheduler_unipc.py 新增 converted-sigma + sigma_min 末步 finiteness 测试。
  • 本地跑 pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py -x。
可能涉及的目录或文件
  • src/diffusers/schedulers/scheduling_dpm_solver_multistep.py(step)
  • src/diffusers/schedulers/scheduling_unipc_multistep.py(step)
  • tests/schedulers/test_scheduler_dpm_multi.py、test_scheduler_unipc.py(新增测试)
验收方式
  • karras 3-steps + final_sigmas_type=sigma_min 末步不再 NaN。
  • pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py -x 全部通过。
开工前问题与风险

向维护者确认

  • 是否同意在 step 中检测 h=0 并强制一阶,而不是在 set_timesteps 中避免重复 sigma?
  • 是否需要同时处理 DPMSolverSinglestep / DEIS / SASolver / EulerDiscrete(Issue 提到它们也有重复 sigma)?

风险

  • 强
  • 制
  • 一
  • 阶
  • 可
  • 能
  • 改
  • 变
  • 末
  • 步
  • 数
  • 值
  • 精
  • 度
  • ,
  • 但
  • I
  • s
  • s
  • u
  • e
  • 已
  • 分
  • 析
  • 一
  • 阶
  • 更
  • 新
  • 在
  • h
  • =
  • 0
  • 时
  • 是
  • 精
  • 确
  • 的
  • (
  • s
  • a
  • m
  • p
  • l
  • e
  • u
  • n
  • c
  • h
  • a
  • n
  • g
  • e
  • d
  • )
  • 。
任务 4可选 · easy · Mac · 1 个晚上

修复 QwenImage21 FlexAttention mask 在 MPS 上因 ~ 操作符编译失败

无人认领0 条评论更新 2026-09-28

用到的专长:MPS 后端适配与编译期问题属于端侧部署范畴;候选人熟悉 TensorRT/量化,MPS 是同类技能。

目标:让 build_qwenimage21_block_causal_mask 在 MPS + torch.compile(fullgraph=True) 下正常编译,不抛 NotImplementedError。

为什么值得长期做:README 明确写 Apple Silicon 支持,MPS 是 Diffusers 一等公民。本任务修复 MPS 后端 FlexAttention 的编译失败,直接提升 Mac 用户体验,且改动极小(~ → torch.logical_not),风险低。

怎么介入:Issue 14889 无 assignee、无 open PR、0 条评论,2026-09-28 创建。可直接认领。
第一个 PR 的边界:第一个 PR 只改 transformer_qwenimage21.py 一行,并新增一个 MPS-only 测试。不动其他模型。
第一步:在本地 MacBook(MPS 可用)跑 Issue 复现脚本,确认 NotImplementedError。
本机怎么复现 / 验证:python -c "import torch; from torch.nn.attention.flex_attention import flex_attention; from diffusers.models.transformers.transformer_qwenimage21 import build_qwenimage21_block_causal_mask; device='mps'; ids=torch.full((97,),-1,dtype=torch.long,device=device); valid=torch.ones(2,97,dtype=torch.bool,device=device); mask=build_qwenimage21_block_causal_mask(ids,valid,2,device); q=torch.randn(2,2,mask.shape[-1],128,device=device); flex_attention(q,q,q,block_mask=mask)" # MPS 上复现 NotImplementedError;修复后通过。
认领留言(英文,可直接贴到 Issue)
I'll take this on my Apple Silicon Mac. The fix is a one-liner: replace ~is_padding with torch.logical_not(is_padding) in build_qwenimage21_block_causal_mask. The MPS FlexAttention lowerer rejects aten.bitwise_not but accepts torch.logical_not for boolean tensors. I'll add a MPS-only regression test using torch.compile(fullgraph=True). PR ready in ~1 evening. Any preference on where the test should live?
大致实施方案
  • 定位 src/diffusers/models/transformers/transformer_qwenimage21.py 中 build_qwenimage21_block_causal_mask 的 return allowed & ~is_padding。
  • 改为 return allowed & torch.logical_not(is_padding)。
  • 在 tests/models/ 新增 test_qwenimage21_mps_flex_attention(或扩展现有测试),在 MPS 可用时跑 torch.compile(fullgraph=True) 验证。
  • 本地在 MPS 上跑该测试确认通过。
可能涉及的目录或文件
  • src/diffusers/models/transformers/transformer_qwenimage21.py(build_qwenimage21_block_causal_mask)
  • tests/models/test_transformer_qwenimage21.py 或新增测试文件(新增)
验收方式
  • 本地 MPS 复现脚本不再抛 NotImplementedError。
  • 新增测试在 MPS 上通过,在 CPU 上 skip。
开工前问题与风险

向维护者确认

  • 是否同意 torch.logical_not 方案?Issue 作者已验证 MPS 上可行。
  • 是否需要同时检查其他模型中类似的 ~boolean 模式?

风险

  • 几
  • 乎
  • 无
  • 风
  • 险
  • :
  • t
  • o
  • r
  • c
  • h
  • .
  • l
  • o
  • g
  • i
  • c
  • a
  • l
  • _
  • n
  • o
  • t
  • 与
  • ~
  • 在
  • 语
  • 义
  • 上
  • 对
  • b
  • o
  • o
  • l
  • e
  • a
  • n
  • t
  • e
  • n
  • s
  • o
  • r
  • 完
  • 全
  • 等
  • 价
  • 。
任务 5可选 · easy · CPU · 1 个晚上

修复 numpy_to_pil 对灰度图 singleton 空间维度的错误 squeeze

无人认领1 条评论更新 2026-09-29

用到的专长:图像预处理与 tensor 形状分析是部署工程师日常;纯 CPU 可验证。

目标:让 numpy_to_pil 在 HWC 灰度分支只 squeeze channel 维度,保留 singleton 的 H 或 W。

为什么值得长期做:图像预处理是 pipeline 最末一环,numpy_to_pil 的 squeeze 错误会导致 H×1 或 1×W 灰度图被压成 1×1。Issue 作者已给出完整复现与修复方案(squeeze(-1)),维护者尚未回应。纯工具函数修复,风险极低。

怎么介入:Issue 14901 无 assignee、无 open PR、1 条评论(作者自己),2026-09-29 创建。可直接认领。
第一个 PR 的边界:第一个 PR 只改 pil_utils.py 和 image_processor.py 的 squeeze 调用,并新增测试。examples 修复视维护者反馈决定是否同 PR。
第一步:在本地 MacBook 跑 Issue 复现脚本(CPU),确认 (1,7,1) 被压成 (1,1)。
本机怎么复现 / 验证:python -c "import numpy as np; from diffusers.utils.pil_utils import numpy_to_pil; x=np.zeros((1,7,1),dtype=np.float32); y=numpy_to_pil(x)[0]; print(y.size)" # 期望 (7,1),实际 (1,7);修复后通过。
认领留言(英文,可直接贴到 Issue)
I'll take this. The fix is to replace image.squeeze() with image.squeeze(-1) in the grayscale branch of numpy_to_pil, and the same pattern in image_processor.py (2 occurrences). I'll also review the unrestricted squeeze in the two Flux2 DreamBooth examples. I'll add a regression test covering (1,1,1), (1,7,1), (7,1,1), and ordinary grayscale/RGB/RGBA cases. CPU-only reproduction. PR in ~1 evening. Should I fix the examples in the same PR or file a separate issue for them?
大致实施方案
  • 定位 src/diffusers/utils/pil_utils.py 中 numpy_to_pil 的 grayscale 分支 Image.fromarray(image.squeeze(), mode="L")。
  • 改为 Image.fromarray(image.squeeze(-1), mode="L")。
  • 检查 src/diffusers/image_processor.py 中两处相同模式并同步修复。
  • 检查 examples/dreambooth/train_dreambooth_lora_flux2_img2img.py 和 train_dreambooth_lora_flux2_klein_img2img.py 中 CHW 灰度 squeeze,改为 channel-axis-specific。
  • 在 tests/ 新增 test_numpy_to_pil_singleton_dims,覆盖 (1,1,1)、(1,7,1)、(7,1,1)、(1,7,3) 等 case。
  • 本地跑 pytest tests/ -k numpy_to_pil 或对应测试文件。
可能涉及的目录或文件
  • src/diffusers/utils/pil_utils.py(numpy_to_pil)
  • src/diffusers/image_processor.py(两处)
  • examples/dreambooth/train_dreambooth_lora_flux2_img2img.py、train_dreambooth_lora_flux2_klein_img2img.py(需 review)
  • tests/utils/test_pil_utils.py 或新增测试文件(新增)
验收方式
  • Issue 复现脚本输出 (7,1) 而非 (1,7)。
  • 新增测试全部通过。
  • ruff check 无新增 warning。
开工前问题与风险

向维护者确认

  • 是否同意 squeeze(-1) 方案?Issue 作者已验证 9 cases。
  • examples 中的 CHW squeeze 是否一并修,还是留给示例作者?

风险

  • 极
  • 低
  • :
  • s
  • q
  • u
  • e
  • e
  • z
  • e
  • (
  • -
  • 1
  • )
  • 只
  • 在
  • c
  • h
  • a
  • n
  • n
  • e
  • l
  • 维
  • 度
  • 为
  • 1
  • 时
  • 生
  • 效
  • ,
  • 对
  • 普
  • 通
  • R
  • G
  • B
  • 图
  • 无
  • 影
  • 响
  • 。
任务 6可选 · easy · Mac · 1 个晚上

新增 DDIM 与 FlowMatchEuler 调度器的 CPU/MPS 跨后端一致性测试

无人认领6 条评论stalecontributions-welcome更新 2026-10-05

用到的专长:跨后端数值一致性测试是性能工程师写 micro-benchmark 的日常;Mac 上可完整验证。

目标:在 tests/schedulers/ 新增 test_cross_backend_consistency.py,覆盖 DDIMScheduler 和 FlowMatchEulerDiscreteScheduler 在 CPU vs MPS 上的 step 输出一致性。

为什么值得长期做:gaps 明确指出 tests/schedulers/ 缺少跨后端数值一致性测试。本任务在 CPU/MPS 上固定 seed 与 timestep,断言 prev_sample 的 max abs error < 1e-5,填补 README 声称的 Apple Silicon 支持与实际测试覆盖之间的鸿沟。与 ownership_target 的「调度与性能基线」方向完全对齐。

怎么介入:Issue 12760 是 contributions-welcome 标签的 stale issue,但正文讨论的是 cache_context 缺失,与本任务不直接相关。本任务作为 proposal 卡,source_url 指向仓库 commits,engagement 写明先开 Issue 提案。
第一个 PR 的边界:第一个 PR 只新增 tests/schedulers/test_cross_backend_consistency.py,不改任何源文件。
第一步:在本地 MacBook 确认 MPS 后端可用(torch.backends.mps.is_available()),阅读 tests/schedulers/ 现有测试结构。
本机怎么复现 / 验证:pytest tests/schedulers/test_cross_backend_consistency.py -x -v # 在 Mac 上直接运行新增测试,MPS 可用则跑 CPU vs MPS 对比,否则 skip。
认领留言(英文,可直接贴到 Issue)
[Proposal] Add CPU/MPS cross-backend consistency tests for schedulers

I'd like to add tests/schedulers/test_cross_backend_consistency.py that fixes seed + timestep and asserts max abs error < 1e-5 between CPU and MPS for DDIMScheduler and FlowMatchEulerDiscreteScheduler. This fills the gap noted in the repo's Apple Silicon support claim vs actual test coverage. All tests run on Mac (MPS) with CPU fallback. Should I open a dedicated issue first, or is a PR with the tests sufficient?
大致实施方案
  • 新建 tests/schedulers/test_cross_backend_consistency.py。
  • 对每个 scheduler(DDIM、FlowMatchEuler),固定 seed、timestep、sample,分别在 CPU 和 MPS 上跑 step,断言 max abs error < 1e-5。
  • MPS 不可用时自动 skip(pytest.mark.skipif)。
  • 本地在 MPS 上跑 pytest tests/schedulers/test_cross_backend_consistency.py -x。
可能涉及的目录或文件
  • tests/schedulers/test_cross_backend_consistency.py(新增)
验收方式
  • 本地 MPS 上测试通过。
  • CPU 上测试通过(MPS skip)。
  • pytest tests/schedulers/ -x 全部通过。
开工前问题与风险

向维护者确认

  • 是否同意 max abs error < 1e-5 的阈值?是否需要覆盖更多 scheduler(DDPM、LCMScheduler)?

风险

  • 极
  • 低
  • :
  • 纯
  • 新
  • 增
  • 测
  • 试
  • ,
  • 不
  • 改
  • 源
  • 逻
  • 辑
  • 。
  • 若
  • M
  • P
  • S
  • 数
  • 值
  • 偏
  • 差
  • 超
  • 过
  • 阈
  • 值
  • ,
  • 会
  • 暴
  • 露
  • 现
  • 有
  • 问
  • 题
  • 而
  • 非
  • 引
  • 入
  • 新
  • b
  • u
  • g
  • 。