← 所有项目

Contribution Tasks

modelscope/DiffSynth-Engine

候选人的 GPU kernel / 训练性能专长与仓库的 attention 后端、scheduler/sampler、LoRA、offload 等模块高度契合;但当前无 GPU,只能落地在纯 Python / CPU / MPS 可验证的测试、文档与兼容性矩阵类贡献,不能做需要 CUDA 运行时的性能调优或 kernel 开发。

当前方向:项目处于快速模型覆盖扩张期(v0.4→v0.7 半年内支持 Wan2.2、Qwen-Image、Qwen-Image-Edit、Wan2.2-S2V、Flux.2、Z-Image、ODTSR),主线是 diffsynth_engine/models 与 pipelines 的新模型接入,以及 NPU/Ascend 平台适配。commit 热点集中在 models、tests/data、pipelines、test_pipelines。

★ 433Fork 516 个候选任务Gemini:LongCat-2.0

更新于 2026-10-08T07:37:23+00:00 · 打开仓库 ↗

一、项目定位

DiffSynth-Engine 是 ModelScope(阿里达摩院 MuseAI)团队维护的高性能扩散模型推理引擎,核心目标是为 DiT 类图像/视频生成模型(Flux、Wan、Qwen-Image、Z-Image 等)提供一套自研、可控、可裁剪的推理管线,减少对 diffusers/k-diffusion 等外部库的依赖。README 明确指出它“carefully re-implemented key components such as sampler and scheduler, without introducing external dependencies on libraries like k-diffusion, ldm, or sgm”,因此它不是简单的 diffusers wrapper,而是把 scheduler、采样器、并行策略、量化/卸载、LoRA 加载都重写了一遍。 项目当前处于快速迭代阶段:从 v0.4.0(2025-08-01)到 v0.7.0(2026-01-19),半年内连续发布多个大版本,每次都以“支持新模型”为主(Wan2.2、Qwen-Image、Qwen-Image-Edit、Wan2.2-S2V、Flux.2-Klein、Z-Image-Omni-Base、ODTSR),说明它正在以“模型覆盖广度 + 推理性能”双轴扩张。代码提交热点集中在 diffsynth_engine/models(10 次)、tests/data(10 次)、diffsynth_engine/pipelines(6 次)、tests/test_pipelines(5 次),说明模型接入与管线测试是主线工作。 与同类项目的差异在于:它同时覆盖图像生成、图像编辑、视频生成、语音驱动视频、超分等多任务,且明确支持 Windows/macOS(Apple Silicon)/Linux 跨平台,对 Apple Silicon 有原生适配(README 列为 requirement)。依赖约束为 torch 2.6–2.8,Python ≥3.10,并引入 yunchang(仅 Linux)做长序列并行。整体定位是“面向生产与研究的扩散模型推理底座”,而非单纯的 demo 仓库。

解决什么问题、给谁用

解决的核心问题:在显存受限或跨平台(含 Apple Silicon)环境下,如何高效、低依赖地运行最新一代扩散模型(尤其是 DiT 架构的 Flux、Wan、Qwen-Image 等),同时支持 FP8/INT8 量化、LoRA、并行推理(flux_parallel、wan_parallel 等 examples 可见)和多种 offloading 策略。 目标用户:1) 需要在端侧/工作站部署扩散模型的研究者与工程师;2) 需要二次开发推理管线、自定义 scheduler/sampler 的性能优化工程师;3) 需要在 Apple Silicon 上做原型验证、再迁移到 NVIDIA GPU 的团队;4) 需要支持多种国产/开源模型(Wan、Qwen-Image、Z-Image)的 ModelScope 生态用户。 典型使用场景:从 ModelScope/HuggingFace 拉取模型(fetch_model),通过 FluxPipelineConfig/WanPipelineConfig 等配置加载,调用 FluxImagePipeline/WanVideoPipeline 等执行 text-to-image、image-to-video、speech-to-video、LoRA 编辑、超分(ODTSR)等任务;也可用 examples 中的 parallel 脚本做多卡并行推理,或用 memory_predict_model 预估显存占用。

同类项目与差别

核心能力

能力在哪成熟度
自研 scheduler/sampler(不依赖 k-diffusion/ldm/sgm)diffsynth_engine/algorithm/ 与 pipelines/ 中各采样逻辑成熟
Flux 系列图像生成(Flux.1 Dev、Flux.2-Klein)diffsynth_engine/models/ 中 flux 相关模型定义,examples/flux_*.py成熟
Wan 系列视频生成(Wan2.2、Wan2.2-S2V 语音驱动视频)diffsynth_engine/models/ 中 wan 相关,examples/wan_*.py成熟
Qwen-Image / Qwen-Image-Edit 图像生成与编辑diffsynth_engine/models/ 中 qwen 相关,examples/qwen_image_*.py成熟
Z-Image(Z-Image-Omni-Base)支持diffsynth_engine/models/ 中 z_image 相关实验
ODTSR 超分工具(Qwen Image Upscaler)examples/odtsr_upscale_tool_example.py,diffsynth_engine/tools/实验
LoRA 加载(兼容 CivitAI 格式)diffsynth_engine/models/ 中 lora 加载逻辑,examples/flux_lora.py、wan_lora.py成熟
FP8/INT8 量化与多种 offloading 策略diffsynth_engine/utils/ 与 models/ 中的量化/offload 实现成熟
多卡并行推理(flux_parallel、wan_parallel)examples/flux_parallel.py、examples/wan_parallel.py,diffsynth_engine/utils/ 中并行相关成熟
Apple Silicon (macOS) 跨平台支持README 列为 requirement,代码中 device 处理与 Apple Silicon 适配成熟
Linux 长序列并行(yunchang)pyproject.toml 中 yunchang 依赖(仅 Linux),diffsynth_engine/utils/ 中并行逻辑实验
显存预测工具(memory_predict_model)examples/memory_predict_model/实验

阶段:快速迭代的成长期(v0.4.0 → v0.7.0 半年内多次大版本,以模型覆盖和性能优化为主线,API 仍有 breaking change 记录,如 v0.4.0 的 from_pretrained 改进)。

技术栈:Python ≥3.10;PyTorch 2.6–2.8;torchvision, numpy, scipy, einops;safetensors, gguf(模型加载);modelscope(模型下载);tokenizers, sentencepiece, ftfy, regex(文本处理);torchsde(随机微分方程采样);pillow, imageio[ffmpeg], opencv-python, moviepy, librosa, scikit-image, trimesh(图像/视频/音频/3D I/O);onnxruntime(可选推理后端);yunchang(仅 Linux,长序列并行);flufl.lock(文件锁);setuptools + setuptools_scm(构建);ruff(lint/format);pytest(测试);pre-commit(提交前检查)

规模:仓库根目录约 13 个顶层条目;diffsynth_engine 包内 211 个文件(含子目录 algorithm/conf/configs/kernels/models/pipelines/processor/tokenizers/tools/utils);tests 目录 156 个文件(含 test_algorithm/test_models/test_pipelines/test_tokenizers/test_tools 等);examples 目录 30 个文件;docs 仅 2 个教程(tutorial.md、tutorial_zh.md)。最近提交时间 2026-05-27,最新 release v0.7.0 于 2026-01-19 发布,v0.6.0 于 2025-09-09,v0.5.0 于 2025-08-27,v0.4.0 于 2025-08-01,显示约每月 1–2 次 release 的节奏。commit_hotspots 显示 models、tests/data、pipelines、test_pipelines 为最活跃区域。README 未标注 star 数,具体 star/fork/issue 数字需进一步查看 GitHub API 确认。

二、架构与代码地图

DiffSynth-Engine 采用经典的分层推理架构,自上而下分为五层: 1. 接入层(Entry Layer):由 pipelines/ 和 tools/ 构成,提供面向用户的统一推理接口。BasePipeline 定义了所有管线的公共生命周期(模型加载、LoRA 绑定、offload 调度、__call__ 入口),各具体管线(FluxImagePipeline、WanVideoPipeline、QwenImagePipeline 等)继承并实现特定模型的加载与推理逻辑。tools/ 层则提供 FluxInpaintingTool、FluxOutpaintingTool、QwenImageUpscalerTool 等高级工具,封装了 ControlNet、IP-Adapter、Redux 等复杂工作流。 2. 调度层(Scheduling Layer):由 algorithm/ 和 configs/ 构成,负责采样算法与配置管理。algorithm/noise_scheduler/ 实现了 Flow Match(flux 系列)和 Stable Diffusion(SD 系列)两套噪声调度器,包含 flow_beta、flow_ddim、recifited_flow、beta、ddim、karras 等多种调度策略。algorithm/sampler/ 实现了对应的采样器(Euler、DPM-Solver++、DEIS、Brownian Tree 等)。configs/ 通过 dataclass 定义了所有管线的配置结构(FluxPipelineConfig、WanPipelineConfig、QwenImagePipelineConfig 等),以及 AttentionConfig、OptimizationConfig、ParallelConfig 等性能相关配置。 3. 执行层(Execution Layer):由 models/ 构成,包含所有模型实现。models/base.py 定义了 PreTrainedModel 和 StateDictConverter 基类,所有模型继承此基类。models/basic/ 提供基础组件(Attention、LoRA、Timestep、TransformerHelper、UNetHelper、VideoSparseAttention)。models/flux/、models/wan/、models/qwen_image/、models/sd/、models/sdxl/、models/sd3/、models/flux2/、models/z_image/、models/hunyuan3d/ 分别实现了各模型家族的 DiT、VAE、TextEncoder、ControlNet 等组件。models/text_encoder/ 和 models/vae/ 提供共享的文本编码器和 VAE 实现。 4. Kernel 层(Kernel Layer):由 kernels/ 和 models/basic/attention.py 构成,负责底层计算优化。models/basic/attention.py 实现了统一的 Attention 接口,支持 FA2/FA3/FA4、SDPA、xFormers、Sage、Sparge、Aiter、VideoSparseAttention 等多种后端,通过 utils/flag.py 动态检测可用性。models/basic/video_sparse_attention.py 实现了视频稀疏注意力机制,支持分布式场景。kernels/ 目录(当前为空或极少文件)预留了自定义 CUDA/Triton kernel 的扩展位置。 5. 工具层(Utility Layer):由 utils/、tokenizers/、processor/ 构成,提供通用工具。utils/ 包含模型加载(loader.py)、量化(fp8_linear.py、gguf.py)、offload(offload.py)、并行(parallel.py、process_group.py)、内存预测(memory/)、缓存(cache.py)、平台检测(platform.py)、环境管理(env.py)等。tokenizers/ 提供 CLIP、T5、Qwen2、Wan 等分词器实现。processor/ 提供 Canny、Depth 等图像预处理。

接入层调度层执行层Kernel 层工具层其他examples → pipelines:调用调用tools → pipelines:封装tools → processor:预处理预处理pipelines → configs:配置配置pipelines → models:加载加载pipelines → algorithm:调度调度pipelines → tokenizers:分词分词models → models/basic:组件组件models/basic → kernels:调用调用models/basic → utils:工具工具configs → utils:检测检测algorithm → models:采样采样utils → models:加载加载utils → pipelines:offloadoffloadtokenizers → models:编码pipelinespipelinestoolstoolsexamplesexamplesconfigsconfigsalgorithmalgorithmtokenizerstokenizersmodelsmodelsprocessorprocessorkernelskernelsutilsutilsmodels/basicmodels/basicteststests
DiffSynth-Engine 五层架构图:接入层(pipelines/tools/examples)→ 调度层(configs/algorithm)→ 执行层(models/tokenizers/processor)→ Kernel 层(kernels/models/basic)→ 工具层(utils)
模块 / 路径职责 · 入口 · 依赖
pipelines
diffsynth_engine/pipelines
约 15 个文件
推理管线入口层,提供面向用户的统一推理接口,管理模型加载、LoRA 绑定、offload 调度、推理执行的全生命周期
入口:BasePipeline, FluxImagePipeline, WanVideoPipeline, QwenImagePipeline, WanSpeech2VideoPipeline, SDImagePipeline, SDXLImagePipeline, Flux2KleinPipeline, ZImagePipeline, ZImageOmniBasePipeline, Hunyuan3DShapePipeline
依赖:configs, models, algorithm, utils, tokenizers
commit 热点第 3 位(6 次),是主线工作。BasePipeline 定义公共接口,各管线继承实现特定逻辑。from_pretrained 是主要工厂方法。
configs
diffsynth_engine/configs
约 5 个文件
配置管理层,通过 dataclass 定义所有管线的配置结构,包括模型路径、数据类型、设备、offload 策略、并行策略、注意力实现等
入口:BaseConfig, FluxPipelineConfig, WanPipelineConfig, QwenImagePipelineConfig, WanSpeech2VideoPipelineConfig, SDPipelineConfig, SDXLPipelineConfig, Flux2KleinPipelineConfig, ZImagePipelineConfig, HunyuanPipelineConfig, AttentionConfig, OptimizationConfig, ParallelConfig, AttnImpl, SpargeAttentionParams, VideoSparseAttentionParams, LoraConfig, ControlNetParams, QwenImageControlNetParams
依赖:utils/flag.py(间接)
commit 热点第 6 位(3 次)。AttnImpl 枚举定义了所有支持的注意力后端。ParallelConfig 包含 ulysses、ring、tp、fsdp、cfg_parallel 等并行参数。
algorithm
diffsynth_engine/algorithm
约 25 个文件
采样算法层,实现噪声调度器和采样器,支持 Flow Match(flux 系列)和 Stable Diffusion(SD 系列)两套体系
入口:base_scheduler.py, flow_beta.py, flow_ddim.py, recifited_flow.py, beta.py, ddim.py, karras.py, linear.py, exponential.py, sgm_uniform.py, flow_match_euler.py, euler.py, dpmpp_2m.py, dpmpp_2m_sde.py, dpmpp_3m_sde.py, deis.py, brownian_tree.py, ddpm.py, epsilon.py, euler_ancestral.py
依赖:torch
README 明确指出「carefully re-implemented key components such as sampler and scheduler, without introducing external dependencies on libraries like k-diffusion, ldm, or sgm」。这是项目的核心竞争力之一。
models
diffsynth_engine/models
约 60+ 个文件
模型实现层,包含所有 DiT、VAE、TextEncoder、ControlNet 等模型组件的实现
入口:PreTrainedModel, StateDictConverter, FluxDiT, FluxVAE, FluxTextEncoder, FluxControlNet, FluxIPAdapter, FluxRedux, WanDiT, WanS2VDit, WanVAE, WanTextEncoder, WanImageEncoder, WanAudioEncoder, QwenImageDiT, QwenImageVAE, SDUNet, SDVAE, SDTextEncoder, SDControlNet, SDXLUNet, SDXLVAE, SDXLTextEncoder, SDXLControlNet, SD3DiT, SD3VAE, SD3TextEncoder, Flux2DiT, Flux2VAE, ZImageDiT, ZImageDiTOmniBase, Hunyuan3DDiT, Hunyuan3DVAE, CLIPEncoderLayer, T5 编码器, Attention, LoRA
依赖:models/basic, utils, configs
commit 热点第 1 位(10 次),是代码变更最频繁的区域。每个模型家族一个子目录(flux/, wan/, qwen_image/, sd/, sdxl/, sd3/, flux2/, z_image/, hunyuan3d/)。
models/basic
diffsynth_engine/models/basic
约 10 个文件
基础组件层,提供所有模型共享的基础模块,包括注意力、LoRA、时间步嵌入、位置编码、Transformer 辅助函数等
入口:Attention, LoRA, Timestep, RelativePositionEmb, TransformerHelper, UNetHelper, VideoSparseAttention
依赖:utils/flag.py, utils/platform.py
attention.py 是核心,实现了统一的 Attention 接口,支持 FA2/FA3/FA4、SDPA、xFormers、Sage、Sparge、Aiter、VideoSparseAttention 等多种后端。
kernels
diffsynth_engine/kernels
约 1-3 个文件
自定义 kernel 层,预留用于 CUDA/Triton kernel 的扩展
入口:需验证(目录存在但文件极少)
依赖:torch
当前目录内容极少,可能是预留扩展位置或 kernel 逻辑分散在 models/basic/attention.py 中。
utils
diffsynth_engine/utils
约 20 个文件
通用工具层,提供模型加载、量化、offload、并行、内存预测、缓存、平台检测、环境管理等基础功能
入口:loader.py, fp8_linear.py, gguf.py, offload.py, parallel.py, process_group.py, cache.py, memory/, platform.py, env.py, flag.py, image.py, video.py, prompt.py, logging.py, lock.py, onnx.py, download.py
依赖:torch, modelscope
commit 热点第 5 位(4 次)。flag.py 动态检测各注意力后端可用性。memory/ 包含线性回归和预测模型,用于显存估算。
tokenizers
diffsynth_engine/tokenizers
约 10 个文件
分词器层,提供 CLIP、T5、Qwen2、Wan 等模型的分词器实现
入口:CLIPTokenizer, T5Tokenizer, Qwen2Tokenizer, WanTokenizer, Qwen2VLProcessor, Qwen2VLImageProcessor
依赖:torch, tokenizers 库
每个分词器继承自 tokenizers/base.py 中的基类。
processor
diffsynth_engine/processor
约 3 个文件
图像预处理层,提供 ControlNet 等所需的图像预处理功能
入口:CannyProcessor, DepthProcessor
依赖:torch, opencv
规模较小,主要服务于 ControlNet 工作流。
tools
diffsynth_engine/tools
约 6 个文件
高级工具层,封装复杂推理工作流,如 inpainting、outpainting、IP-Adapter、Redux、ControlNet 替换、超分等
入口:FluxInpaintingTool, FluxOutpaintingTool, FluxIPAdapterRefTool, FluxReduxRefTool, FluxReplaceByControlTool, QwenImageUpscalerTool
依赖:pipelines, models, processor
commit 热点第 8 位(2 次)。QwenImageUpscalerTool 对应 ODTSR 超分模型,是 v0.7.0 新增功能。
examples
examples
约 30 个文件
示例代码层,提供各模型的使用示例,包括基础推理、LoRA、并行推理、性能基准测试等
入口:flux_text_to_image.py, flux_lora.py, flux_parallel.py, wan_text_to_video.py, wan_image_to_video.py, wan_lora.py, wan_speech_to_video.py, wan_dmd_text_to_video.py, wan_flf_to_video.py, qwen_image_edit.py, qwen_image_eligen.py, qwen_image_parallel.py, sdxl_text_to_image.py, flux2_klein_image.py, odtsr_upscale_tool_example.py, model_perf_benchmark.py
依赖:pipelines, configs, utils
examples 是理解管线使用方式的最佳入口。flux_parallel.py、qwen_image_parallel.py、wan 系列展示了并行推理用法。
tests
tests
约 156 个文件
测试层,提供模型、管线、算法、分词器、工具的单元测试
入口:test_flux_dit.py, test_flux_text_encoder.py, test_flux_vae.py, test_wan_vae.py, test_qwen_image_dit.py, test_sd_unet.py, test_sdxl_unet.py, test_z_image_dit.py, test_flux_image.py, test_wan_video.py, test_qwen_image.py, test_sd_image.py, test_sdxl_image.py, test_sampler.py, test_scheduler.py, test_clip.py, test_t5.py, test_qwen2.py, test_flux_tools.py
依赖:所有其他模块
commit 热点第 2 位(tests/data 10 次),测试覆盖较全面。test_pipelines 是管线测试的核心。
目录树(按文件数)
  • diffsynth_engine/ 211 个文件
    __init__.py, algorithm, conf, configs, kernels, models, pipelines, processor, tokenizers, tools, utils
  • tests/ 156 个文件
    __init__.py, common, data, test_algorithm, test_models, test_pipelines, test_tokenizers, test_tools
  • examples/ 30 个文件
    flux_lora.py, flux_parallel.py, flux_text_to_image.py, input, memory_predict_model, model_benchmark_readme.md, model_perf_benchmark.py, odtsr_upscale_tool_example.py, qwen_image_edit.py, qwen_image_edit_2511.py, qwen_image_eligen.py, qwen_image_parallel.py, sdxl_text_to_image.py, wan_dmd_image_to_video.py
  • assets/ 3 个文件
    dingtalk.png, showcase.jpeg, tongyi.svg
  • docs/ 2 个文件
    tutorial.md, tutorial_zh.md
  • .gitattributes/ 1 个文件
  • .github/ 1 个文件
    workflows
  • .gitignore/ 1 个文件
  • .pre-commit-config.yaml/ 1 个文件
  • LICENSE/ 1 个文件
  • MANIFEST.in/ 1 个文件
  • README.md/ 1 个文件
  • pyproject.toml/ 1 个文件
  • setup.py/ 1 个文件

一次调用怎么流过这些模块

一次典型的推理请求(以 Flux 文生图为例)的数据流如下: 1. 用户调用入口:用户通过 FluxImagePipeline.__call__() 传入 prompt 参数,触发推理流程。该方法定义在 diffsynth_engine/pipelines/flux_image.py 中,继承自 BasePipeline。 2. 配置解析与初始化:管线通过 FluxPipelineConfig(定义在 diffsynth_engine/configs/pipeline.py)获取模型路径、数据类型、设备、offload 策略、并行策略、注意力实现等配置。from_pretrained 工厂方法(v0.4.0 改进过)负责根据配置加载所有模型组件。 3. 文本编码:prompt 被送入 tokenizers/clip.py 或 tokenizers/t5.py 进行分词,生成 token ids。然后传入 models/flux/flux_text_encoder.py 或 models/text_encoder/ 中的文本编码器,生成 text embeddings(Tensor)。 4. 噪声初始化与调度:根据配置选择噪声调度器(如 algorithm/noise_scheduler/flow_match/flow_beta.py),初始化 latent noise(Tensor)。调度器定义了噪声水平随时间步的变化规律。 5. 采样循环(核心):进入采样循环,每个时间步: a. 将 latent、text embeddings、timestep 送入 DiT 模型(models/flux/flux_dit.py)。 b. DiT 内部通过 models/basic/attention.py 中的 Attention 层执行自注意力和交叉注意力,Attention 根据 AttnImpl 配置选择 FA2/FA3/FA4/SDPA/xFormers/Sage/Sparge/Aiter 等后端。 c. 如果启用 ControlNet,通过 models/flux/flux_controlnet.py 注入额外条件。 d. 如果启用 LoRA,通过 models/basic/lora.py 进行权重融合。 e. 采样器(如 algorithm/sampler/flow_match/flow_match_euler.py)根据 DiT 输出更新 latent。 6. Offload 调度:在显存受限场景下,utils/offload.py 根据 offload_mode 配置,在时间步之间将模型组件在 CPU/GPU 之间移动,或启用磁盘卸载(offload_to_disk)。 7. 并行推理:如果配置了并行策略(ParallelConfig),通过 utils/parallel.py 和 utils/process_group.py 实现序列并行(Ulysses、Ring)、张量并行(TP)、FSDP、CFG 并行等。 8. VAE 解码:采样完成后,latent 被送入 models/flux/flux_vae.py 或 models/vae/vae.py 进行解码,生成图像 Tensor。支持 VAE tiling(vae_tiled、vae_tile_size、vae_tile_stride)以降低显存占用。 9. 后处理与输出:通过 utils/image.py 将 Tensor 转换为 PIL Image,返回给用户。 关键数据结构:FluxPipelineConfig(配置)、text_embeddings(Tensor)、latent(Tensor)、timestep(Tensor)、ControlNetParams(控制参数)、LoRAConfig(LoRA 参数)。调度点:BasePipeline.__call__(入口)、噪声调度器(时间步规划)、采样器(latent 更新)、offload 调度器(显存管理)、并行策略(分布式计算)。

1用户调用FluxImagePipeline.__call__ · diffsynth_engine/pipelines/flux_image.py2配置解析FluxPipelineConfig · diffsynth_engine/configs/pipeline.py3文本编码CLIPEncoderLayer / T5 · diffsynth_engine/models/text_encoder/4噪声初始化flow_beta scheduler · diffsynth_engine/algorithm/noise_scheduler/flow_match/flow_b5DiT 推理FluxDiT + Attention · diffsynth_engine/models/flux/flux_dit.py6采样更新flow_match_euler sampler · diffsynth_engine/algorithm/sampler/flow_match/flow_match_eu7Offload 调度offload · diffsynth_engine/utils/offload.py8VAE 解码FluxVAE · diffsynth_engine/models/flux/flux_vae.py9图像输出utils/image.py · diffsynth_engine/utils/image.py

关键类型与函数

名称路径用途
BasePipelinediffsynth_engine/pipelines/base.py所有推理管线的基类,定义了模型加载、LoRA 绑定、offload 调度、__call__ 入口等公共生命周期
FluxImagePipelinediffsynth_engine/pipelines/flux_image.pyFlux 文生图管线,继承 BasePipeline,实现 Flux 模型的加载与推理逻辑
WanVideoPipelinediffsynth_engine/pipelines/wan_video.pyWan 视频生成管线,支持文生视频、图生视频等多种模式
QwenImagePipelinediffsynth_engine/pipelines/qwen_image.pyQwen-Image 文生图管线,支持复杂文本渲染
PreTrainedModeldiffsynth_engine/models/base.py所有模型的基类,定义了 StateDictConverter 接口和模型加载规范
StateDictConverterdiffsynth_engine/models/base.py状态字典转换器,用于将不同格式(safetensors、gguf 等)的权重转换为模型可用格式
Attentiondiffsynth_engine/models/basic/attention.py统一的注意力层实现,支持 FA2/FA3/FA4、SDPA、xFormers、Sage、Sparge、Aiter、VideoSparseAttention 等多种后端
AttnImpldiffsynth_engine/configs/pipeline.py注意力后端枚举,定义了 AUTO、EAGER、FA2、FA3、FA3_FP8、FA4、AITER、AITER_FP8、XFORMERS、SDPA、SAGE、SPARGE、VSA 等选项
FluxPipelineConfigdiffsynth_engine/configs/pipeline.pyFlux 管线配置,包含模型路径、数据类型、设备、offload 策略、并行策略等
ParallelConfigdiffsynth_engine/configs/pipeline.py并行策略配置,包含 parallelism、cfg_parallel、sp_ulysses_degree、sp_ring_degree、tp_degree、use_fsdp 等参数
OptimizationConfigdiffsynth_engine/configs/pipeline.py优化配置,包含 use_fp8_linear、use_fbcache、fbcache_relative_l1_threshold、use_torch_compile 等
LoRAdiffsynth_engine/models/basic/lora.pyLoRA 权重融合模块,支持加载和绑定 LoRA 权重
VideoSparseAttentiondiffsynth_engine/models/basic/video_sparse_attention.py视频稀疏注意力实现,支持分布式场景,通过 vsa 库或自定义实现
ControlNetParamsdiffsynth_engine/configs/controlnet.pyControlNet 参数封装,包含 image、scale、model、mask、control_start、control_end 等
FluxDiTdiffsynth_engine/models/flux/flux_dit.pyFlux DiT 模型实现,是 Flux 推理的核心组件
WanDiTdiffsynth_engine/models/wan/wan_dit.pyWan DiT 模型实现,支持视频生成
QwenImageDiTdiffsynth_engine/models/qwen_image/qwen_image_dit.pyQwen-Image DiT 模型实现
fp8_lineardiffsynth_engine/utils/fp8_linear.pyFP8 线性层实现,用于模型量化加速
offloaddiffsynth_engine/utils/offload.py模型 offload 调度器,支持 CPU offload、磁盘卸载等策略
flagdiffsynth_engine/utils/flag.py运行时标志检测,动态判断各注意力后端(FA2/FA3/FA4/xFormers/Sage/Sparge/Aiter/VSA)是否可用

扩展点

最近在动的地方

建议阅读顺序

  1. 1. README.md - 了解项目定位、核心特性、支持的模型列表
  2. 2. diffsynth_engine/pipelines/base.py - 理解 BasePipeline 的公共接口和生命周期
  3. 3. diffsynth_engine/configs/pipeline.py - 理解配置结构和各模型家族的 PipelineConfig
  4. 4. diffsynth_engine/models/base.py - 理解 PreTrainedModel 和 StateDictConverter 基类
  5. 5. diffsynth_engine/models/basic/attention.py - 理解统一的 Attention 接口和多后端调度
  6. 6. diffsynth_engine/pipelines/flux_image.py - 以 Flux 为例理解具体管线的实现
  7. 7. diffsynth_engine/models/flux/flux_dit.py - 理解 DiT 模型的实现方式
  8. 8. diffsynth_engine/algorithm/noise_scheduler/flow_match/flow_beta.py - 理解 Flow Match 调度器
  9. 9. diffsynth_engine/algorithm/sampler/flow_match/flow_match_euler.py - 理解采样器实现
  10. 10. diffsynth_engine/utils/offload.py - 理解 offload 策略和显存管理
  11. 11. diffsynth_engine/utils/parallel.py - 理解并行推理实现
  12. 12. examples/flux_text_to_image.py - 通过示例理解使用方式

三、本地跑起来(没有 GPU 的 Mac)

安装

  1. # 1. 克隆仓库并创建 Python 3.10+ 虚拟环境(Apple Silicon 推荐使用 Miniforge 或 venv)
  2. git clone https://github.com/modelscope/diffsynth-engine.git && cd diffsynth-engine
  3. python3 -m venv .venv && source .venv/bin/activate
  4. # 2. 安装 PyTorch(Apple Silicon 支持 MPS 后端,pip 安装即可)
  5. pip install torch torchvision --index-url https://download.pytorch.org/whl/cpu # 或默认安装(含 MPS)
  6. # 3. 安装项目核心依赖(跳过 yunchang,因其仅限 Linux)
  7. pip install -e . # 自动安装 pyproject.toml 中的依赖,yunchang 因 sys_platform 限制会被跳过
  8. # 4. 安装开发依赖(可选,用于测试和 lint)
  9. pip install -e '.[dev]'
  10. # 5. 安装 pre-commit hooks(可选)
  11. pre-commit install
  12. # 注意:flash_attn、xformers、sageattention、aiter 等 CUDA 库在 macOS 上无法安装,需跳过或忽略
  13. # 注意:onnxruntime、opencv-python、moviepy、librosa 等可正常安装

哪些路径能真跑

['# 可在 CPU / MPS 上执行的路径:', '# - diffsynth_engine/algorithm/(scheduler 和 sampler 纯 Python 实现,无 CUDA 依赖)', '# - diffsynth_engine/configs/(配置 dataclass,纯 Python)', '# - diffsynth_engine/tokenizers/(基于 HuggingFace tokenizers,CPU 可用)', '# - diffsynth_engine/models/text_encoder/clip.py、t5.py(模型前向可用 CPU,但慢)', '# - diffsynth_engine/models/basic/attention.py 中的 eager_attn 和 sdpa_attn(PyTorch 原生)', '# - diffsynth_engine/utils/(大部分工具函数,如 image.py、video.py、loader.py)', '# - tests/test_algorithm/、tests/test_tokenizers/(纯 CPU 测试)', '# 需要 CUDA / GPU 的路径(只能读代码或 Colab T4 验证):', '# - diffsynth_engine/kernels/(自定义 CUDA kernel)', '# - diffsynth_engine/models/flux/、wan/、qwen_image/ 等 DiT 模型(依赖 flash_attn、CUDA)', '# - diffsynth_engine/models/basic/video_sparse_attention.py(依赖 vsa CUDA 库)', '# - diffsynth_engine/utils/fp8_linear.py(FP8 量化需 Hopper+ GPU)', '# - diffsynth_engine/utils/parallel.py(分布式并行需多 GPU)', '# - tests/test_pipelines/、tests/test_models/ 中的大部分测试(需模型权重和 GPU)']

最小可运行

  1. # 1. 测试 scheduler 和 sampler(纯 CPU,无需模型权重)
  2. python -c "from diffsynth_engine.algorithm.noise_scheduler.flow_match import flow_euler; print('scheduler ok')"
  3. # 2. 测试 tokenizer(CPU 可用)
  4. python -c "from diffsynth_engine.tokenizers import CLIPTokenizer; print('tokenizer ok')"
  5. # 3. 测试配置加载(纯 Python)
  6. python -c "from diffsynth_engine.configs import FluxPipelineConfig; cfg = FluxPipelineConfig.basic_config(model_path='test', device='cpu'); print(cfg)"
  7. # 4. 测试模型初始化(小模型,CPU 慢但可运行)
  8. python -c "from diffsynth_engine.models.text_encoder.clip import CLIPEncoderLayer; layer = CLIPEncoderLayer(embed_dim=64, intermediate_size=128, device='cpu', dtype=torch.float32); print(layer)"
  9. # 5. 测试工具函数
  10. python -c "from diffsynth_engine.utils.image import load_image; print('utils ok')"
  11. # 注意:完整推理 pipeline(如 FluxImagePipeline)需要模型权重和 GPU,建议在 Colab T4 上验证

测试

['# 测试框架:pytest', '# 测试目录:tests/,包含 test_algorithm/、test_models/、test_pipelines/、test_tokenizers/、test_tools/', '# 仅跑 CPU 子集(跳过 GPU 测试):', 'pytest tests/test_algorithm/ -v # scheduler 和 sampler 测试,纯 CPU', 'pytest tests/test_tokenizers/ -v # tokenizer 测试,纯 CPU', '# 跳过 GPU 测试(需模型权重和 CUDA):', '# pytest tests/test_models/ -v # 需要 GPU', '# pytest tests/test_pipelines/ -v # 需要 GPU 和模型权重', '# 预计耗时:CPU 子集约 1-5 分钟(取决于模型加载)']

调试

CI

['# CI 框架:GitHub Actions(.github/workflows/)', '# 测试内容:', '# - 代码风格检查(ruff)', '# - 单元测试(pytest,可能包含 GPU 测试)', '# - 构建验证(pip install)', '# PR 可能被卡住的情况:', '# - 代码风格不符合 ruff 配置(line-length=120)', '# - 缺少测试或测试失败', '# - 依赖冲突(如 torch 版本不在 2.6-2.8 范围)', '# - 文档或类型注解问题']

坑

四、维护者与社区

高频迭代。从 releases 看,v0.4.0(2025-08-01)→ v0.4.1(2025-08-04)→ v0.5.0(2025-08-27)→ v0.6.0(2025-09-09)→ v0.7.0(2026-01-19),平均约 1–2 个月一个 minor release,patch 版本穿插其间。commit 证据显示提交热点集中在 diffsynth_engine/models(10 次)、tests/data(10 次)、diffsynth_engine/pipelines(6 次)、tests/test_pipelines(5 次),说明主线工作是“接入新模型 + 配套测试”。最近一次有记录的 release 是 v0.7.0(2026-01-19),其后仍有 commit 活动(如 2026-05-27 的 docs fix、2026-03-27 的 fa2/fa4 冲突修复),表明维护在继续但 release 节奏略有放缓。

谁角色依据
MuseAI x ModelScope项目所有者与核心维护者pyproject.toml 中 authors = [{ name = "MuseAI x ModelScope" }];README 引用作者为 Zhipeng Di, Guoxuan Zhu, Zhongjie Duan, Zihao Chu, Yingda Chen, Weiyi Lu;仓库位于 modelscope/ 组织下。
社区贡献者(NPU/Ascend 方向)华为昇腾 NPU 适配的主要推动者PR #264、#265、#270、#272 均为 feat(npu) 或 Refactor/npu,涉及 Ascend platform abstraction、Ulysses SP、long-context attention、fused DiT ops、MindIE-SD compile fusion 等,说明有专门的贡献者在做 NPU 后端。
社区贡献者(模型接入与 bugfix)模型支持与问题修复PR #250(fix wan_s2v_dit attn_kwargs,fixes #221)、#257(fix qwen image tile mask weights)、#237(add reference image KV cache for QwenImageDiT);commit 中可见“support flux.2-klein (#227)”、“support Z-Image-Omni-Base (#226)”、“add qwen image upscaler tool(ODTSR) (#229)”等,说明有多位贡献者负责新模型接入与 bug 修复。
社区贡献者(文档与体验)文档与安装体验维护commit 中有“docs: fix typos in README ('buidling', 'varous')”、“docs: address review — add 'methods', drop trailing whitespace”、“fix: constrain torch version (#232)”、“fix api argument error (#231)”,表明有人负责文档与依赖约束维护。

流程与 Review 风格

贡献流程相对轻量,但文档尚不完善。README Contributing 章节仅写:“We welcome contributions to DiffSynth-Engine. After Install from source, we recommand developers install this project using following command to setup the development environment. pip install -e '.[dev]' pre-commit install”,并注明“TODO: Please refer to [CONTRIBUTING.md](./CONTRIBUTING.md) for more details”——但仓库中实际不存在 CONTRIBUTING.md(root_paths 中未列出),说明贡献指南尚未落地。pyproject.toml 配置了 ruff(line-length=120, target-version=py310, select E4/E7/E9/F)和 pre-commit(.pre-commit-config.yaml 存在),因此代码风格检查是自动化的。未发现 CLA 或 DCO 要求(README、CONTRIBUTING 均未提及)。从 PR 模式看,贡献者直接提 PR,维护者 review 后合入,没有明显的 RFC 或 Issue 前置流程要求。

务实、偏快速迭代。从 PR 标题和内容看,review 关注点集中在:1) 正确性(如 #250 直接 fixes #221,#257 修正 tile mask weights);2) 性能(如 #267 “change sdpa backend to flash attention”、#237 “add reference image KV cache”);3) 平台兼容性(如 #232 “constrain torch version”、#272 “npu transformer cleanup”)。Issue 响应方面,部分 issue 有互动(#199 PyTorch version problem 有 6 条评论,#224 MAC Ultra load_lora 报错有 2 条评论),但也有 issue 长期无回复(#266 performance regression on Ring SP 0 条评论,#235 batch inference support 0 条评论),说明维护者对高优先级 bug 会响应,但对功能请求或边缘问题响应较慢。docs 相关的 review 较细致(commit “docs: address review — add 'methods', drop trailing whitespace”),表明对对外文档质量有要求。

渠道

这里的规矩

维护者现在最想要的帮助

五、切入方案

建议长期负责:diffsynth_engine/algorithm/ + tests/test_algorithm/ + 跨平台(CPU/MPS/CUDA)attention 后端兼容性层
algorithm/ 是项目“自研 sampler/scheduler”的核心卖点,README 明确强调无外部依赖;该层纯 Python 实现,完全无 GPU 起步门槛;同时它向上承接 pipelines、向下依赖 attention,是天然的枢纽模块。长期负责此模块可让你从“写测试”过渡到“主导采样算法演进”,是通向 reviewer/维护者身份的短路径。

它现在缺什么(你无 GPU 也能补)

缺口依据为什么是你
缺少 CPU/MPS 可运行的端到端正确性测试套件tests/test_pipelines/ 下所有测试文件(如 test_flux_image.py、test_wan_video.py)均依赖 CUDA 设备,tests/common/test_case.py 中测试基类默认 device='cuda';runbook 中列出的 no_gpu_paths 仅覆盖 algorithm/、configs/、tokenizers/ 等局部模块,缺少一条从 pipelines 到 models 的 CPU/MPS 回归路径。你擅长算子融合与内核正确性验证,可在无 GPU 条件下用 torch.device('cpu') 和 MPS 后端为 pipelines 层编写轻量级 shape/dtype/梯度测试,直接填补项目“跨平台但无跨平台 CI”的缺口。
缺少 scheduler/sampler 与 diffusers 的数值对照基准README 明确声明“carefully re-implemented key components such as sampler and scheduler, without introducing external dependencies on libraries like k-diffusion, ldm, or sgm”,但 tests/test_algorithm/ 仅有 test_scheduler.py 与 test_sampler.py 的基本结构,未发现与 diffusers 或原版 k-diffusion 的数值对齐测试;pyproject.toml 中 dev 依赖包含 diffusers==0.31.0,说明项目具备对照条件但未启用。你熟悉扩散模型采样算法,可在 CPU 上用 diffusers==0.31.0 作为参考实现,为 flow_match、stable_diffusion 两套 scheduler/sampler 建立数值一致性测试,这是纯 Python 工作,完全无 GPU 依赖。
缺少 Apple Silicon 原生性能基准与回归脚本README 将 Apple Silicon 列为 requirement,examples/model_perf_benchmark.py 存在但默认面向 CUDA;commit 中无 MPS 相关性能追踪记录;utils/memory/ 下有 memory_predict_model.py 但仅做显存预测,无 MPS 耗时基线。你手边就是 MacBook Air M 系列,可基于 examples/model_perf_benchmark.py 扩展 MPS 路径,建立 attention、vae、dit 单层在 MPS 下的耗时基线,为后续优化提供可量化锚点。
缺少 attention 后端在 CPU/MPS 下的可用性矩阵文档diffsynth_engine/models/basic/attention.py 中实现了 eager_attn、sdpa_attn、flash_attn2/3/4、xformers、sage_attn、sparge_attn、aiter、video_sparse_attn 等多后端,但 README 与 docs/ 未说明哪些后端在 CPU、MPS、CUDA 上分别可用;utils/flag.py 仅做布尔标记,无设备-后端兼容性表。你可编写一个脚本遍历所有 AttnImpl 枚举值,在 cpu/mps/cuda 上分别执行一次前向,生成兼容性矩阵文档或测试,这是低风险、高可见度的贡献。
缺少新模型接入的标准化 checklist 与模板commit 热点显示 diffsynth_engine/models 与 pipelines 是主线工作,半年内新增 Flux2、Z-Image、Qwen-Image-Edit、Wan2.2-S2V 等;但仓库根目录无 CONTRIBUTING.md(README 引用但文件缺失),无新模型接入模板,PR 流程依赖维护者口头指导。你可基于现有 Flux/Wan/Qwen-Image 的接入模式,提炼一份“新模型接入 checklist”(StateDictConverter、Config、Pipeline、测试、example),以文档 PR 形式提交,降低社区贡献门槛。
缺少 GGUF 量化路径的测试覆盖tests/test_pipelines/test_wan_video_gguf.py 存在,utils/gguf.py 提供 GGUF 加载,但仅 wan_video 一条路径有测试;models/ 下 flux、qwen_image、z_image 等均未发现对应 GGUF 测试;commit 中无 GGUF 扩展记录。你可为 flux 或 qwen_image 的 GGUF 路径补充测试,或在 CPU 上验证 GGUF 反量化与原始权重的数值一致性,这是纯 Python 工作。
缺少 LoRA 加载的跨模型单元测试models/basic/lora.py 实现 LoRA 基类,pipelines/base.py 提供 load_lora,但 tests/ 下无 test_lora.py;tests/test_pipelines 中仅在具体管线测试中隐式调用 LoRA,缺少针对 LoRA 合并/卸载/缩放的独立测试。你可为 LoRA 的 merge/unload/scale 逻辑编写 CPU 上的单元测试,覆盖线性层与卷积层,这是低风险、高价值的正确性保障。
缺少 offload 策略在 CPU 下的行为验证utils/offload.py 实现模型卸载逻辑,configs/pipeline.py 中 offload_mode/offload_to_disk 参数存在,但 tests/ 下无 test_offload.py;runbook 中 no_gpu_paths 未提及 offload 模块。

第 1–30 天:看懂并露面

第 31–60 天:稳定产出

第 61–90 天:接管一块

第一批 PR

题目范围为什么安全
test: add CPU regression tests for flow_match and stable_diffusion schedulerstests/test_algorithm/test_scheduler.py纯测试新增,不修改生产代码;使用 diffusers==0.31.0(已在 dev 依赖中)作为参考实现;CPU 运行无 GPU 依赖
test: add DPM-Solver++ and DEIS sampler unit tests on CPUtests/test_algorithm/test_sampler.py仅添加测试用例,验证采样器单步/多步的数值稳定性;不涉及模型权重加载,仅需随机张量
test: add attention backend compatibility matrix scripttests/test_models/ 或新文件 tests/test_attention_backends.py脚本化测试,遍历 AttnImpl 枚举并记录可用性;失败时标记为 skip 而非 break,不影响现有 CI
docs: add MPS/CPU performance benchmark example for attention layersexamples/ 新增 mps_attention_benchmark.py新增示例文件,不修改现有代码;基于 model_perf_benchmark.py 扩展 MPS 路径
test: add LoRA merge/unload/scale unit teststests/test_models/test_lora.py(新建)针对 models/basic/lora.py 的纯 Python 单元测试;使用随机初始化的线性层,无需 GPU

怎么知道自己站住了

风险与对策
  • algorithm/ 模块的数学实现可能与 diffusers 存在刻意差异(如数值精度截断),导致数值一致性测试难以对齐:在测试中设置合理的 atol/rtol 容差,并在测试文档中注明“对齐目标”而非“完全相等”;与维护者确认差异是否有意为之
  • MPS 后端对某些算子(如某些 attention 变体)支持不完善,导致测试在 MPS 上失败而非代码 bug:在测试中区分 device 能力,对 MPS 不支持的算子标记为 skip 或 expected failure;在兼容性矩阵中明确记录
  • 维护者对测试 PR 的优先级不高,导致 review 周期长:先提交小批量、高价值的测试(如单个 scheduler),快速建立信任;同时在 issue 中明确说明测试覆盖的缺口与用户价值
  • 项目处于快速迭代期(v0.4→v0.7 半年),你的测试可能因上游重构而频繁 break:将测试聚焦于稳定接口(scheduler/sampler 的公开方法),避免依赖内部实现细节;在 CI 中设置合理的维护策略
  • Apple Silicon 16GB 内存限制可能导致某些测试(如大模型加载)OOM:测试中使用极小模型(如 1x1 输入、单层网络)或 mock 权重;避免在 CPU/MPS 上加载完整 Flux/Wan 模型
  • 仓库缺少 CONTRIBUTING.md,导致你的文档类 PR 缺乏明确的合并路径:
  • attention 后端涉及多个第三方库(flash_attn、xformers、sage、aiter),在 macOS 上无法安装,导致兼容性矩阵不完整:在脚本中捕获 ImportError 并标记为 unavailable,而非测试失败;文档中明确列出各后端的平台限制
  • 你的长期目标模块(algorithm/)可能由核心维护者直接掌控,社区贡献空间有限:通过高质量测试建立信任后,主动申请成为该模块的 reviewer;同时关注 pipelines/ 与 models/ 中依赖 algorithm 的扩展点

六、怎么介入这个项目

建议顺序

成为长期维护者的路径
  • diffsynth_engine/algorithm/ + tests/test_algorithm/
  • diffsynth_engine/models/basic/attention.py 的 CPU/MPS 后端路径
  • tests/test_models/ 新增测试文件

七、任务卡

任务 1优先 · medium · CPU · 2–3 个晚上

test: 为 flow_match/stable_diffusion scheduler 添加 CPU 回归测试

无人认领6 条评论更新 2026-02-03

用到的专长:扩散模型采样算法与数值正确性验证

目标:在 tests/test_algorithm/test_scheduler.py 中新增 CPU 测试,验证 DiffSynth 的 flow_match、beta、exponential 等 scheduler 与 diffusers==0.31.0(已在 dev 依赖中)的数值一致性。

为什么值得长期做:README 明确声明 scheduler/sampler 是核心竞争力(carefully re-implemented without external dependencies),但 tests/test_algorithm/ 缺乏与参考实现的数值对照;这是维护者最可能长期依赖的测试防线。

怎么介入:Issue #199 讨论的是 torch 版本兼容性,但维护者在评论中关注测试覆盖;本任务从测试角度切入,不与现有 open PR 冲突。
第一个 PR 的边界:仅修改 tests/test_algorithm/test_scheduler.py,新增 TestSchedulerCPU 类;不改动生产代码。
第一步:阅读 tests/test_algorithm/test_scheduler.py 现有结构,确认测试基类与 device 默认值;阅读 diffsynth_engine/algorithm/ 下 flow_match_euler.py、beta.py、exponential.py 的接口签名。
本机怎么复现 / 验证:pip install -e '.[dev]' && pytest tests/test_algorithm/test_scheduler.py::TestSchedulerCPU -v --device=cpu
认领留言(英文,可直接贴到 Issue)
I'd like to add CPU regression tests for the flow_match and stable_diffusion schedulers in tests/test_algorithm/test_scheduler.py, using diffusers==0.31.0 (already in dev deps) as the reference implementation. The tests will verify numerical consistency of sigma/timestep values on CPU, with no GPU dependency. May I proceed? I'll have a PR ready in 2-3 days.
大致实施方案
  • 在 tests/test_algorithm/test_scheduler.py 中新增 TestSchedulerCPU 类,device='cpu'
  • 使用 diffusers.schedulers 作为参考实现,固定 timesteps 与 generator
  • 对比 DiffSynth scheduler 与 diffusers scheduler 在相同输入下的 sigma / timestep 数值,容差 1e-6
  • 覆盖 flow_match、beta、exponential、sgm_uniform 等核心 scheduler
  • 在 Mac 上运行 pytest tests/test_algorithm/test_scheduler.py -v 验证
可能涉及的目录或文件
  • tests/test_algorithm/test_scheduler.py(修改)
  • diffsynth_engine/algorithm/(只读)
验收方式
  • pytest tests/test_algorithm/test_scheduler.py -v 在 Mac CPU 上全部通过
  • 新增测试文件:tests/test_algorithm/test_scheduler.py(修改,新增 TestSchedulerCPU 类)
开工前问题与风险

向维护者确认

  • 是否接受 diffusers==0.31.0 作为参考实现?还是希望与原版 k-diffusion 对照?
  • 容差阈值 1e-6 是否合适?

风险

  • diffusers 版本差异可能导致某些 scheduler 接口不匹配,需确认 pyproject.toml 中 dev 依赖版本
任务 2优先 · medium · CPU · 2–3 个晚上

test: 为 DPM-Solver++/DEIS sampler 添加 CPU 单元测试

无人认领6 条评论更新 2026-02-03

用到的专长:扩散模型采样算法与算子正确性验证

目标:在 tests/test_algorithm/test_sampler.py 中新增 CPU 测试,验证 dpmpp_2m、dpmpp_3m_sde、deis 等 sampler 在随机噪声输入下的单步输出 shape/dtype 正确性与数值稳定性。

为什么值得长期做:sampler 是项目核心卖点之一,但 tests/test_algorithm/test_sampler.py 仅有基本结构;补充单步/多步数值稳定性测试可直接提升代码可靠性。

怎么介入:Issue #199 关注 torch 版本问题,但测试补充是正交改进;无 open PR 冲突。
第一个 PR 的边界:仅修改 tests/test_algorithm/test_sampler.py,新增 TestSamplerCPU 类;不改动 sampler 实现。
第一步:阅读 tests/test_algorithm/test_sampler.py 现有结构;阅读 diffsynth_engine/algorithm/ 下 dpmpp_2m.py、deis.py 的 step() 接口。
本机怎么复现 / 验证:pip install -e '.[dev]' && pytest tests/test_algorithm/test_sampler.py::TestSamplerCPU -v --device=cpu
认领留言(英文,可直接贴到 Issue)
I'd like to add CPU unit tests for DPM-Solver++ (dpmpp_2m, dpmpp_2m_sde, dpmpp_3m_sde) and DEIS samplers in tests/test_algorithm/test_sampler.py. The tests will verify single-step output shape, dtype, and numerical stability (no NaN/Inf) using random tensors on CPU. No GPU needed. May I proceed? PR in 2-3 days.
大致实施方案
  • 在 tests/test_algorithm/test_sampler.py 中新增 TestSamplerCPU 类
  • 构造随机 noise_pred 与 latent 张量(shape [1, 4, 8, 8]),device='cpu'
  • 调用 sampler.step() 验证输出 shape、dtype、无 NaN/Inf
  • 覆盖 dpmpp_2m、dpmpp_2m_sde、dpmpp_3m_sde、deis、brownian_tree
  • 运行 pytest tests/test_algorithm/test_sampler.py -v 验证
可能涉及的目录或文件
  • tests/test_algorithm/test_sampler.py(修改)
  • diffsynth_engine/algorithm/(只读)
验收方式
  • pytest tests/test_algorithm/test_sampler.py::TestSamplerCPU -v 在 Mac CPU 上通过
  • 新增测试文件:tests/test_algorithm/test_sampler.py(修改,新增 TestSamplerCPU 类)
开工前问题与风险

向维护者确认

  • 是否需要验证 sampler 的多步收敛性?还是仅验证单步接口正确性?

风险

  • 某些 sampler 可能依赖 scheduler 状态,需确认测试是否需要 mock scheduler
任务 3可选 · low · Mac · 1–2 个晚上

test: 添加 attention 后端 CPU/MPS/CUDA 兼容性矩阵脚本

无人认领2 条评论更新 2026-01-19

用到的专长:attention 后端与跨平台适配

目标:创建 tests/test_models/test_attention_backends.py,遍历 AttnImpl 枚举,在 cpu/mps 上执行一次前向,生成兼容性矩阵文档或测试报告。

为什么值得长期做:attention.py 支持 8+ 种后端(FA2/FA3/FA4/SDPA/xFormers/Sage/Sparge/Aiter/VSA),但 README 与 docs 未说明各后端在 CPU/MPS/CUDA 的可用性;这是跨平台用户的高频困惑点。

怎么介入:Issue #224 讨论的是 LoRA 加载错误,但根因涉及模型结构与 attention 调用;本任务从测试角度补充后端兼容性信息,不与现有 PR 冲突。
第一个 PR 的边界:新建 tests/test_models/test_attention_backends.py;不修改现有 attention 实现。
第一步:阅读 diffsynth_engine/models/basic/attention.py 中 AttnImpl 枚举与各后端实现;阅读 diffsynth_engine/utils/flag.py 了解现有可用性检测逻辑。
本机怎么复现 / 验证:pip install -e '.[dev]' && pytest tests/test_models/test_attention_backends.py -v --device=cpu && pytest tests/test_models/test_attention_backends.py -v --device=mps
认领留言(英文,可直接贴到 Issue)
I'd like to create tests/test_models/test_attention_backends.py to build a compatibility matrix for all AttnImpl backends (FA2/FA3/FA4/SDPA/xFormers/Sage/Sparge/Aiter/VSA) on CPU and MPS. The script will attempt a forward pass for each backend and report availability. This helps cross-platform users understand which backends work on their device. No GPU needed. May I proceed? PR in 1-2 days.
大致实施方案
  • 新建 tests/test_models/test_attention_backends.py
  • 遍历 AttnImpl 枚举值,构造小规模 Q/K/V 张量(shape [1, 8, 4, 64])
  • 在 device='cpu' 和 device='mps'(若可用)上尝试调用 attention
  • 记录每个后端是否可用、错误类型,输出兼容性矩阵
  • 对不可用后端标记为 pytest.skip 而非 fail
  • 运行 pytest tests/test_models/test_attention_backends.py -v 验证
可能涉及的目录或文件
  • tests/test_models/test_attention_backends.py(新建)
  • diffsynth_engine/models/basic/attention.py(只读)
  • diffsynth_engine/utils/flag.py(只读)
验收方式
  • pytest tests/test_models/test_attention_backends.py -v 在 Mac CPU 上通过
  • 新增测试文件:tests/test_models/test_attention_backends.py
开工前问题与风险

向维护者确认

  • 是否需要将兼容性矩阵输出为 Markdown 文档?还是仅作为测试报告?
  • 是否需要在 CI 中运行此脚本?

风险

  • MPS 后端在某些操作上可能 fallback 到 CPU,需区分"原生支持"与"fallback"
任务 4优先 · low · CPU · 1–2 个晚上

test: 添加 LoRA merge/unload/scale 单元测试

无人认领2 条评论更新 2026-01-19

用到的专长:LoRA 机制与算子正确性验证

目标:新建 tests/test_models/test_lora.py,为 models/basic/lora.py 的 merge/unload/scale 逻辑编写 CPU 单元测试,覆盖线性层与卷积层。

为什么值得长期做:LoRA 是项目核心功能(README 强调 extensive model support + LoRA),但 tests/ 下无 test_lora.py;Issue #224 的 LoRA 加载错误暴露了测试缺口。

怎么介入:Issue #224 的 LoRA 错误涉及模型属性缺失,但本任务聚焦 LoRA 基础逻辑测试,不与现有 PR 冲突。
第一个 PR 的边界:新建 tests/test_models/test_lora.py;不修改 LoRA 实现。
第一步:阅读 diffsynth_engine/models/basic/lora.py 中 LoRA 基类的 merge/unload/scale 接口;阅读 pipelines/base.py 中 load_lora 的调用路径。
本机怎么复现 / 验证:pip install -e '.[dev]' && pytest tests/test_models/test_lora.py -v --device=cpu
认领留言(英文,可直接贴到 Issue)
I'd like to create tests/test_models/test_lora.py to add CPU unit tests for the LoRA merge/unload/scale logic in diffsynth_engine/models/basic/lora.py. The tests will use random Linear and Conv2d layers to verify that merge changes weights, unload restores original weights, and scale correctly adjusts LoRA contribution. No GPU needed. May I proceed? PR in 1-2 days.
大致实施方案
  • 新建 tests/test_models/test_lora.py
  • 构造 torch.nn.Linear 与 torch.nn.Conv2d 层,device='cpu'
  • 创建 LoRA 实例(rank=4, alpha=1.0),执行 merge 后验证权重变化
  • 执行 unload 后验证权重恢复原始值
  • 验证 scale 参数对 LoRA 权重的影响
  • 运行 pytest tests/test_models/test_lora.py -v 验证
可能涉及的目录或文件
  • tests/test_models/test_lora.py(新建)
  • diffsynth_engine/models/basic/lora.py(只读)
验收方式
  • pytest tests/test_models/test_lora.py -v 在 Mac CPU 上通过
  • 新增测试文件:tests/test_models/test_lora.py
开工前问题与风险

向维护者确认

  • 是否需要测试 LoRA 与 DiT 模型的实际集成?还是仅测试基础 LoRA 逻辑?

风险

  • LoRA 实现可能依赖特定模型结构(如 key 映射),需确认测试是否需要 mock 模型
任务 5可选 · medium · CPU · 2–3 个晚上

test: 为 offload 策略添加 CPU 行为验证

无人认领6 条评论更新 2026-02-03

用到的专长:内存管理与 offload 调度逻辑

目标:新建 tests/test_utils/test_offload.py,在 CPU 上用 meta device + disk offload 模拟低显存场景,验证 offload 调度逻辑。

为什么值得长期做:offload 是项目核心卖点(versatile resource management),但 tests/ 下无 test_offload.py;验证 offload 调度逻辑可直接提升低显存场景的可靠性。

怎么介入:Issue #199 讨论 torch 版本,但 offload 测试是正交改进;无 open PR 冲突。
第一个 PR 的边界:新建 tests/test_utils/test_offload.py;不修改 offload 实现。
第一步:阅读 diffsynth_engine/utils/offload.py 中 offload 调度逻辑;阅读 configs/pipeline.py 中 offload_mode/offload_to_disk 参数。
本机怎么复现 / 验证:pip install -e '.[dev]' && pytest tests/test_utils/test_offload.py -v --device=cpu
认领留言(英文,可直接贴到 Issue)
I'd like to create tests/test_utils/test_offload.py to verify the offload scheduling logic in diffsynth_engine/utils/offload.py on CPU. The tests will use meta device and disk offload to simulate low-memory scenarios, verifying that model parameters are correctly moved and restored. No GPU needed. May I proceed? PR in 2-3 days.
大致实施方案
  • 新建 tests/test_utils/test_offload.py
  • 构造一个小型模型(如 nn.Sequential),使用 meta device 初始化
  • 调用 offload 函数,验证模型参数是否正确移动到目标设备
  • 验证 disk offload 时参数是否序列化到磁盘并可在需要时恢复
  • 运行 pytest tests/test_utils/test_offload.py -v 验证
可能涉及的目录或文件
  • tests/test_utils/test_offload.py(新建)
  • diffsynth_engine/utils/offload.py(只读)
验收方式
  • pytest tests/test_utils/test_offload.py -v 在 Mac CPU 上通过
  • 新增测试文件:tests/test_utils/test_offload.py
开工前问题与风险

向维护者确认

  • 是否需要测试 offload 与 pipeline 的集成?还是仅测试 offload 工具函数?

风险

  • disk offload 可能依赖特定文件系统路径,需确认测试环境兼容性
任务 6候补 · low · Mac · 1–2 个晚上

docs: 添加 MPS/CPU attention 性能基准示例

无人认领2 条评论更新 2026-01-19

用到的专长:性能基准与 Apple Silicon 适配

目标:新建 examples/mps_attention_benchmark.py,基于 model_perf_benchmark.py 扩展 MPS 路径,建立 attention 层在 MPS 下的耗时基线。

为什么值得长期做:README 将 Apple Silicon 列为 requirement,但 examples/model_perf_benchmark.py 默认面向 CUDA;MPS 性能基线可直接帮助 Mac 用户优化推理。

怎么介入:Issue #224 涉及 MPS 平台问题,本任务从性能基线角度补充;无 open PR 冲突。
第一个 PR 的边界:新建 examples/mps_attention_benchmark.py;不修改现有代码。
第一步:阅读 examples/model_perf_benchmark.py 的基准测试结构;阅读 diffsynth_engine/models/basic/attention.py 中各后端的 MPS 可用性。
本机怎么复现 / 验证:pip install -e '.[dev]' && python examples/mps_attention_benchmark.py --device=mps && python examples/mps_attention_benchmark.py --device=cpu
认领留言(英文,可直接贴到 Issue)
I'd like to create examples/mps_attention_benchmark.py to establish performance baselines for attention layers on Apple Silicon (MPS). Based on examples/model_perf_benchmark.py, the script will benchmark eager_attn and sdpa_attn at various sequence lengths and output a timing table. This helps Mac users optimize inference. No GPU needed. May I proceed? PR in 1-2 days.
大致实施方案
  • 新建 examples/mps_attention_benchmark.py
  • 复用 model_perf_benchmark.py 的测试框架,device 默认 'mps'(若可用)否则 'cpu'
  • 测试 eager_attn、sdpa_attn 在不同 sequence length 下的耗时
  • 输出耗时表格,保存为 Markdown 或打印到终端
  • 在 Mac 上运行 python examples/mps_attention_benchmark.py 验证
可能涉及的目录或文件
  • examples/mps_attention_benchmark.py(新建)
  • examples/model_perf_benchmark.py(参考)
验收方式
  • python examples/mps_attention_benchmark.py 在 Mac 上成功运行并输出耗时表格
  • 新增示例文件:examples/mps_attention_benchmark.py
开工前问题与风险

向维护者确认

  • 是否需要将基准结果保存为文件?还是仅打印到终端?
  • 是否需要覆盖 VAE 和 DiT 单层的基准?

风险

  • MPS 在某些 attention 实现上可能不支持,需添加 fallback 逻辑