Composing diffusion models through a shared latent space
InstantFusion maps the latents of different diffusion models into a shared latent space, so their denoising trajectories can be aligned and composed.
Latent alignment encoder (LAE). Sigma-conditioned encoders and decoders map both models’ latents into the shared space. Training combines reconstruction, cross-reconstruction, latent alignment, and denoising-velocity alignment losses.
Applications. Once the models share a latent space, a trajectory can be handed off from one model to another mid-denoising. This enables:
- Acceleration: hand off between models with different step budgets.
- Multi-preference composition: combine models optimized for different preferences.
- Cross-model OPD: on-policy distillation of model B’s velocity field v from model A in the shared latent space.
This repository includes LAE training, OPD training, and LAE-based model handoff inference.
The examples below compare the baseline with the LAE reconstruction results for Qwen-Image, FLUX.1, and SD3. The figure also reports mean SSIM over 100 images for each model.
### Cross-model denoising handoffThese examples show the results of switching models and continuing denoising with the second model. The figure includes model pairings beyond the SD3/Qwen-Image inference implementation currently available in this repository.
Trained InstantFusion checkpoints are available on Hugging Face.
git clone https://github.com/InternLM/InstantFusion.git
cd InstantFusion
conda create -n instantfusion python=3.10 -y
conda activate instantfusion
pip install -r requirements.txtUse a compatible DiffSynth-Studio checkout in the same environment. The training scripts can import it through DIFFSYNTH_ROOT:
export DIFFSYNTH_ROOT=/path/to/DiffSynth-StudioProvide a Qwen-Image model directory containing transformer/, text_encoder/, vae/, and tokenizer/, plus an SD3 single-file checkpoint. Model weights are not bundled with this repository.
We use VisPrompt5M for training, which is not included in this repository.
The training data is supplied locally and is not included in this repository:
DATASET_ROOT/
└── train/
├── metadata.csv
└── images/
└── example.png
metadata.csv needs file_name and text columns. Image paths in file_name are relative to DATASET_ROOT/train/:
file_name,text
images/example.png,"A lighthouse above a quiet sea at sunrise"--prepare-prompt-cache creates the required SD3 and Qwen-Image prompt embeddings before training. Training uses ./prompt_cache unless --prompt-cache is set. The cache is runtime data and should stay outside version control.
bash scripts/train_lae.sh \
--qwen-path /path/to/Qwen-Image \
--sd3-path /path/to/sd3.safetensors \
--dataset-path /path/to/DATASET_ROOT \
--output-path ./outputs/lae \
--prepare-prompt-cache \
--size 512 --precision bf16 \
--steps-per-epoch 1000 --epochs 5The LAE checkpoint stores the shared-latent bridge parameters.
Use the LAE checkpoint from the previous stage. OPD trains an SD3 LoRA with Qwen-Image as the frozen teacher. By default it uses 20 sampling steps and supervises zero-based steps 1,2,4,8. Change these with --sampling-steps and --supervised-step-ids when needed.
bash scripts/train_opd.sh \
--qwen-path /path/to/Qwen-Image \
--sd3-path /path/to/sd3.safetensors \
--bridge-checkpoint /path/to/lae.ckpt \
--dataset-path /path/to/DATASET_ROOT \
--output-path ./outputs/opd \
--size 512 --precision bf16 \
--steps-per-epoch 1000 --epochs 5OPD reuses the LAE prompt cache by default. If you use a different dataset for OPD, add --prepare-prompt-cache to populate embeddings for its prompts.
bash scripts/infer.sh \
--qwen-path /path/to/Qwen-Image \
--sd3-path /path/to/sd3.safetensors \
--lae-checkpoint /path/to/lae.ckpt \
--prompt "A lighthouse above a quiet sea at sunrise" \
--output-dir ./outputs/inferenceThe command-line interfaces expose paths and common controls. Put less frequently changed trainer or model settings in a JSON file and pass --config path/to/config.json to the relevant command above. JSON uses the Python setting names with underscores; model and dataset paths remain command-line arguments. For example, an OPD config can contain:
{
"learning_rate": 0.00001,
"lora_rank": 64,
"reference_weight": 0.1
}To reuse the training prompt cache during inference, an inference config can contain:
{
"prompt_embedding_dir": "./prompt_cache",
"sd3_cfg": 4.0,
"qwen_cfg": 4.0
}Explicit command-line options take precedence over JSON settings. Training commands also accept the previous underscore-style option names, though --help displays the shorter hyphen-style names. Run bash scripts/train_lae.sh --help, bash scripts/train_opd.sh --help, or bash scripts/infer.sh --help for the compact option lists.
fig/ Method and result figures (PNG and source PDF)
scripts/train_lae.sh LAE training entry point
scripts/train_opd.sh OPD training entry point
scripts/infer.sh LAE-based inference entry point
src/instantfusion/ Bridge, model-loading, cache, training, and inference modules
The paper link and BibTeX entry will be added when the paper is available.



