6.1 KiB
FLUX.2 Setup in ComfyUI — Full Reproduction Recipe
Hardware Context (This Session)
- Proxmox CT 204, ROCm 6.2, PyTorch 2.6.0 (ROCm build)
- AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (gfx1150)
HSA_OVERRIDE_GFX_VERSION=11.0.0required for gfx1150- 18GB "VRAM" reported; ~8.5 GB disk free after downloads
Timeline
| Step | Result |
|---|---|
| FLUX.1-dev E2E | ✅ Works in 46s (512×512, 4 steps) |
| FLUX.2 with FLUX.1 CLIP/T5 | ❌ mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096) |
| FLUX.2 with Qwen3VL + ae.safetensors | ❌ expected input[1, 128, 32, 32] to have 16 channels, but got 128 channels instead |
| FLUX.2 with Qwen3VL + taef2 | ✅ Success ~110s (1024×1024, 4 steps) |
Required Models
1. UNet (GGUF)
# ~5.4 GB
wget https://huggingface.co/unsloth/FLUX.2-klein-9B-GGUF/resolve/main/flux-2-klein-9b-Q4_K_S.gguf
# → models/unet/flux-2-klein-9b-Q4_K_S.gguf
2. Text Encoder (Qwen3-VL-8B, GGUF)
CRITICAL: FLUX.2 uses Qwen3-VL-8B, NOT CLIP-L + T5XXL.
# ~4.7 GB — Q4_K_M quantization
wget https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF/resolve/main/Qwen3VL-8B-Instruct-Q4_K_M.gguf
# → models/clip/Qwen3VL-8B-Instruct-Q4_K_M.gguf
Alternative: Qwen3-VL-8B-Instruct-Q8_0.gguf (~8.3 GB) for higher quality.
The GGUF detection logic in ComfyUI (comfy/sd.py:detect_te_model) identifies this model by:
model.layers.0.self_attn.q_proj.weightshape =[4096, 4096]num_hidden_layers = 36→ maps toTEModel.QWEN3VL_8B
3. VAE — TWO OPTIONS
Option A: TAEF2 (Tiny AutoEncoder, fast, lower quality)
# ~5 MB total, public, no auth needed
cd models/vae_approx
curl -LO https://github.com/madebyollin/taesd/raw/main/taef2_encoder.pth
curl -LO https://github.com/madebyollin/taesd/raw/main/taef2_decoder.pth
# Reference as vae_name: "taef2" in VAELoader node
Option B: Full FLUX.2 VAE (flux2_vae.safetensors)
# ~160 MB, gated repo (BFL license acceptance required)
# From: black-forest-labs/FLUX.2-klein-9B vae/diffusion_pytorch_model.safetensors
# Save as: models/vae/flux2_vae.safetensors
The FLUX.1 VAE (ae.safetensors, 16 channels) is incompatible with FLUX.2 (128 channels).
ComfyUI Workflow (API JSON)
{
"1": {
"inputs": {"unet_name": "flux-2-klein-9b-Q4_K_S.gguf"},
"class_type": "UnetLoaderGGUF"
},
"2": {
"inputs": {
"clip_name": "Qwen3VL-8B-Instruct-Q4_K_M.gguf",
"type": "flux2"
},
"class_type": "CLIPLoaderGGUF"
},
"3": {
"inputs": {"vae_name": "taef2"},
"class_type": "VAELoader"
},
"4": {
"inputs": {"width": 1024, "height": 1024, "batch_size": 1},
"class_type": "EmptyFlux2LatentImage"
},
"5": {
"inputs": {
"text": "A beautiful mountain lake at sunset, photorealistic, 4k",
"clip": ["2", 0]
},
"class_type": "CLIPTextEncode"
},
"6": {
"inputs": {"text": "", "clip": ["2", 0]},
"class_type": "CLIPTextEncode"
},
"7": {
"inputs": {
"seed": 42,
"steps": 4,
"cfg": 1.0,
"sampler_name": "euler",
"scheduler": "simple",
"denoise": 1.0,
"model": ["1", 0],
"positive": ["5", 0],
"negative": ["6", 0],
"latent_image": ["4", 0]
},
"class_type": "KSampler"
},
"8": {
"inputs": {"samples": ["7", 0], "vae": ["3", 0]},
"class_type": "VAEDecode"
},
"9": {
"inputs": {"filename_prefix": "flux2_out", "images": ["8", 0]},
"class_type": "SaveImage"
}
}
Key node differences from FLUX.1
| Node | FLUX.1 | FLUX.2 |
|---|---|---|
| CLIP loader | DualCLIPLoaderGGUF (clip_l + t5xxl) |
CLIPLoaderGGUF (single Qwen3VL) |
| CLIP type | "flux" |
"flux2" |
| Empty latent | EmptyLatentImage (16 channels) |
EmptyFlux2LatentImage (128 channels) |
| VAE | ae.safetensors |
taef2 or flux2_vae.safetensors |
ROCm gfx1150 Workaround
gfx1150 (RDNA3.5, Radeon 890M) is not fully supported by ROCm 6.2. Set override before launch:
export HSA_OVERRIDE_GFX_VERSION=11.0.0
python main.py --listen 0.0.0.0 --port 8188
Without this, CLIP encoding fails with:
RuntimeError: HIP error: invalid device function
Performance (gfx1150, ROCm 6.2)
| Pipeline | Resolution | Steps | Time |
|---|---|---|---|
| FLUX.1-dev Q4_K_S + CLIP/T5 + ae.safetensors | 512×512 | 4 | ~46s |
| FLUX.2-klein Q4_K_S + Qwen3VL Q4_K_M + taef2 | 1024×1024 | 4 | ~110s |
FLUX.2 is slower despite fewer parameters because Qwen3VL-8B text encoding is heavier than CLIP+T5, and the TAEF2 decode takes extra time.
Disk Space Budget
| Component | Size |
|---|---|
| FLUX.2 UNet Q4_K_S | 5.4 GB |
| Qwen3VL-8B Q4_K_M | 4.7 GB |
| taef2 encoder+decoder | 5 MB |
| Total | ~10.1 GB |
Full FLUX.2 VAE adds ~160 MB. Q8_0 text encoder adds ~3.6 GB more.
Common Errors
mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096)
Cause: Using CLIP-L + T5XXL (FLUX.1 encoders) with FLUX.2 UNet.
Fix: Switch to Qwen3VL-8B-Instruct-Q4_K_M.gguf + CLIPLoaderGGUF with "type": "flux2".
expected input[1, 128, 32, 32] to have 16 channels, but got 128 channels instead
Cause: Using FLUX.1 VAE (ae.safetensors) with FLUX.2 latents.
Fix: Use taef2 or download dedicated flux2_vae.safetensors.
HIP error: invalid device function
Cause: ROCm doesn't natively support gfx1150.
Fix: export HSA_OVERRIDE_GFX_VERSION=11.0.0 before starting ComfyUI.
RuntimeError: The expanded size of the tensor (512) must match ...
Cause: Using EmptyLatentImage (16-channel) instead of EmptyFlux2LatentImage (128-channel).
Fix: Change node class to EmptyFlux2LatentImage.
References
- ComfyUI
comfy/model_base.pylines 1057-1066:class Flux2(Flux) - ComfyUI
comfy/latent_formats.pyline 192:class Flux2(LatentFormat)withlatent_channels = 128 - ComfyUI
comfy/sd.pylines 1631-1634: FLUX2 CLIPType routes toklein_tewith Qwen3VL - TAEF2 source: https://github.com/madebyollin/taesd
- Unsloth GGUF: https://huggingface.co/unsloth/FLUX.2-klein-9B-GGUF
- Qwen GGUF: https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF