Files
hermes-skills/creative/comfyui/references/flux2-setup.md
T

6.1 KiB
Raw Blame History

FLUX.2 Setup in ComfyUI — Full Reproduction Recipe

Hardware Context (This Session)

  • Proxmox CT 204, ROCm 6.2, PyTorch 2.6.0 (ROCm build)
  • AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (gfx1150)
  • HSA_OVERRIDE_GFX_VERSION=11.0.0 required for gfx1150
  • 18GB "VRAM" reported; ~8.5 GB disk free after downloads

Timeline

Step Result
FLUX.1-dev E2E Works in 46s (512×512, 4 steps)
FLUX.2 with FLUX.1 CLIP/T5 mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096)
FLUX.2 with Qwen3VL + ae.safetensors expected input[1, 128, 32, 32] to have 16 channels, but got 128 channels instead
FLUX.2 with Qwen3VL + taef2 Success ~110s (1024×1024, 4 steps)

Required Models

1. UNet (GGUF)

# ~5.4 GB
wget https://huggingface.co/unsloth/FLUX.2-klein-9B-GGUF/resolve/main/flux-2-klein-9b-Q4_K_S.gguf
# → models/unet/flux-2-klein-9b-Q4_K_S.gguf

2. Text Encoder (Qwen3-VL-8B, GGUF)

CRITICAL: FLUX.2 uses Qwen3-VL-8B, NOT CLIP-L + T5XXL.

# ~4.7 GB — Q4_K_M quantization
wget https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct-GGUF/resolve/main/Qwen3VL-8B-Instruct-Q4_K_M.gguf
# → models/clip/Qwen3VL-8B-Instruct-Q4_K_M.gguf

Alternative: Qwen3-VL-8B-Instruct-Q8_0.gguf (~8.3 GB) for higher quality.

The GGUF detection logic in ComfyUI (comfy/sd.py:detect_te_model) identifies this model by:

  • model.layers.0.self_attn.q_proj.weight shape = [4096, 4096]
  • num_hidden_layers = 36 → maps to TEModel.QWEN3VL_8B

3. VAE — TWO OPTIONS

Option A: TAEF2 (Tiny AutoEncoder, fast, lower quality)

# ~5 MB total, public, no auth needed
cd models/vae_approx
curl -LO https://github.com/madebyollin/taesd/raw/main/taef2_encoder.pth
curl -LO https://github.com/madebyollin/taesd/raw/main/taef2_decoder.pth
# Reference as vae_name: "taef2" in VAELoader node

Option B: Full FLUX.2 VAE (flux2_vae.safetensors)

# ~160 MB, gated repo (BFL license acceptance required)
# From: black-forest-labs/FLUX.2-klein-9B vae/diffusion_pytorch_model.safetensors
# Save as: models/vae/flux2_vae.safetensors

The FLUX.1 VAE (ae.safetensors, 16 channels) is incompatible with FLUX.2 (128 channels).

ComfyUI Workflow (API JSON)

{
  "1": {
    "inputs": {"unet_name": "flux-2-klein-9b-Q4_K_S.gguf"},
    "class_type": "UnetLoaderGGUF"
  },
  "2": {
    "inputs": {
      "clip_name": "Qwen3VL-8B-Instruct-Q4_K_M.gguf",
      "type": "flux2"
    },
    "class_type": "CLIPLoaderGGUF"
  },
  "3": {
    "inputs": {"vae_name": "taef2"},
    "class_type": "VAELoader"
  },
  "4": {
    "inputs": {"width": 1024, "height": 1024, "batch_size": 1},
    "class_type": "EmptyFlux2LatentImage"
  },
  "5": {
    "inputs": {
      "text": "A beautiful mountain lake at sunset, photorealistic, 4k",
      "clip": ["2", 0]
    },
    "class_type": "CLIPTextEncode"
  },
  "6": {
    "inputs": {"text": "", "clip": ["2", 0]},
    "class_type": "CLIPTextEncode"
  },
  "7": {
    "inputs": {
      "seed": 42,
      "steps": 4,
      "cfg": 1.0,
      "sampler_name": "euler",
      "scheduler": "simple",
      "denoise": 1.0,
      "model": ["1", 0],
      "positive": ["5", 0],
      "negative": ["6", 0],
      "latent_image": ["4", 0]
    },
    "class_type": "KSampler"
  },
  "8": {
    "inputs": {"samples": ["7", 0], "vae": ["3", 0]},
    "class_type": "VAEDecode"
  },
  "9": {
    "inputs": {"filename_prefix": "flux2_out", "images": ["8", 0]},
    "class_type": "SaveImage"
  }
}

Key node differences from FLUX.1

Node FLUX.1 FLUX.2
CLIP loader DualCLIPLoaderGGUF (clip_l + t5xxl) CLIPLoaderGGUF (single Qwen3VL)
CLIP type "flux" "flux2"
Empty latent EmptyLatentImage (16 channels) EmptyFlux2LatentImage (128 channels)
VAE ae.safetensors taef2 or flux2_vae.safetensors

ROCm gfx1150 Workaround

gfx1150 (RDNA3.5, Radeon 890M) is not fully supported by ROCm 6.2. Set override before launch:

export HSA_OVERRIDE_GFX_VERSION=11.0.0
python main.py --listen 0.0.0.0 --port 8188

Without this, CLIP encoding fails with:

RuntimeError: HIP error: invalid device function

Performance (gfx1150, ROCm 6.2)

Pipeline Resolution Steps Time
FLUX.1-dev Q4_K_S + CLIP/T5 + ae.safetensors 512×512 4 ~46s
FLUX.2-klein Q4_K_S + Qwen3VL Q4_K_M + taef2 1024×1024 4 ~110s

FLUX.2 is slower despite fewer parameters because Qwen3VL-8B text encoding is heavier than CLIP+T5, and the TAEF2 decode takes extra time.

Disk Space Budget

Component Size
FLUX.2 UNet Q4_K_S 5.4 GB
Qwen3VL-8B Q4_K_M 4.7 GB
taef2 encoder+decoder 5 MB
Total ~10.1 GB

Full FLUX.2 VAE adds ~160 MB. Q8_0 text encoder adds ~3.6 GB more.

Common Errors

mat1 and mat2 shapes cannot be multiplied (512x4096 and 12288x4096)

Cause: Using CLIP-L + T5XXL (FLUX.1 encoders) with FLUX.2 UNet. Fix: Switch to Qwen3VL-8B-Instruct-Q4_K_M.gguf + CLIPLoaderGGUF with "type": "flux2".

expected input[1, 128, 32, 32] to have 16 channels, but got 128 channels instead

Cause: Using FLUX.1 VAE (ae.safetensors) with FLUX.2 latents. Fix: Use taef2 or download dedicated flux2_vae.safetensors.

HIP error: invalid device function

Cause: ROCm doesn't natively support gfx1150. Fix: export HSA_OVERRIDE_GFX_VERSION=11.0.0 before starting ComfyUI.

RuntimeError: The expanded size of the tensor (512) must match ...

Cause: Using EmptyLatentImage (16-channel) instead of EmptyFlux2LatentImage (128-channel). Fix: Change node class to EmptyFlux2LatentImage.

References