Qwen-Image-2.1-PE-T2I-Pocket-2B

Built with Qwen — non-commercial research use (Qwen Research License, see LICENSE in this repo)

A 2B text-only prompt rewriter for Qwen/Qwen-Image-2.1: it turns a short user request into the long English description and aspect ratio the image model expects from its prompt-rewriting step — without the 9B thinking model Qwen/Qwen-Image-2.1-PE-T2I, its 1,700-word system prompt, or any thinking tokens.

It is a full fine-tune of Qwen/Qwen3.5-2B (Apache 2.0, no architectural change), distilled from teacher outputs: {user: raw request} → {assistant: one compact JSON object}. No system prompt and thinking are disabled in the chat template, so the rewrite behaviour — including the JSON format — is baked into the weights.

{"rewritten_prompt": "...", "wh_ratio": "3:2"}

Teacher, data and filter flags: ML-Intern-lab/Qwen-Image-2.1-rewriter-distill. Training data is the hard-filtered subset (1,776 pairs) of 8,797 teacher-labelled requests.

Usage

import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from diffusers import QwenImage21Pipeline  # diffusers from git main: 0.41.0.dev0+

tok = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B")
model = AutoModelForCausalLM.from_pretrained(
    "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B",
    torch_dtype=torch.bfloat16,
    device_map="cuda",
)

def rewrite(request: str) -> dict:
    """Raw request -> {"rewritten_prompt": str, "wh_ratio": str}.

    No system prompt, thinking disabled: the chat template emits an empty
    think block before the answer, exactly as at training time, so the model
    continues directly with the JSON object.
    """
    prompt = tok.apply_chat_template(
        [{"role": "user", "content": request}],
        add_generation_prompt=True, tokenize=False,
    )
    ids = tok(prompt, return_tensors="pt").to(model.device)
    gen = model.generate(
        **ids,
        max_new_tokens=1024,
        do_sample=True, temperature=1.0, top_p=0.95, top_k=20,
        pad_token_id=tok.pad_token_id or tok.eos_token_id,
    )
    text = tok.decode(gen[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)
    return json.loads(text)

# Render with Qwen-Image-2.1 (diffusers from git main; torchvision must be importable)
pipe = QwenImage21Pipeline.from_pretrained(
    "Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16, device_map="cuda"
)

# ~1 megapixel (height, width), multiples of 32 — same table as the image eval
RATIO_SIZES = {
    "1:1": (1024, 1024), "3:2": (1248, 832), "2:3": (832, 1248),
    "16:9": (1376, 768), "9:16": (768, 1376), "4:3": (1184, 896),
    "3:4": (896, 1184), "2:1": (1472, 736), "1:2": (736, 1472),
    "21:9": (1568, 672), "9:21": (672, 1568), "4:5": (928, 1152),
    "5:4": (1152, 928), "3:1": (1728, 576), "1:3": (576, 1760),
}

r = rewrite("a photo of a red bicycle leaning against a bakery door")
w, h = RATIO_SIZES.get(r["wh_ratio"], (1248, 832))
image = pipe(
    r["rewritten_prompt"], height=h, width=w,
    num_inference_steps=40,          # default 40 steps, true_cfg_scale 1.0 = no CFG
).images[0]
image.save("bicycle.png")

Evaluation

Measured on the 300 held-out eval requests of the distill dataset, generating every arm with the teacher's own sampling protocol (temperature 1.0, top_p 0.95, top_k 20, seed 0). Baseline = the untuned 2B given the teacher's full system prompt with thinking enabled (cut at 80 rows by the eval job's time budget — a partial floor, not a ceiling).

metric teacher (9B) student-0.8B this model baseline-2B
rows 300 300 300 80
valid JSON rate 100.0% 99.7% 100.0% 77.5%
allowed-ratio rate 100.0% 99.3% 99.7% 20.0%
text fidelity (quoted strings verbatim) 53.1% 53.1% 60.2% 3.7%
fidelity by script Arabic 66.7%, Devanagari 0.0%, Han 60.0%, Japanese 66.7%, Latin 52.0% Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 33.3%, Latin 57.1% Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 50.0%, Latin 64.3% Japanese 0.0%, Latin 3.9%
ratio agreement w/ teacher 100.0% 57.7% 67.3% 8.8%
gen tokens mean / median 1630.8 / 1536.5 453.1 / 456.5 482.8 / 462.0 6106.4 / 6144.0
latency s mean / median n/a 2.81 / 2.61 3.22 / 3.12 29.90 / 29.81

The 2B is the better student across the board: format compliance on par with the teacher, the best quoted-text fidelity of the two (60.2%), and the closest ratio agreement (67.3%) — at ~30% of the teacher's token cost.

Limitations

  • Non-commercial. Trained on Qwen-Image-2.1-PE-T2I outputs under the Qwen Research License: research/evaluation use only; commercial use requires a separate licence from Qwen (see NOTICE).
  • English output for any input. Like the teacher, it rewrites non-English requests into English descriptions; quoted in-image text is copied verbatim in the original script at roughly the teacher's rate (weaker for Arabic/Devanagari than Latin — see fidelity by script).
  • Shorter than the teacher. Trained on rewrites of 80–400 words (fits max_length 1024); the teacher writes up to ~4× longer.
  • Format is learned, not guaranteed. ~0-0.3% of generations are not valid JSON; ~0.3% of ratios fall outside the allowed set.
  • Ratio choice. A user-stated ratio was enforced in training, but the model does not always follow it, and it often picks a different reasonable ratio than the teacher would.
  • Not a general chat model. Single task; baked-in instructions, no system prompt.
  • No GGUF export for this size in this project (the 0.8B has one; the 2B runs fine on small GPUs at ~3.2 s/rewrite in bf16).

Training

TRL 1.13.0 SFTTrainer, full fine-tune, bf16, 2 epochs, lr 1e-5 cosine (warmup 3%), effective batch 32 (8×4), max_length 1024, assistant-only loss, gradient checkpointing, seed 42. Final train loss 1.462 (best of the two students; the 0.8B plateaued at 1.693). Trackio: https://huggingface.co/spaces/ML-Intern-lab/Qwen-Image-2.1-rewriter-distill-trackio

License

This model is a derivative of Qwen-Image-2.1-PE-T2I outputs and is distributed under the Qwen Research License Agreement — see LICENSE (full text) and NOTICE (attribution per §3(c), modification statement per §3(b), and the Built with Qwen mark per §4(b)). Non-commercial use only.

Downloads last month
976
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(410)
this model

Spaces using ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B 2

Collection including ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B