Instructions to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B") model = AutoModelForCausalLM.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B
- SGLang
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B with Docker Model Runner:
docker model run hf.co/ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B
Qwen-Image-2.1-PE-T2I-Pocket-2B
Built with Qwen — non-commercial research use (Qwen Research License, see LICENSE in this repo)
A 2B text-only prompt rewriter for Qwen/Qwen-Image-2.1: it turns a short user request into the long English description and aspect ratio the image model expects from its prompt-rewriting step — without the 9B thinking model Qwen/Qwen-Image-2.1-PE-T2I, its 1,700-word system prompt, or any thinking tokens.
It is a full fine-tune of Qwen/Qwen3.5-2B (Apache 2.0, no architectural change), distilled from teacher outputs: {user: raw request} → {assistant: one compact JSON object}. No system prompt and thinking are disabled in the chat template, so the rewrite behaviour — including the JSON format — is baked into the weights.
{"rewritten_prompt": "...", "wh_ratio": "3:2"}
Teacher, data and filter flags: ML-Intern-lab/Qwen-Image-2.1-rewriter-distill. Training data is the hard-filtered subset (1,776 pairs) of 8,797 teacher-labelled requests.
Usage
import json
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from diffusers import QwenImage21Pipeline # diffusers from git main: 0.41.0.dev0+
tok = AutoTokenizer.from_pretrained("ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B")
model = AutoModelForCausalLM.from_pretrained(
"ML-Intern-lab/Qwen-Image-2.1-PE-T2I-Pocket-2B",
torch_dtype=torch.bfloat16,
device_map="cuda",
)
def rewrite(request: str) -> dict:
"""Raw request -> {"rewritten_prompt": str, "wh_ratio": str}.
No system prompt, thinking disabled: the chat template emits an empty
think block before the answer, exactly as at training time, so the model
continues directly with the JSON object.
"""
prompt = tok.apply_chat_template(
[{"role": "user", "content": request}],
add_generation_prompt=True, tokenize=False,
)
ids = tok(prompt, return_tensors="pt").to(model.device)
gen = model.generate(
**ids,
max_new_tokens=1024,
do_sample=True, temperature=1.0, top_p=0.95, top_k=20,
pad_token_id=tok.pad_token_id or tok.eos_token_id,
)
text = tok.decode(gen[0][ids["input_ids"].shape[1]:], skip_special_tokens=True)
return json.loads(text)
# Render with Qwen-Image-2.1 (diffusers from git main; torchvision must be importable)
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16, device_map="cuda"
)
# ~1 megapixel (height, width), multiples of 32 — same table as the image eval
RATIO_SIZES = {
"1:1": (1024, 1024), "3:2": (1248, 832), "2:3": (832, 1248),
"16:9": (1376, 768), "9:16": (768, 1376), "4:3": (1184, 896),
"3:4": (896, 1184), "2:1": (1472, 736), "1:2": (736, 1472),
"21:9": (1568, 672), "9:21": (672, 1568), "4:5": (928, 1152),
"5:4": (1152, 928), "3:1": (1728, 576), "1:3": (576, 1760),
}
r = rewrite("a photo of a red bicycle leaning against a bakery door")
w, h = RATIO_SIZES.get(r["wh_ratio"], (1248, 832))
image = pipe(
r["rewritten_prompt"], height=h, width=w,
num_inference_steps=40, # default 40 steps, true_cfg_scale 1.0 = no CFG
).images[0]
image.save("bicycle.png")
Evaluation
Measured on the 300 held-out eval requests of the distill dataset, generating every arm with the teacher's own sampling protocol (temperature 1.0, top_p 0.95, top_k 20, seed 0). Baseline = the untuned 2B given the teacher's full system prompt with thinking enabled (cut at 80 rows by the eval job's time budget — a partial floor, not a ceiling).
| metric | teacher (9B) | student-0.8B | this model | baseline-2B |
|---|---|---|---|---|
| rows | 300 | 300 | 300 | 80 |
| valid JSON rate | 100.0% | 99.7% | 100.0% | 77.5% |
| allowed-ratio rate | 100.0% | 99.3% | 99.7% | 20.0% |
| text fidelity (quoted strings verbatim) | 53.1% | 53.1% | 60.2% | 3.7% |
| fidelity by script | Arabic 66.7%, Devanagari 0.0%, Han 60.0%, Japanese 66.7%, Latin 52.0% | Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 33.3%, Latin 57.1% | Arabic 33.3%, Devanagari 0.0%, Han 20.0%, Japanese 50.0%, Latin 64.3% | Japanese 0.0%, Latin 3.9% |
| ratio agreement w/ teacher | 100.0% | 57.7% | 67.3% | 8.8% |
| gen tokens mean / median | 1630.8 / 1536.5 | 453.1 / 456.5 | 482.8 / 462.0 | 6106.4 / 6144.0 |
| latency s mean / median | n/a | 2.81 / 2.61 | 3.22 / 3.12 | 29.90 / 29.81 |
The 2B is the better student across the board: format compliance on par with the teacher, the best quoted-text fidelity of the two (60.2%), and the closest ratio agreement (67.3%) — at ~30% of the teacher's token cost.
Limitations
- Non-commercial. Trained on Qwen-Image-2.1-PE-T2I outputs under the Qwen Research License: research/evaluation use only; commercial use requires a separate licence from Qwen (see NOTICE).
- English output for any input. Like the teacher, it rewrites non-English requests into English descriptions; quoted in-image text is copied verbatim in the original script at roughly the teacher's rate (weaker for Arabic/Devanagari than Latin — see fidelity by script).
- Shorter than the teacher. Trained on rewrites of 80–400 words (fits
max_length 1024); the teacher writes up to ~4× longer. - Format is learned, not guaranteed. ~0-0.3% of generations are not valid JSON; ~0.3% of ratios fall outside the allowed set.
- Ratio choice. A user-stated ratio was enforced in training, but the model does not always follow it, and it often picks a different reasonable ratio than the teacher would.
- Not a general chat model. Single task; baked-in instructions, no system prompt.
- No GGUF export for this size in this project (the 0.8B has one; the 2B runs fine on small GPUs at ~3.2 s/rewrite in bf16).
Training
TRL 1.13.0 SFTTrainer, full fine-tune, bf16, 2 epochs, lr 1e-5 cosine (warmup 3%), effective batch 32 (8×4), max_length 1024, assistant-only loss, gradient checkpointing, seed 42. Final train loss 1.462 (best of the two students; the 0.8B plateaued at 1.693). Trackio: https://huggingface.co/spaces/ML-Intern-lab/Qwen-Image-2.1-rewriter-distill-trackio
License
This model is a derivative of Qwen-Image-2.1-PE-T2I outputs and is distributed under the Qwen Research License Agreement — see LICENSE (full text) and NOTICE (attribution per §3(c), modification statement per §3(b), and the Built with Qwen mark per §4(b)). Non-commercial use only.
- Downloads last month
- 976