StarVLA QwenPI_v3 for SIMPLE: Joint 10-Task Policy (100k)
This repository contains the 100,000-step StarVLA QwenPI_v3 checkpoint for
the SIMPLE humanoid
simulation benchmark. A single policy is jointly trained on all ten tasks and
is evaluated on every task without task-specific fine-tuning.
The result below uses SIMPLE's task-native episode limits and success thresholds.
Model details
| Item | Value |
|---|---|
| Framework | StarVLA QwenPI_v3 |
| Base VLM | Qwen/Qwen3-VL-4B-Instruct |
| Training data | Combined ten-task SIMPLE/Psi0 mixture (simple_psi0_all_ten_tasks) |
| Data format | LeRobot v2.1 |
| Observation | One 224 x 224 egocentric RGB image, language instruction, 32-D robot state |
| Action representation | Absolute joint position (abs_qpos), 36 dimensions |
| Action horizon | 30 steps |
| Action model | 36-layer LayerwiseFM DiT, hidden size 1024, 16 attention heads |
| Policy inference | Flow matching, 4 inference steps, repeated diffusion 4 |
| Normalization | Min-max |
| Checkpoint format | Complete StarVLA framework state dict (.pt) |
The ten jointly trained tasks are:
G1WholebodyXMovePickTeleop-v0G1WholebodyBendPickMP-v0G1WholebodyHandoverTeleop-v0G1WholebodyLocomotionPickBetweenTablesTeleop-v0G1WholebodyTabletopGraspMP-v0G1WholebodyXMoveBendPickTeleop-v0G1WholebodyCloseDoorTeleop-v0G1WholebodyOpenOvenTeleop-v0G1WholebodyOpenFaucetTeleop-v0G1WholebodyPickAndPlaceAndHugContainerTeleop-v0
Checkpoint and training configuration
The included checkpoint is:
checkpoints/steps_100000_pytorch_model.pt
This is the 100k intermediate snapshot of a run configured for 200,000 optimizer steps; checkpoints through 200k were recorded by the original run. The bundle here intentionally contains only the 100k snapshot.
| Setting | Value |
|---|---|
| Per-GPU batch size | 32 |
| Gradient accumulation | 1 |
| Base VLM learning rate | 1e-5 |
| VLM interface learning rate | 1e-5 |
| Action-head learning rate | 1e-4 |
| Schedule | Cosine, 5,000 warmup steps, 5e-7 minimum LR |
| Loss | VLA loss 1.0; VLM language-model loss 0.0 |
| Frozen module | qwen_vl_interface |
| Gradient checkpointing | Enabled |
| Save interval | 10,000 optimizer steps |
| Random seed | 42 |
The checkpoint is not a LoRA or adapter-only artifact. It must be restored by
StarVLA rather than loaded directly with
transformers.AutoModel.from_pretrained().
SIMPLE native-profile evaluation
Evaluation uses three environment levels and 10 episodes per level, giving 30 episodes per task and 300 episodes overall. The task-native limits and success thresholds are:
| Task | Maximum episode steps | Success threshold |
|---|---|---|
| XMovePickTeleop | 800 | 0.9 |
| BendPickMP | 400 | 0.9 |
| HandoverTeleop | 800 | 0.9 |
| LocomotionPickBetweenTablesTeleop | 1200 | 0.4 |
| TabletopGraspMP | 800 | 0.9 |
| XMoveBendPickTeleop | 800 | 0.8 |
| CloseDoorTeleop | 1000 | 0.5 |
| OpenOvenTeleop | 1000 | 0.5 |
| OpenFaucetTeleop | 1000 | 0.2 |
| PickAndPlaceAndHugContainerTeleop | 1000 | 0.9 |
Each cell below is Level 0 | Level 1 | Level 2, with each level scored out
of 10 episodes.
| Task | Successes by level | Total |
|---|---|---|
| XMovePickTeleop | 10 | 10 | 9 | 29/30 |
| BendPickMP | 8 | 6 | 6 | 20/30 |
| HandoverTeleop | 9 | 9 | 10 | 28/30 |
| LocomotionPickBetweenTablesTeleop | 8 | 7 | 4 | 19/30 |
| TabletopGraspMP | 10 | 9 | 7 | 26/30 |
| XMoveBendPickTeleop | 9 | 8 | 9 | 26/30 |
| CloseDoorTeleop | 10 | 10 | 10 | 30/30 |
| OpenOvenTeleop | 9 | 9 | 9 | 27/30 |
| OpenFaucetTeleop | 5 | 9 | 6 | 20/30 |
| PickAndPlaceAndHugContainerTeleop | 7 | 5 | 8 | 20/30 |
| Overall | 245/300 (81.67%) |
Result provenance
Nine task rows come from the complete 100k SIMPLE-native evaluation run dated
2026-07-29. XMovePick is taken from the instruction-aligned rerun dated
2026-07-30. The original environment instruction said “bend to pick up the
cracker box”, while this policy's training metadata uses “move forward to pick
up the cracker box”. The aligned rerun uses the training instruction and
replaces the mismatched XMovePick result (1 | 2 | 0). Without this replacement,
the native-profile aggregate is 219/300 (73.00%).
The XMoveBendPick Level-2 launcher returned a nonzero exit status after writing all ten episode outcomes. The reported 9/10 is recomputed directly from those ten unique completed episode records. Failed or incomplete retry directories are excluded.
Loading and usage
Preserve the following logical layout after restoring the archive:
MODEL_DIR/
|-- README.md
|-- config.yaml
|-- config.full.yaml
|-- dataset_statistics.json
`-- checkpoints/
`-- steps_100000_pytorch_model.pt
Keep the two configuration files and dataset_statistics.json beside the
checkpoints/ directory. Start the StarVLA policy server with
steps_100000_pytorch_model.pt, then run SIMPLE in its separate simulator
environment using the task-native settings listed above.
Limitations
- Results cover one checkpoint and one set of 10 episodes per task and level; confidence intervals and multi-seed variance were not measured.
- The instruction-aligned XMovePick row corrects an evaluation prompt mismatch and should be reported with that provenance.
- Results are simulator measurements and do not establish real-robot safety or performance.
- Camera calibration, robot embodiment, instruction wording, action conventions, and scene distribution shifts can materially affect behavior.
This model is intended for research evaluation. Do not deploy it directly on a physical robot without task-specific validation and independent safety controls.
Download and restore the checkpoint
The 100k checkpoint is published as 3 numbered byte-stream parts,
not independent PyTorch checkpoints. Restore the original .pt before loading:
hf download StarVLA/Qwen3-VL-4B-PI-SIMPLE --local-dir ./Qwen3-VL-4B-PI-SIMPLE
cd Qwen3-VL-4B-PI-SIMPLE
python merge_checkpoint.py
Keep all parts in their manifest order. The helper verifies the restored
checkpoint's byte count and SHA256. Restoration needs at least
11 GB of additional free space on a filesystem supporting files larger
than 4 GB. SHA256SUMS refers to the restored checkpoint.
The configuration uses public base-model identifiers and relative data/output directories. Download the base model from its public repository or override the base-model fields for your deployment. Training topology, machine addresses, account identifiers, and private filesystem paths are omitted from this release.
- Downloads last month
- 14
Model tree for StarVLA/Qwen3-VL-4B-PI-SIMPLE
Base model
Qwen/Qwen3-VL-4B-Instruct