StarVLA QwenPI_v3 for SIMPLE: Joint 10-Task Policy (100k)

This repository contains the 100,000-step StarVLA QwenPI_v3 checkpoint for the SIMPLE humanoid simulation benchmark. A single policy is jointly trained on all ten tasks and is evaluated on every task without task-specific fine-tuning.

The result below uses SIMPLE's task-native episode limits and success thresholds.

Model details

Item Value
Framework StarVLA QwenPI_v3
Base VLM Qwen/Qwen3-VL-4B-Instruct
Training data Combined ten-task SIMPLE/Psi0 mixture (simple_psi0_all_ten_tasks)
Data format LeRobot v2.1
Observation One 224 x 224 egocentric RGB image, language instruction, 32-D robot state
Action representation Absolute joint position (abs_qpos), 36 dimensions
Action horizon 30 steps
Action model 36-layer LayerwiseFM DiT, hidden size 1024, 16 attention heads
Policy inference Flow matching, 4 inference steps, repeated diffusion 4
Normalization Min-max
Checkpoint format Complete StarVLA framework state dict (.pt)

The ten jointly trained tasks are:

  • G1WholebodyXMovePickTeleop-v0
  • G1WholebodyBendPickMP-v0
  • G1WholebodyHandoverTeleop-v0
  • G1WholebodyLocomotionPickBetweenTablesTeleop-v0
  • G1WholebodyTabletopGraspMP-v0
  • G1WholebodyXMoveBendPickTeleop-v0
  • G1WholebodyCloseDoorTeleop-v0
  • G1WholebodyOpenOvenTeleop-v0
  • G1WholebodyOpenFaucetTeleop-v0
  • G1WholebodyPickAndPlaceAndHugContainerTeleop-v0

Checkpoint and training configuration

The included checkpoint is:

checkpoints/steps_100000_pytorch_model.pt

This is the 100k intermediate snapshot of a run configured for 200,000 optimizer steps; checkpoints through 200k were recorded by the original run. The bundle here intentionally contains only the 100k snapshot.

Setting Value
Per-GPU batch size 32
Gradient accumulation 1
Base VLM learning rate 1e-5
VLM interface learning rate 1e-5
Action-head learning rate 1e-4
Schedule Cosine, 5,000 warmup steps, 5e-7 minimum LR
Loss VLA loss 1.0; VLM language-model loss 0.0
Frozen module qwen_vl_interface
Gradient checkpointing Enabled
Save interval 10,000 optimizer steps
Random seed 42

The checkpoint is not a LoRA or adapter-only artifact. It must be restored by StarVLA rather than loaded directly with transformers.AutoModel.from_pretrained().

SIMPLE native-profile evaluation

Evaluation uses three environment levels and 10 episodes per level, giving 30 episodes per task and 300 episodes overall. The task-native limits and success thresholds are:

Task Maximum episode steps Success threshold
XMovePickTeleop 800 0.9
BendPickMP 400 0.9
HandoverTeleop 800 0.9
LocomotionPickBetweenTablesTeleop 1200 0.4
TabletopGraspMP 800 0.9
XMoveBendPickTeleop 800 0.8
CloseDoorTeleop 1000 0.5
OpenOvenTeleop 1000 0.5
OpenFaucetTeleop 1000 0.2
PickAndPlaceAndHugContainerTeleop 1000 0.9

Each cell below is Level 0 | Level 1 | Level 2, with each level scored out of 10 episodes.

Task Successes by level Total
XMovePickTeleop 10 | 10 | 9 29/30
BendPickMP 8 | 6 | 6 20/30
HandoverTeleop 9 | 9 | 10 28/30
LocomotionPickBetweenTablesTeleop 8 | 7 | 4 19/30
TabletopGraspMP 10 | 9 | 7 26/30
XMoveBendPickTeleop 9 | 8 | 9 26/30
CloseDoorTeleop 10 | 10 | 10 30/30
OpenOvenTeleop 9 | 9 | 9 27/30
OpenFaucetTeleop 5 | 9 | 6 20/30
PickAndPlaceAndHugContainerTeleop 7 | 5 | 8 20/30
Overall 245/300 (81.67%)

Result provenance

Nine task rows come from the complete 100k SIMPLE-native evaluation run dated 2026-07-29. XMovePick is taken from the instruction-aligned rerun dated 2026-07-30. The original environment instruction said “bend to pick up the cracker box”, while this policy's training metadata uses “move forward to pick up the cracker box”. The aligned rerun uses the training instruction and replaces the mismatched XMovePick result (1 | 2 | 0). Without this replacement, the native-profile aggregate is 219/300 (73.00%).

The XMoveBendPick Level-2 launcher returned a nonzero exit status after writing all ten episode outcomes. The reported 9/10 is recomputed directly from those ten unique completed episode records. Failed or incomplete retry directories are excluded.

Loading and usage

Preserve the following logical layout after restoring the archive:

MODEL_DIR/
|-- README.md
|-- config.yaml
|-- config.full.yaml
|-- dataset_statistics.json
`-- checkpoints/
    `-- steps_100000_pytorch_model.pt

Keep the two configuration files and dataset_statistics.json beside the checkpoints/ directory. Start the StarVLA policy server with steps_100000_pytorch_model.pt, then run SIMPLE in its separate simulator environment using the task-native settings listed above.

Limitations

  • Results cover one checkpoint and one set of 10 episodes per task and level; confidence intervals and multi-seed variance were not measured.
  • The instruction-aligned XMovePick row corrects an evaluation prompt mismatch and should be reported with that provenance.
  • Results are simulator measurements and do not establish real-robot safety or performance.
  • Camera calibration, robot embodiment, instruction wording, action conventions, and scene distribution shifts can materially affect behavior.

This model is intended for research evaluation. Do not deploy it directly on a physical robot without task-specific validation and independent safety controls.

Download and restore the checkpoint

The 100k checkpoint is published as 3 numbered byte-stream parts, not independent PyTorch checkpoints. Restore the original .pt before loading:

hf download StarVLA/Qwen3-VL-4B-PI-SIMPLE --local-dir ./Qwen3-VL-4B-PI-SIMPLE
cd Qwen3-VL-4B-PI-SIMPLE
python merge_checkpoint.py

Keep all parts in their manifest order. The helper verifies the restored checkpoint's byte count and SHA256. Restoration needs at least 11 GB of additional free space on a filesystem supporting files larger than 4 GB. SHA256SUMS refers to the restored checkpoint.

The configuration uses public base-model identifiers and relative data/output directories. Download the base model from its public repository or override the base-model fields for your deployment. Training topology, machine addresses, account identifiers, and private filesystem paths are omitted from this release.

Downloads last month
14
Video Preview
loading

Model tree for StarVLA/Qwen3-VL-4B-PI-SIMPLE

Finetuned
(470)
this model

Dataset used to train StarVLA/Qwen3-VL-4B-PI-SIMPLE