SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
Abstract
SynthGait-19k is a large synthetic video dataset for gait analysis that enables benchmarking of video-based gait estimation and shows synthetic supervision transfers to real data.
Accurate estimation of clinically meaningful gait parameters from monocular video is important for scalable mobility assessment, yet progress is limited by the small scale, restricted viewpoints, and limited visual diversity of existing datasets. We introduce SynthGait-19k, a physically grounded synthetic video dataset containing 19,272 walking videos derived from 6,427 MoCap sequences across 437 subjects, with paired SMPL motion and annotations for six gait parameters. To construct the dataset, we develop Gait2Vid, which unifies heterogeneous MoCap recordings through SMPL and synthesizes diverse RGB walking videos under controllable viewpoints and scene appearances. We assess the generated videos for consistency with their conditioning gait kinematics and validate extracted gait events against force-platform measurements. Using SynthGait-19K, we benchmark direct RGB, pose-based, biomechanical, and human-mesh-recovery approaches and analyze viewpoint, training-data scale, and synthetic-to-real domain shift. We also introduce GaitXFormer as a direct RGB reference model for estimating gait parameters. Synthetic supervision transfers effectively to real videos across both GaitXFormer and a pose-based architecture, demonstrating utility across different representations. We further find that spatial gait parameters are more sensitive to visual domain shift and that improved HMR reconstruction alone does not necessarily translate to improved downstream gait estimation.
Community
Annotated walking videos with reliable gait ground truth are difficult to collect at scale, partly due to privacy concerns and the need for specialized motion-capture setups. We introduce SynthGait-19K, a synthetic dataset of 19,272 RGB walking videos generated from 6,427 MoCap sequences across 437 subjects, with ground-truth annotations for six clinically relevant gait parameters.
Our Gait2Vid pipeline unifies MoCap data from multiple datasets and converts it into diverse, controllable RGB walking videos while preserving the underlying gait kinematics. We also introduce GaitXFormer, a direct RGB-to-gait model that estimates gait parameters from video without relying on intermediate pose or mesh representations and can run in real time.
Experiments show that training with SynthGait-19K transfers effectively to real-world videos, providing a scalable way to develop video-based gait analysis systems.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Decoding Children's Gait Behavior (2026)
- Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction (2026)
- Seeing the World and the Self from Egocentric Video (2026)
- TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views (2026)
- Kirin: Animal Motion Generation from In-the-Wild Video (2026)
- MuyBridge: Mobile Human Center-of-Mass Estimation from Monocular Video via Sparse Fusion (2026)
- HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.08108 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper