Human motion-language alignment

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

A diverse, balanced, and multi-granular testbed for measuring robust motion-language alignment beyond narrow datasets and fixed caption styles.

1 Shanghai Jiao Tong University 2 Beijing Zhongguancun Academy * Equal contribution ‡ Project leader † Corresponding author

Paper Video
3,390Motion sequences
10,170Verified captions
118Fine-grained categories
Concise · Standard · Fine-grained Text granularities
Overview

One benchmark, many ways to move.

MRBench brings together motion capture, in-the-wild videos, synthetic videos, and generated motions—each paired with concise, standard, and fine-grained text.

Why MRBench?

Existing benchmarks hide critical failure modes.

Homogeneous motion sources, long-tailed categories, and repetitive captions can inflate in-domain performance while obscuring true cross-modal understanding.

01

Heterogeneous motion

Four sources probe transfer across capture conditions and generation domains.

MoCap43.8% In-the-wild26.5% Synthetic22.7% Generated7.0%
02

Balanced semantics

A hierarchical taxonomy covers 7 coarse, 33 middle-level, and 118 fine-grained categories.

03

Multi-granular language

Every motion has visually grounded descriptions at three explicit semantic levels.

Concise 9.2 words Standard 16.5 words Fine-grained 43.6 words
Statistical comparison of MRBench, HumanML3D, and KIT-ML
Category coverage, caption length, and retrieval ambiguity across benchmarks. View full-resolution PDF ↗
Construction

From a large motion pool to a reliable benchmark.

A multi-stage pipeline filters weak captions, balances semantic coverage, verifies cross-modal alignment, and produces motion-grounded descriptions.

01

Candidate filtering

Remove generic, repetitive, and non-kinematic descriptions.

02

Balanced sampling

Use a hierarchical taxonomy to enforce diverse category coverage.

03

Alignment verification

Assess whether the described motion is observable and unambiguous.

04

Text expansion

Generate three granularities, followed by careful human validation.

Granularity-Aware Retrieval

Adapt to the query while preserving the anchor.

We freeze a standard-caption-aligned dual encoder and add lightweight, granularity-specific motion extractors and text adapters. LLM-rewritten concise and fine-grained captions provide pseudo-supervision for the new branches.

  • Frozen global branch preserves standard-caption alignment
  • Attention-based motion extractors target each text granularity
  • Calibrated score fusion enables mixed-granularity ranking
Granularity-aware motion-text retrieval framework
View full-resolution PDF ↗
Key Findings

MRBench reveals what in-domain evaluation misses.

01

Generalization gap

Models strong on HumanML3D degrade consistently when evaluated across MRBench’s broader motion domains.

02

Granularity sensitivity

Concise and fine-grained queries expose distinct weaknesses hidden by standard descriptions.

03

Better mixed retrieval

Our lightweight adaptation improves mixed-granularity retrieval while retaining the standard-caption branch.

Supplementary Video

See MRBench in motion.

A visual introduction to why current retrieval benchmarks fall short, how MRBench is constructed, and how our model handles descriptions at different levels of detail.

Benchmark

Diverse motion, precise language

Explore 3,390 motions from four sources, 118 fine-grained categories, and 10,170 motion-verified captions at three granularities.

Construction

Reliable by design

Follow the filtering, taxonomy-guided sampling, semantic verification, visually grounded rewriting, and human validation pipeline.

Retrieval

Alignment across granularities

See how lightweight granularity-specific branches improve non-standard and mixed-granularity retrieval while preserving standard alignment.

Citation

Found MRBench useful?

Please cite our work. The paper is now available.

@misc{liu2026mrbenchcomprehensivebenchmarkhuman,
  title         = {MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval},
  author        = {Fulong Liu and Liang Xu and Chengqun Yang and Yuhao Zhang and Yichao Yan and Xiaokang Yang},
  year          = {2026},
  eprint        = {2608.07993},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2608.07993},
}