Heterogeneous motion
Four sources probe transfer across capture conditions and generation domains.
MRBench
A diverse, balanced, and multi-granular testbed for measuring robust motion-language alignment beyond narrow datasets and fixed caption styles.
1 Shanghai Jiao Tong University 2 Beijing Zhongguancun Academy * Equal contribution ‡ Project leader † Corresponding author
MRBench brings together motion capture, in-the-wild videos, synthetic videos, and generated motions—each paired with concise, standard, and fine-grained text.
Homogeneous motion sources, long-tailed categories, and repetitive captions can inflate in-domain performance while obscuring true cross-modal understanding.
Four sources probe transfer across capture conditions and generation domains.
A hierarchical taxonomy covers 7 coarse, 33 middle-level, and 118 fine-grained categories.
Every motion has visually grounded descriptions at three explicit semantic levels.
A multi-stage pipeline filters weak captions, balances semantic coverage, verifies cross-modal alignment, and produces motion-grounded descriptions.
Remove generic, repetitive, and non-kinematic descriptions.
Use a hierarchical taxonomy to enforce diverse category coverage.
Assess whether the described motion is observable and unambiguous.
Generate three granularities, followed by careful human validation.
We freeze a standard-caption-aligned dual encoder and add lightweight, granularity-specific motion extractors and text adapters. LLM-rewritten concise and fine-grained captions provide pseudo-supervision for the new branches.
Models strong on HumanML3D degrade consistently when evaluated across MRBench’s broader motion domains.
Concise and fine-grained queries expose distinct weaknesses hidden by standard descriptions.
Our lightweight adaptation improves mixed-granularity retrieval while retaining the standard-caption branch.
A visual introduction to why current retrieval benchmarks fall short, how MRBench is constructed, and how our model handles descriptions at different levels of detail.
Explore 3,390 motions from four sources, 118 fine-grained categories, and 10,170 motion-verified captions at three granularities.
Follow the filtering, taxonomy-guided sampling, semantic verification, visually grounded rewriting, and human validation pipeline.
See how lightweight granularity-specific branches improve non-standard and mixed-granularity retrieval while preserving standard alignment.
Please cite our work. The paper is now available.
@misc{liu2026mrbenchcomprehensivebenchmarkhuman,
title = {MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval},
author = {Fulong Liu and Liang Xu and Chengqun Yang and Yuhao Zhang and Yichao Yan and Xiaokang Yang},
year = {2026},
eprint = {2608.07993},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2608.07993},
}