MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation

Ruiqi Li1*, Xuanyi Liu1*, Sijia Li2*, Haofeng Wang1, Yuxin Liu2, Feng Xie2, Songchao Tan2, Shiqi Wang3, Hanwei Zhu4†, Yizong Wang1†, Chuanmin Jia1, Siwei Ma1
1Peking University
2University of Science and Technology Beijing
3City University of Hong Kong
4Nanyang Technological University

*Equal Contribution   Corresponding Authors

Overview

Main figure

(Left) Evolution of Generation Paradigms: Unlike traditional Text-to-Video (explicit instruction alignment) or Video Reasoning (objective physical rules), our paradigm requires models to autonomously deduce behaviors from latent cognitive states. As illustrated, the model must prioritize the man's subjective belief over objective reality to generate the correct reasoning-driven action. (Right) MindWorldBench Taxonomy: Grounded in the BDI-P framework, we systematically categorize mental variables into three primary dimensions—Perception, Belief, and Desire—to comprehensively evaluate Theory of Mind capabilities in video generation.

Abstract

Current image-to-video models achieve visual realism and physical plausibility, but reasoning about latent mental states remains unexplored. Since actions are often driven by beliefs, desires, and perceptions rather than explicit instructions, we introduce MindWorldBench to evaluate mental-state-conditioned video generation. We formalize a mental-state-to-behavior reasoning task in which models must generate actions from world states and latent mental variables without action prompts. MindWorldBench adopts zero-action prompting with counterfactual prompt design to isolate the causal effects of mental states, and an automated evaluation pipeline measures video quality, commonsense plausibility, and mental-state consistency. Evaluations across eight models show that, despite strong visual fidelity, current systems fail to align behaviors with latent mental states, revealing an Omniscient Bias in which models default to objective world states rather than a human subject's beliefs. These results expose a gap between video generation and cognitive reasoning, and highlight the need for explicit mental-state modeling in future generative systems.

Dataset

Dataset show figure

Representative examples from MindWorldBench. The benchmark spans the three mental dimensions: Belief, Desire, and Perception, covering diverse subcategories. Each example includes an input image and a structured mental-state prompt, followed by a question probing expected actions, highlighting the need for models to infer behavior from latent mental states.

Evaluation Metrics

Evaluation metric explanation figure

Illustration of evaluation metrics in a typical task. The left panel shows the input scene and mental-state prompt. The middle column indicates Intention Accuracy, whether the model-generated action matches the agent's intended goal. The right column shows World-State Consistency, whether the post-action environment remains correct. Check marks denote correct predictions, and crosses denote errors.

Results

BibTeX

@inproceedings{li2026mindworldbench,
  title={MindWorldBench: Evaluating Mental-State-to-Behavior Reasoning in Image-to-Video Generation},
  author={Ruiqi Li and Xuanyi Liu and Sijia Li and Yuxin Liu and Feng Xie and Haofeng Wang and Songchao Tan and Shiqi Wang and Hanwei Zhu and Yizong Wang and Chuanmin Jia and Siwei Ma},
  year={2026}
}