PHYRE: A New Benchmark for Physical Reasoning

Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, Ross Girshick

2019-08-15NeurIPS 2019 12Visual Reasoning

Abstract

Understanding and reasoning about physics is an important ability of intelligent agents. We develop the PHYRE benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment. The benchmark is designed to encourage the development of learning algorithms that are sample-efficient and generalize well across puzzles. We test several modern learning algorithms on PHYRE and find that these algorithms fall short in solving the puzzles efficiently. We expect that PHYRE will encourage the development of novel sample-efficient agents that learn efficient but useful models of physics. For code and to play PHYRE for yourself, please visit https://player.phyre.ai.

Results

Task	Dataset	Metric	Value	Model
Visual Reasoning	PHYRE-1B-Within	AUCCESS	77.6	DQN
Visual Reasoning	PHYRE-1B-Cross	AUCCESS	36.8	DQN

Related Papers

LaViPlan : Language-Guided Visual Path Planning with RLVR2025-07-17 Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning2025-07-15 PyVision: Agentic Vision with Dynamic Tooling2025-07-10 Orchestrator-Agent Trust: A Modular Agentic AI Visual Classification System with Trust-Aware Orchestration and RAG-Based Reasoning2025-07-09 MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning2025-07-09 Skywork-R1V3 Technical Report2025-07-08 High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning2025-07-08 Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning2025-07-07