UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models

1University of California San Diego   2University of Southern California
3University of Illinois Urbana-Champaign   4Carnegie Mellon University
*Equal contribution.

Abstract

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent textual and visual modalities are aligned. To investigate this question, we use reasoning-guided image generation as a diagnostic task, where models produce textual reasoning first and then generate images.

We introduce UReason, a benchmark for evaluating reasoning-to-generation alignment in this paradigm, consisting of 2,000 human-curated and human-verified instances spanning five reasoning-intensive tasks: Code, Arithmetic, Spatial, Attribute and Text. To enable controlled analysis, we develop an evaluation framework that compares direct generation, reasoning-guided generation and de-contextualized generation, which conditions only on the refined prompt extracted from reasoning.

Across eight widely used open-source UMMs, while we find that reasoning-guided generation yields improvements over direct generation, somewhat surprisingly, de-contextualized generation consistently outperforms reasoning-guided generation by a large margin.

Our further analyses suggest that the intended visual semantics in textual reasoning are not reliably reflected in the generated images, despite their unified design and training. Overall, UReason serves as a practical litmus test for reasoning-to-generation alignment and provides a challenging benchmark for developing next-generation, more tightly aligned UMMs.

The UReason Benchmark

UReason shifts the paradigm from description to deduction. Unlike traditional text-to-image benchmarks that evaluate descriptive prompts with an emphasis on aesthetic fidelity, the target content here is not stated verbatim and must be inferred from the input scenario, requiring multi-step reasoning such as state tracking and distractor suppression. The benchmark consists of 2,000 human-curated and human-verified instances spanning five tasks with 30 fine-grained subcategories:

  • Code Reasoning: Interpreting and executing code (e.g., HTML, Python) to render visual outputs.
  • Arithmetic Reasoning: Tracking object quantities through mathematical reasoning.
  • Spatial Reasoning: Inferring complex layouts from implicit spatial cues and logical constraints.
  • Attribute Reasoning: Tracking state transitions to determine final object properties.
  • Text Reasoning: Deriving text strings via logical rules rather than direct quotation.

Each instance is paired with an instance-specific verifiable criterion (e.g., exact counts, specific spatial arrangements) that enables objective, scalable evaluation.

Curation. Human experts first establish a fine-grained taxonomy over the 5 tasks and manually construct 500 validated seed instances. We then scale coverage through human-guided, LLM-assisted augmentation that systematically varies the number of reasoning steps, target visual entities and narrative context, with multi-round human verification, expanding UReason to 2,000 instances.

Splits. The full test set contains all 2,000 instances; testmini is a 500-instance subset (100 per task) for rapid validation during model development. Results reported below are on testmini; the full test set shows consistent trends.

UReason Tasks.

Fig 1. Representative UReason instances covering Code, Arithmetic, Spatial, Attribute, and Text reasoning.

Evaluation Framework

To rigorously diagnose the impact of reasoning on image generation, we introduce the UReason Evaluation Toolkit. This framework implements a controlled ablation protocol to isolate the effectiveness of reasoning from potential interference. We evaluate models across three distinct settings:

  1. Direct Generation: The baseline setting where the model generates images directly from the original prompt.
  2. Reasoning-Guided Generation: The model generates a Chain-of-Thought (CoT) reasoning trace first, and then generates the image conditioned on the full context (Prompt + Reasoning).
  3. De-contextualized Generation: The model performs reasoning to derive a refined prompt, but the intermediate thoughts are discarded. The image is generated conditioning only on the refined prompt.

Settings 2 and 3 encode the same model-produced visual semantics in different contextual formats, so in principle they should perform comparably. Any gap between them isolates the reasoning-to-generation alignment gap: the model derives the right visual intent in text but fails to carry it through to pixels. Generated images are scored against each instance's ground-truth criterion by Qwen3-VL-235B-A22B as an automated verifier.

Evaluation Framework.

Fig 2. Overview of the UReason evaluation framework comparing Direct, Reasoning-Guided, and De-contextualized settings.

Main Results

We evaluate 8 open-source UMMs on UReason across all 3 evaluation settings. The UReason leaderboard reports Visual Verification Accuracy (%) on testmini and Performance Gain (Δ) over the previous setting.

Key Findings

  • Direct prompting performs poorly on implicit targets. Overall accuracy ranges from 4.2% to 8.2%. The target content is intentionally implicit, so direct text-to-image mapping is fundamentally insufficient to solve UReason.
  • Chain-of-thought reasoning helps. Every one of the eight models improves from Direct to Reasoning-Guided Generation, with gains from +0.4% (T2I-R1) to +19.2% (UniCoT-v2).
  • De-contextualized generation helps far more. Conditioning only on the refined prompt beats Reasoning-Guided Generation for every model, reaching +44.8% for Bagel and +43.2% for SRUM — even though both settings encode the same model-produced visual semantics. The same ordering holds on the full 2,000-instance test set (+8.6 to +44.4 points).
  • The bottleneck is not reasoning quality. Judged against the ground-truth criteria, reasoning chains are largely correct (Bagel 93.4%, SRUM 91.6%). Error analysis attributes 75.4% of Bagel's failures to task-specific execution errors despite correct reasoning.
  • Contextual interference is a contributing factor. An attention analysis on Bagel shows the intermediate reasoning trace keeps drawing more than half the attention paid to the refined prompt, competing with the final visual specification. A length-controlled ablation rules out sequence length alone: repeating the refined prompt 8× costs at most 6.0 points, far short of the 44.8-point gap.

Citation

@article{yang2026ureason,
  title   = {UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models},
  author  = {Yang, Cheng and Shi, Chufan and Shui, Bo and Wu, Yaokang and Tao, Muzi and
             Wang, Huijuan and Lee, Ivan Yee and Liu, Yong and Ma, Xuezhe and
             Berg-Kirkpatrick, Taylor},
  journal = {arXiv preprint arXiv:2602.08336},
  year    = {2026}
}

📬 Contact Us

If you have any inquiries about UReason, feel free to reach out to us at ureason2026@gmail.com, chy085@ucsd.edu, chufansh@usc.edu.