VISIONXLAB · REASONING-INFORMED VISUAL EDITING

Visual reasoning,
measured.

How well can visual editors understand what comes next?
Explore, compare, and understand the RISE benchmark results.

Six connected dimensions of visual reasoningRISE++REASON · EDIT · EVALUATETemporalCausalSpatialLogicalCounterfactualHybridFrom implicit intent to precise visual change.
ENGLISH RESULTS · FULL BENCHMARK

The reasoning leaderboard

Instruction language
Columns
Click any score to sort · Select up to 3 models to compare

A CLOSER LOOK

Different strengths. Different reasoning.

Full-benchmark highlights
UNDERSTAND THE BENCHMARK

More than following an instruction.

Read the paper ↗

RISE asks editors to infer the intended outcome, then make the right visual change. Explore the reasoning capabilities behind each score.

THREE EVALUATION DIMENSIONS

Every applicable dimension matters.

01

Instruction Reasoning

Does the image realize the intended reasoning outcome?

02

Appearance Consistency

Does unrelated content remain consistent with the input?

03

Visual Plausibility

Is the image visually coherent? Logical tasks are excluded; intentionally impossible counterfactual outcomes are not penalized.

How is accuracy different from evaluation scores?

How should I compare single- and multi-image models?

What does the rank mean?

Ranks reflect the selected metric across all models in the current benchmark and language. Higher scores rank first; equal scores share a rank. Filtering preserves these ranks, and missing results appear last.

How can I evaluate a model or contribute results?

Download the full dataset, follow the evaluation instructions, and report the model version, input setting, judge configuration, and generated outputs. Use the repository to discuss or submit reproducible results.

BUILT FOR REPRODUCIBLE RESEARCH

Use RISE in your work.

Explore the data, inspect model outputs, and cite the benchmark.

Download leaderboard JSON ↓
MODEL COMPARISON