Visual reasoning,
measured.
How well can visual editors understand what comes next?
Explore, compare, and understand the RISE benchmark results.
The reasoning leaderboard
No matching models
Try another model name or clear the filters.
Different strengths. Different reasoning.
More than following an instruction.
RISE asks editors to infer the intended outcome, then make the right visual change. Explore the reasoning capabilities behind each score.
Every applicable dimension matters.
Instruction Reasoning
Does the image realize the intended reasoning outcome?
Appearance Consistency
Does unrelated content remain consistent with the input?
Visual Plausibility
Is the image visually coherent? Logical tasks are excluded; intentionally impossible counterfactual outcomes are not penalized.
How is accuracy different from evaluation scores?
How should I compare single- and multi-image models?
What does the rank mean?
Ranks reflect the selected metric across all models in the current benchmark and language. Higher scores rank first; equal scores share a rank. Filtering preserves these ranks, and missing results appear last.
How can I evaluate a model or contribute results?
Download the full dataset, follow the evaluation instructions, and report the model version, input setting, judge configuration, and generated outputs. Use the repository to discuss or submit reproducible results.
Use RISE in your work.
Explore the data, inspect model outputs, and cite the benchmark.