VDiff-Bench A Benchmark for Fine-Grained Image Difference Identification
Qualitative examples
10 Challenging Visual Difference Categories
Benchmark at a glance
Can models identify fine-grained visual differences?
MLLMs perform strongly on general visual understanding yet often miss small changes between closely related images. VDiff-Bench tests this skill with 1,756 four-way questions across ten semantic, textual, and low-level categories. Each question pairs one verified difference with two hard alternatives and a no-difference distractor, enabling exact, judge-free scoring. Across 11 MLLMs, performance remains brittle, especially on noise, texture, and other low-level changes.
Why VDiff-Bench
Controlled visual verification
Prior benchmarks cover surveillance, rendered, and other paired images but rely on open-ended outputs scored by caption metrics or learned judges. VDiff-Bench spans real, edited, rendered, and 2D puzzle images, covering semantic, OCR/text, photometric, noise/resolution, and texture changes. Models must distinguish the observed change from plausible alternatives, enabling exact, judge-free scoring.
| Benchmark | Pairs | Image regime | Target-change coverage | Evaluation design | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Real | Edited | Rendered | 2D puzzle | Semantic | OCR/text | Global photo. | Noise/res. | Texture | Hard alternatives | Exact scoring | ||
| Spot-the-Diff | 13,192 | ✓ | × | × | × | ✓ | × | × | × | × | × | × |
| CLEVR-Change | 79,606 | × | × | ✓ | × | ✓ | × | × | × | ✓ | × | × |
| OmniDiff | 15,598 | ✓ | × | ✓ | × | ✓ | ✓ | ✓ | × | × | × | × |
| DiffCap-Bench | 1,075 | ✓ | ✓ | ✓ | × | ✓ | ✓ | ✓ | ✓ | ✓ | × | × |
| VDiff-Bench (ours) | 1,543 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
Checks mark explicitly covered dimensions. “Global photo.” includes whole-image color and illumination; “Noise/res.” includes noise and resolution degradation. Hard alternatives are plausible competing changes; exact scoring uses direct answer-key matching without caption metrics or learned judges.
Dataset statistics
Ten complementary change categories
Hover or focus a segment for its definition, group, and question count.
Model performance
Strong averages can hide perceptual gaps
Overall accuracy and low-level accuracy on whole-image color, noise, texture, and illumination changes.
Category explorer
Compare models on a specific change
Full results
VDiff-Bench leaderboard
Accuracy (%). Uniform four-way guessing yields 25.0%; excluding the always-false no-difference option yields 33.3%.