VDiff-Bench A Benchmark for Fine-Grained Image Difference Identification

Authors

Yixin Wan Tianle Zheng Kai-Wei Chang

University of California, Los Angeles

1,756questions
1,543image pairs
10change categories
11evaluated MLLMs

Qualitative examples

10 Challenging Visual Difference Categories

Benchmark at a glance

Can models identify fine-grained visual differences?

MLLMs perform strongly on general visual understanding yet often miss small changes between closely related images. VDiff-Bench tests this skill with 1,756 four-way questions across ten semantic, textual, and low-level categories. Each question pairs one verified difference with two hard alternatives and a no-difference distractor, enabling exact, judge-free scoring. Across 11 MLLMs, performance remains brittle, especially on noise, texture, and other low-level changes.

Qualitative comparison of VDiff-Bench with CLEVR-Change, Spot-the-Diff, and OmniDiff
Figure 1. VDiff-Bench compared with prior visual-difference benchmarks and representative model responses.

Why VDiff-Bench

Controlled visual verification

Prior benchmarks cover surveillance, rendered, and other paired images but rely on open-ended outputs scored by caption metrics or learned judges. VDiff-Bench spans real, edited, rendered, and 2D puzzle images, covering semantic, OCR/text, photometric, noise/resolution, and texture changes. Models must distinguish the observed change from plausible alternatives, enabling exact, judge-free scoring.

Coverage and evaluation design of visual difference identification benchmarks
Benchmark Pairs Image regime Target-change coverage Evaluation design
RealEditedRendered2D puzzle SemanticOCR/textGlobal photo.Noise/res.Texture Hard alternativesExact scoring
Spot-the-Diff13,192×××××××××
CLEVR-Change79,606××××××××
OmniDiff15,598××××××
DiffCap-Bench1,075×××
VDiff-Bench (ours)1,543

Checks mark explicitly covered dimensions. “Global photo.” includes whole-image color and illumination; “Noise/res.” includes noise and resolution degradation. Hard alternatives are plausible competing changes; exact scoring uses direct answer-key matching without caption metrics or learned judges.

Dataset statistics

Ten complementary change categories

Hover or focus a segment for its definition, group, and question count.

Semantic or textualLow-level

Model performance

Strong averages can hide perceptual gaps

Overall accuracy and low-level accuracy on whole-image color, noise, texture, and illumination changes.

Overall accuracy Low-level accuracy Guessing baselines

Category explorer

Compare models on a specific change

Category accuracyGuessing baselines

Full results

VDiff-Bench leaderboard

Model accuracy by VDiff-Bench change category

Accuracy (%). Uniform four-way guessing yields 25.0%; excluding the always-false no-difference option yields 33.3%.