Back to Blog

Edge AI

An image the judge was told to ignore moved one label in five

Author
Dev SanghviFounder & CEO, DHI
Published
2026-10-08
Read time
6 min read
Updated
2026-10-08

What the paper did

"It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them" (arXiv 2609.37863) was posted on September 29, 2026. It was written for the Trust-AI-Eval workshop at NeurIPS 2026, and its subject is the habit of using vision-language models in place of human annotators.

The authors built a test called MIST. It has 200 English sentences, each containing a phrase that can be read figuratively or literally, as "kicked the bucket" can. The judge has to label how the phrase is used, from the sentence alone. Each sentence is shown three ways: with an image that matches the reading, with a misleading image that depicts the opposite reading, and with no image. The instructions say to ignore the image. Thirteen models from seven families acted as judges.

What it found

An aligned image changed 20.5 percent of the labels. A misleading image changed 19.4 percent. That is about one label in five whichever image was shown, and the two numbers were close for every judge.

For comparison, deleting the instruction to ignore the image, with the image left in place, changed 11.6 percent. A plain-language instruction moved fewer labels than the picture did.

The part that is easy to miss: in 1,776 of 10,400 cases (one sentence, one prompt, one judge) the label differed between the two images. Of those, only 37 percent moved toward the sense the misleading image showed, and the rest moved away from it. The label does not follow what the picture shows. It moves because a picture is there.

Accuracy did not notice. Agreement with the human majority was 54.4 percent with no image, 53.7 percent with an aligned one and 54.7 percent with a misleading one. The changes cancel out in the total, so a score computed from overall agreement cannot see them.

Why a labelling test matters to a camera system

It is a different job from ours. In the paper the image is a distraction. In a verifier that looks at a frame and answers whether an alert is real, the image is the evidence.

What transfers is the method, and the method is an invariance test. Hold the right answer fixed. Change something that should not matter. Count how many answers flip, and report that number next to accuracy.

For a second-stage verifier, the things that should not matter could be the camera's name in the prompt, the order of the alert fields, the wording of the question, or an unrelated second frame. That list is ours, not the paper's. The paper is also bounded by its own limits: all items are English and drawn from one corpus of idioms. Nothing in it says anything about a safety deployment, and we are not claiming it does.

What we have and have not done

DHI runs an optional second-stage vision-language verifier as the last gate on an alert. It ships off, and when it cannot reach a judgment, the alert is kept rather than dropped.

In September we wrote about a paper on whether such models use the visual evidence at all, and said we had no published measurement of our own verifier's visual sensitivity. This is a second test with the same status. We have published no result from either on our verifier, and nothing above should be read as one.

When you evaluate a checker model, do you test what happens when the context changes and the right answer should not?

Sources

  • Omar, Jabarin, Ibraheem, "It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them," arXiv 2609.37863, submitted September 29, 2026: arxiv.org/abs/2609.37863. Figures are from the abstract and the paper body (all thirteen judges; Appendix F).
  • Edge AI
  • VLM
  • Evaluation
  • Verification