Edge AI
An image the judge was told to ignore moved one label in five
- Author
- Dev SanghviFounder & CEO, DHI
- Published
- 2026-10-08
- Read time
- 6 min read
- Updated
- 2026-10-08
What the paper did
"It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them" (arXiv 2609.37863) was posted on September 29, 2026. It was written for the Trust-AI-Eval workshop at NeurIPS 2026, and its subject is the habit of using vision-language models in place of human annotators.
The authors built a test called MIST. It has 200 English sentences, each containing a phrase that can be read figuratively or literally, as "kicked the bucket" can. The judge has to label how the phrase is used, from the sentence alone. Each sentence is shown three ways: with an image that matches the reading, with a misleading image that depicts the opposite reading, and with no image. The instructions say to ignore the image. Thirteen models from seven families acted as judges.
What it found
An aligned image changed 20.5 percent of the labels. A misleading image changed 19.4 percent. That is about one label in five whichever image was shown, and the two numbers were close for every judge.
For comparison, deleting the instruction to ignore the image, with the image left in place, changed 11.6 percent. A plain-language instruction moved fewer labels than the picture did.
The part that is easy to miss: in 1,776 of 10,400 cases (one sentence, one prompt, one judge) the label differed between the two images. Of those, only 37 percent moved toward the sense the misleading image showed, and the rest moved away from it. The label does not follow what the picture shows. It moves because a picture is there.
Accuracy did not notice. Agreement with the human majority was 54.4 percent with no image, 53.7 percent with an aligned one and 54.7 percent with a misleading one. The changes cancel out in the total, so a score computed from overall agreement cannot see them.
Why a labelling test matters to a camera system
It is a different job from ours. In the paper the image is a distraction. In a verifier that looks at a frame and answers whether an alert is real, the image is the evidence.
What transfers is the method, and the method is an invariance test. Hold the right answer fixed. Change something that should not matter. Count how many answers flip, and report that number next to accuracy.
For a second-stage verifier, the things that should not matter could be the camera's name in the prompt, the order of the alert fields, the wording of the question, or an unrelated second frame. That list is ours, not the paper's. The paper is also bounded by its own limits: all items are English and drawn from one corpus of idioms. Nothing in it says anything about a safety deployment, and we are not claiming it does.
What we have and have not done
DHI runs an optional second-stage vision-language verifier as the last gate on an alert. It ships off, and when it cannot reach a judgment, the alert is kept rather than dropped.
In September we wrote about a paper on whether such models use the visual evidence at all, and said we had no published measurement of our own verifier's visual sensitivity. This is a second test with the same status. We have published no result from either on our verifier, and nothing above should be read as one.
When you evaluate a checker model, do you test what happens when the context changes and the right answer should not?
Sources
- Omar, Jabarin, Ibraheem, "It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them," arXiv 2609.37863, submitted September 29, 2026: arxiv.org/abs/2609.37863. Figures are from the abstract and the paper body (all thirteen judges; Appendix F).
- Edge AI
- VLM
- Evaluation
- Verification
Continue exploring.
- Edge AIAsk the model lessA new tracking paper reformulates a hard vision task as a yes or no question instead of free text generation. DHI's own alert verifier made the same bet, and the paper shows us the next step we have not taken.
- Edge AIWhen your verifier stops lookingA new paper finds vision-language models often answer without using the image evidence in front of them. That is a direct risk for any system, DHI's included, that uses a VLM as a verification gate.
- Edge AIOur thermal model's evaluation was wrong twice before it told us the truthA thermal perception backbone looked broken, then looked useless, then turned out to be neither. The bug was never in the model. It was in the test we used to judge it, and in a data pipeline that let RGB photos into a corpus we called thermal.
See what real edge AI looks like on your cameras.
Start with one camera that matters. We will run a 30-day live validation on the CCTV and VMS you already have, and you keep every frame on-premise.
Best follow-up: bring the single feed that keeps you up at night.
- Request a demoSee the flow on a real operating scenario and scope a pilot around one facility or corridor.
- See deployment architectureReview camera ingest, edge inference, alert routing, and what stays on-premises.
- Get the implementation checklistDownload the deployment checklist buyers use before green-lighting an industrial AI pilot.
- Talk to an engineerBring camera count, VMS constraints, latency expectations, and privacy requirements to a technical review.