Multimodal grounding

When an image is input, force the model to cite visible evidence before inventing scene details.

When to use

Vision Q&A, edit briefs, or any task that starts from a reference frame.

How to apply

  1. Require: (1) list visible facts, (2) state uncertainties, (3) only then propose edits or answers.

Glossary

grounding
Tying claims to visible evidence in an image before inventing scene details.
visible facts
What can actually be seen in the frame — listed before interpretation.
uncertainties
What the model cannot see or resolve — stated explicitly before editing or answering.