Multimodal grounding
When an image is input, force the model to cite visible evidence before inventing scene details.
When to use
Vision Q&A, edit briefs, or any task that starts from a reference frame.
How to apply
Require: (1) list visible facts, (2) state uncertainties, (3) only then propose edits or answers.
Glossary
- grounding
- Tying claims to visible evidence in an image before inventing scene details.
- visible facts
- What can actually be seen in the frame — listed before interpretation.
- uncertainties
- What the model cannot see or resolve — stated explicitly before editing or answering.
