Multimodal Reasoning: Keeping the Answer Connected to the Image
ICoT reinserts selected image evidence while a model reasons. Explore the actual examples, modest benchmark gains and a practical evaluation path.
7
Select a figure to open it at full size.
A vision-language model can select the correct answer while giving an explanation that misdescribes the picture. That matters when a product uses the explanation to justify an inspection, answer a customer or guide a later action. A plausible paragraph is not evidence that the model kept looking at the relevant details.
Interleaved-Modal Chain-of-Thought, or ICoT, investigates a specific response to this problem: reintroduce selected image information as the model generates its explanation. Jun Gao and colleagues test this on two open vision-language models. Their contribution is a change to the generation process, with modest benchmark improvements and instructive examples of where ordinary text-only reasoning goes wrong.
What the researchers change
A conventional multimodal reasoning prompt provides an image and asks the model to work through the answer in words. Subsequent steps are textual, even when the question depends on a small visual detail. ICoT adds selected visual representations between those steps.
The system first retains the image’s internal patch representations. As the model generates a response, a designated signal token—such as a newline—triggers a selection operation. The researchers inspect how strongly that token attends to different image patches, average the attention information across layers, and select a fixed number of highly attended patches.
Before putting those patches back into the sequence, the system restores their relative spatial order. This preserves information about how the selected pieces were arranged in the original image. Generation then continues with both the written explanation and this refreshed visual input. The loop can repeat at another signal token.
Gao et al., Figure 2. Attention selects which existing image representations are reinserted during generation; it does not independently verify that those details are correct. Original paper. Select the image for full size.
This is more than adding “look carefully” to a prompt. It needs access to the model’s attention information and generation internals. The paper avoids additional model training, but the runtime still changes and carries memory and computation overhead. A hosted model API that returns only text does not automatically expose what this implementation requires.
A correct answer can hide a broken explanation
The first row of the researchers’ original case study asks what an inflatable castle, crayons and a parachute have in common. The correct choice is that they are colorful. The text-only response gets that choice right while incorrectly describing all three objects as colored pencils.
ICoT’s illustrated response instead identifies the different objects, revisits selected image regions and connects their shared colorfulness to the answer. The practical improvement in this example is not the multiple-choice result—it was already correct—but whether the explanation refers to the actual image.
Gao et al., Figure 3. The selected examples show different failure modes; the first also demonstrates why answer accuracy alone can miss an incorrect explanation. Original paper. Select the image for full size.
The other rows make different mistakes visible. A person flying a kite becomes an unsupported claim that the scene is a kite festival. A troll statue becomes a story about honoring a local legend. In the interleaved examples, the model attends to the kites or street sign before reaching the benchmark answer.
These are selected illustrations rather than a count of how frequently the method prevents such errors. They also do not make attention maps proof of faithful reasoning. A model can attend to a relevant region and still interpret it incorrectly. What the examples provide is a better evaluation question: does each substantive claim in the explanation have support in the visible evidence?
What the benchmark results establish
The researchers test Chameleon-7B and Qwen2-VL-7B-Instruct on visual reasoning questions, science questions and image-based instruction responses. They compare direct answers, ordinary textual reasoning and ICoT, both without a worked demonstration and with one demonstration.
In the one-demonstration Chameleon results, ICoT scores 32.3 on the multimodal reasoning benchmark compared with 31.4 for the next-best listed method. On the science questions, the comparison is 53.4 versus 51.3. For Qwen2-VL, the corresponding gains are smaller: 46.0 versus 45.7 and 65.4 versus 64.9. These are useful improvements to investigate, rather than evidence that multimodal reasoning is now broadly reliable. The main results table also shows that asking for a textual chain of reasoning does not consistently beat answering directly.
The paper’s larger relative improvement headline comes from a text-overlap metric on generated responses. That metric rewards similarity to reference wording; it is not a direct measure of factual accuracy or successful business decisions. A product team should not translate it into a percentage improvement in trustworthy answers.
Removing the attention-driven selection component weakens the reported results, supporting the idea that selecting relevant patches contributes to the method. But the exact selection size matters. The supplementary experiments vary the number of patches and show that adding more is not uniformly better: too few can omit needed information, while too many add irrelevant material and computation.
The implementation tradeoff
ICoT depends on when visual evidence is reintroduced as well as what is selected. A frequent signal token can trigger too many insertions and reduce response quality. The one-demonstration setup also uses a manually prepared example with relevant visual details, so it should not be mistaken for a fully automatic demonstration-building pipeline.
The authors explore copying cached visual information to reduce repeated processing. That variation changes results slightly, and the paper does not provide an end-to-end latency comparison sufficient to promise that ICoT is faster or cheaper. Measure those costs directly if the application has a response-time budget.
For a screenshot assistant, a sensible experiment would compare direct answering, ordinary reasoning and visual reinsertion on the same questions. Include cases where a correct answer can be reached for the wrong reason, small text must be read, and the image does not support any confident conclusion. Score the answer and its evidence separately. This proposed evaluation follows the paper’s failure examples; the paper itself does not validate a deployed screenshot product.
Implementation Frameworks
The authors’ ICoT repository is the appropriate starting point for reproducing the model-specific generation changes. Check the supported model and dependency setup before adapting it. The essential capability is access to attention and visual representations during decoding, not simply support for image inputs.
Begin with a fixed collection of images and questions. Save outputs from the direct-answer baseline, then add the published method with its matching settings. Record the selected visual evidence, unsupported explanatory claims, answer correctness, latency and peak memory. Change the signal frequency or patch count on a development set and retain untouched cases for the final comparison.
If the existing API does not allow these internal changes, an explicit crop-and-reinspect workflow may be worth a separate experiment. It is a different technique and should be labeled and tested as such, rather than described as a reproduction of ICoT. Sometimes a direct answer with a quoted visual detail will be more useful than a long generated explanation.
The important product lesson is to evaluate the connection between an answer and the evidence offered for it. A correct final label can conceal a description that would mislead the next person or system in the workflow.
ICoT offers a concrete mechanism for strengthening that connection. Its value should be judged against simpler baselines on the errors that matter to the application, including the additional runtime cost. Longer reasoning is not the objective; better-supported reasoning is.
Original Research
Interleaved-Modal Chain-of-Thought, Jun Gao, Yongqi Li, Ziqiang Cao and Wenjie Li. Version 2, March 17, 2025. Original Figures 2 and 3 reproduced above.