Skip to content
All articles

Multimodal AI

Multimodal AI starts with evidence, not modality

Sending every photo, recording and document to the largest model is the common multimodal mistake. Start from the decision, treat each signal as evidence with an origin, and let policy decide.

· 3 min read

Pipeline: sources (camera, mic, file, sensor) to observations, to alignment by source, page and time, to a decision: answer, abstain or review. Example: should we accept a shipment, using a photo of a dented carton, a cold-chain temperature reading, a driver's note about hard braking and the purchase order. An evidence ledger records capture time, transform, confidence, policy, source and digest for every signal.

A receiving clerk has to decide whether a refrigerated shipment is safe to accept. Four signals are available:

  • a photograph showing a dented carton and liquid on the floor,
  • a temperature sensor reporting the cold-chain reading,
  • a voice note from the driver describing a hard braking event,
  • the purchase order, which defines the acceptable quantity and packaging.

None of these is the truth by itself. Each is evidence: produced by a source, captured at a time, transformed by a pipeline, and interpreted under a policy.

The common mistake

The tempting design sends every available file to the largest model and asks for a verdict. More media does not make that verdict better. It adds irrelevant signal, conflicting versions of the same document, data the system was not allowed to use, content crafted to manipulate the model, latency and cost.

Write the decision first

Before choosing a single model, write down:

  1. the decision to be made,
  2. the observations it requires,
  3. the sources that are permitted,
  4. the uncertainty that is acceptable,
  5. what authority the output has: a final answer, or a recommendation for a person.

Only then select the modalities and the model capabilities that can supply those observations.

Four stages and a ledger

Sources such as cameras, microphones, files and sensors produce raw evidence. Observations extract regions, words, values and events from it without losing track of where each came from. Alignment joins the observations by source, position on a page, location and time. A bounded decision then answers, abstains, or routes the case to human review, under a policy that behaves the same way every time.

Running underneath, an evidence ledger records for every signal when it was captured, how it was transformed, how confident the extraction was, which policy applied, and a digest of the original.

Mental Model

A modality is a channel. Evidence is an observation that supports a claim, with an origin and known limitations. A product should reason over packages of evidence, not over anonymous blobs of media.

In the shipment example, that means the system can say not just "reject" but which signals led there, when each was recorded, and what a reviewer should look at first.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy