Architecture

Evidence and the knowledge graph

Every extracted value points back to the moment in the source that produced it. Evidence is an architectural invariant: it is carried by the graph the reasoning step reads, so a citation cannot be lost or bolted on after the fact.

Lirovo · architecture / evidence + graphlirovo.aisource momentstranscript spant=00:12:04 · audiovision spant=00:03:20 · visualknowledge graphnodeclaimnodeentitynodeeventresult JSONconforms to schemaD1 mirrornodes · edges · evidencespanspanreasonmirrordata flow
Pass A builds a temporal graph whose nodes carry a span (timestamp + modality). Reasoning carries those pointers into the result. The graph is canonical in R2 and mirrored into D1.

Evidence is mandatory#

Every extracted value points back to a source moment, identified by a timestamp and a modality, either audio or visual. This is not a post-processing pass. The pointer lives on the graph node that produces the value, and the reasoning step carries it through into the final JSON, so the chain from a field in the result back to a second of footage is never broken.

The temporal knowledge graph#

Pass A fuses the transcript and the per-frame vision descriptions into a compact temporal knowledge graph: nodes (claims, entities, events) and the edges between them. Each node carries a span (timestamp) and a modality. The Workflow writes two views to R2, the canonical graph and a compact view, and mirrors the nodes, edges, and evidence into D1 so the graph is queryable with SQL.

Spans are always populated
When a model omits a timestamp on a node that has supporting evidence, the engine backfills it deterministically from the evidence span before writing, so every spanful node carries a timestamp in both the canonical and the compact graph.

Why reasoning reads the graph#

The reasoning step reads the compact graph, not the raw transcript. That is a deliberate choice with three payoffs. The compact graph is a fraction of the size of the transcript plus vision text, so it is cheaper to reason over and fits the context window comfortably. It is already structured and de-duplicated, so the model reasons over claims rather than re-reading prose, which is more accurate. And because the graph is stored, re-running the same video against a different schema reuses the graph and only re-runs reasoning, instead of paying for the whole pipeline again.

Retrievable artifacts#

Every intermediate a job produces is retrievable, so you can inspect exactly which moment produced any value:

  • Transcript. The time-aligned ASR output.
  • Frames. The deduplicated frame inventory.
  • Vision. The per-frame vision descriptions.
  • Knowledge graph. The nodes and edges with evidence-linked spans.

The REST API and the MCP server both expose these, including the evidence chain at GET /v1/jobs/{id}/evidence and the graph at GET /v1/jobs/{id}/graph.