Architecture

Overview

Lirovo turns a video into typed JSON that conforms to a schema you define, with every value linked back to the source moment that produced it. The whole platform runs on Cloudflare, multi-tenant from the first row, with evidence as an invariant.

Lirovo · architecture / data flowlirovo.aisource videosurfacesREST APIdevelopers · SDKMCP serverAI agentssubmit jobextraction engineWorkflowdurable stepsContainerffmpeg · framesAI Gatewaymodel callsPass Aknowledge graph10-stage pipeline · checkpointed · idempotentstorageR2artifactsD1metadatawebhookHMAC signedartifactsmetadatadeliverdata flowindirectentry point
Source video enters at a surface, runs through one durable engine, and lands as artifacts plus metadata, with optional delivery.

What the system has to do#

A caller hands Lirovo a video and a JSON Schema and expects three things back: JSON that actually conforms to the schema, a citation for every value (the timestamp and modality it came from), and a price that makes sense for a long video. The hard parts are that videos are large and slow to process, the work is bursty, model providers fail independently, and the same minute of footage may need to be re-read against a different schema later.

The architecture below is the answer to those forces: durable orchestration so a long job survives restarts, a separate runtime for the heavy media work, a single gateway in front of every model, and an intermediate graph that is cheap to reason over and reusable across schemas.

The high-level shape#

The platform is a single path through five layers, shown in the diagram above:

  • Surfaces. A REST API and an MCP server, both served by one Hono Worker (packages/api). The same handlers sit under both surfaces, so an SDK call and an agent tool call run identical code.
  • Orchestration. Each extraction is a Cloudflare Workflow: a durable, retryable, checkpointed sequence of steps that resumes from the last successful stage rather than restarting the job.
  • Compute. Heavy media work runs in a Cloudflare Container (packages/engine). Transcription, vision, reasoning, and the graph build run in the Worker and reach models through AI Gateway.
  • Storage. Artifacts (media, frames, transcript, vision, graph, results) live in R2. Relational metadata (tenants, jobs, schemas, destinations, the graph mirror, the evidence chain) lives in D1.
  • Delivery. On completion the result is optionally pushed to a destination, such as an HMAC-signed webhook, in its own retrying Workflow.

Design principles#

Five decisions shape everything else. Each page below expands on one of them.

One platform, no SaaS in the critical path

Compute, storage, the database, and the model gateway are all Cloudflare. That keeps data close to where it is processed, keeps the attack surface small, and puts every request behind a global network with always-on network-layer protection before it reaches our code.

Durable orchestration over hand-rolled state

A 90-minute video is a long-running, multi-step job. Modelling it as a Workflow gives checkpointing, automatic retries, and resume-from-failure for free, instead of tracking step state by hand in a queue.

Evidence is architectural, not a feature

Every extracted value points back to a source span and modality. This is built into the graph the reasoning step reads, so a citation cannot be bolted on or lost.

A provider-agnostic model layer

Models are reached only through AI Gateway, behind a stable interface with a per-request override. Swapping a model or adding a fallback is a configuration change, not a rewrite.

Multi-tenancy is an invariant
Every artifact, row, and queue entry is scoped to a tenant, and no query joins across tenants. Tenancy is denormalized into even leaf tables. It is part of the data model, not a flag layered on top.

Explore the architecture#

Each layer and cross-cutting concern has its own page.