Overview
Lirovo turns a video into typed JSON that conforms to a schema you define, with every value linked back to the source moment that produced it. The whole platform runs on Cloudflare, multi-tenant from the first row, with evidence as an invariant.
What the system has to do#
A caller hands Lirovo a video and a JSON Schema and expects three things back: JSON that actually conforms to the schema, a citation for every value (the timestamp and modality it came from), and a price that makes sense for a long video. The hard parts are that videos are large and slow to process, the work is bursty, model providers fail independently, and the same minute of footage may need to be re-read against a different schema later.
The architecture below is the answer to those forces: durable orchestration so a long job survives restarts, a separate runtime for the heavy media work, a single gateway in front of every model, and an intermediate graph that is cheap to reason over and reusable across schemas.
The high-level shape#
The platform is a single path through five layers, shown in the diagram above:
- Surfaces. A REST API and an MCP server, both served by one Hono Worker (
packages/api). The same handlers sit under both surfaces, so an SDK call and an agent tool call run identical code. - Orchestration. Each extraction is a Cloudflare Workflow: a durable, retryable, checkpointed sequence of steps that resumes from the last successful stage rather than restarting the job.
- Compute. Heavy media work runs in a Cloudflare Container (
packages/engine). Transcription, vision, reasoning, and the graph build run in the Worker and reach models through AI Gateway. - Storage. Artifacts (media, frames, transcript, vision, graph, results) live in R2. Relational metadata (tenants, jobs, schemas, destinations, the graph mirror, the evidence chain) lives in D1.
- Delivery. On completion the result is optionally pushed to a destination, such as an HMAC-signed webhook, in its own retrying Workflow.
Design principles#
Five decisions shape everything else. Each page below expands on one of them.
One platform, no SaaS in the critical path
Compute, storage, the database, and the model gateway are all Cloudflare. That keeps data close to where it is processed, keeps the attack surface small, and puts every request behind a global network with always-on network-layer protection before it reaches our code.
Durable orchestration over hand-rolled state
A 90-minute video is a long-running, multi-step job. Modelling it as a Workflow gives checkpointing, automatic retries, and resume-from-failure for free, instead of tracking step state by hand in a queue.
Evidence is architectural, not a feature
Every extracted value points back to a source span and modality. This is built into the graph the reasoning step reads, so a citation cannot be bolted on or lost.
A provider-agnostic model layer
Models are reached only through AI Gateway, behind a stable interface with a per-request override. Swapping a model or adding a fallback is a configuration change, not a rewrite.
Explore the architecture#
Each layer and cross-cutting concern has its own page.
The ordered stages of a job, the two-runtime boundary, the job lifecycle, retries, and latency.
Each primitive mapped to its role, why a Container, the scaling model, and observability.
The provider-agnostic stack, AI Gateway as the choke point, fallback, and per-request overrides.
Spans and modality, the temporal knowledge graph, and why reasoning reads the graph.
Tenant isolation, admission guards, circuit breakers, idempotency, signed webhooks, backups.