Architecture

Built on Cloudflare

Every piece of the platform maps to a Cloudflare primitive, with no external SaaS in the critical path. One Worker hosts the surfaces and the orchestration; a Container handles the work the Workers runtime cannot.

Lirovo · architecture / deploymentlirovo.airequestWorker (packages/api)REST APIMCP serverextraction Workflowdurable orchestrationASR · vision · reason · Pass AContainerffmpeg · scene · pHashD1metadataR2artifactsAI Gatewaymodel routingQueue + DO admissionplanned · ADR-010dispatchrowsobjectsmodel callsadmissiondata flowindirectentry point
One Hono Worker hosts the API, the MCP server, and the Workflow, and binds to the media Container, D1, R2, and AI Gateway. The Queue plus per-tenant Durable Object admission layer is designed (ADR-010) and not yet shipped.

Why one platform#

Running compute, storage, the database, and the model gateway on a single platform is a deliberate trade. It costs some best-of-breed flexibility, and it buys three things that matter more for this workload: data stays close to the compute that processes it (no cross-cloud egress for large media), the attack surface stays small (one identity and network boundary instead of a patchwork), and every request sits behind a global network with always-on network-layer protection before it reaches our code. Durable execution and the object store are native here, so the orchestration layer is a primitive rather than a service to run.

Primitives and their roles#

Workers
compute

The REST API and the MCP server run on one Hono Worker. ASR, vision, reasoning, and Pass A also execute in the Worker via the Workflow.

Workflows
orchestration

Each extraction is a durable, retryable, checkpointed Workflow. A crash resumes from the last successful step rather than re-running the whole job.

Containers
compute

Heavy media work (yt-dlp fetch, FFmpeg normalize and remux, scene detection, perceptual-hash dedup) runs in a stateless Container that reads and writes R2 by key.

D1
database

Relational metadata: tenants, API keys, jobs, schemas, destinations, the graph mirror (nodes and edges), and the evidence chain.

R2
object storage

All artifacts: source and normalized media, frames, transcript, vision descriptions, the canonical knowledge graph, and extraction results.

AI Gateway
model routing

The single choke point for model calls. Routing plus per-call logging of provider, model, latency, tokens, and cost.

Queues + Durable Objects
scaling (planned)

The planned scaling layer from ADR-010: per-tenant Durable Object admission feeding a Cloudflare Queue ahead of the Container pool. Designed, not yet shipped.

Why a separate Container for media#

The Workers runtime cannot run arbitrary binaries or sustained CPU-bound work, and the media stages are exactly that: shelling out to yt-dlp, running FFmpeg, and hashing frames. Those stages run in a Cloudflare Container instead. The Container is stateless: it reads inputs from R2 by key and writes outputs back by key, so it can be scaled, retried, or replaced without any coordination, while the durable job state stays in the Workflow, R2, and D1.

Scaling model#

The binding constraint under load is model throughput, not the Container pool: the provider rate limit on the reasoning and graph calls is reached well before compute runs out. The scaling design (ADR-010) is built around that fact rather than around raw concurrency.

It is three layers: per-tenant Durable Object admission decides whether a tenant may start another job, a Cloudflare Queue absorbs the burst, and the Container pool drains it. Per-plan concurrency caps keep one tenant from starving the rest. The admission and queue layers are designed and tracked; the per-job pipeline ships today.

Designed, not yet shipped
The Queue plus per-tenant Durable Object admission layer is specified in ADR-010 and not yet implemented. The per-job Workflow, Container, and storage path described elsewhere in this section is live.

Observability#

Every request gets a trace id at API ingress that propagates through the Workflow, the Container, and the result write. It is returned on every response as x-lirovo-request-id. Logs are structured (pino), and errors are captured in Sentry with the trace id attached, so a single id locates a request across the whole pipeline. Every model call is metered at the gateway with provider, model, latency, tokens, and cost.