§
— Reference Build

The system we run
on ourselves.

Before we route anyone else’s AI spend, we ran the architecture on our own operations. This is the teardown: what it is, what it runs, and why the cost-routing holds up.

01
— The architecture

Local carries the load. Cloud earns its keep.

The fleet is hybrid by design. Open-weight models running on owned hardware handle the bulk of the day-to-day work. Cloud frontier models are never the default — they’re routed in through a broker only for the tasks that justify their cost. A model-tiering policy decides which model each job gets, guard checks audit what the agents do, and a memory layer keeps context across sessions.

Local-first bulk
Open-weight models on owned hardware
Brokered escalation
Cloud frontier only where it pays
Model tiering
Right-sized model per task
Guard checks
Automated auditing of every run
Memory
Context that persists across sessions
02
— What it runs, every day

Real operations, 24/7.

01
Client CRM
Tracks relationships, follow-ups, and the day-to-day messaging that keeps client work moving.
02
Catalog sync
Keeps the product catalog for an e-commerce watch store current, without manual re-entry.
03
Monitoring & health checks
Watches the systems that run the business and flags trouble before it spreads.
04
Weekly reporting
Compiles the recurring reports that used to eat hours — on a schedule, unattended.
03
— Why cost-routing works

The router is plumbing. The gate is the product.

Routing work to cheaper models only saves money if quality holds. The part that makes it safe isn’t the router — it’s the evaluation gate built from your real tasks, so you can see whether a smaller model actually did the job before you trust it with the job.

That’s the whole discipline: measure the traffic, classify it, prove the cheaper path on your own work, then route. Escalate to a frontier model only where the eval says you need it. It’s the same cost-routing architecture the industry began marketing heavily in 2026 — the difference is that it has been running here in production, not sitting in a slide.

A founder who runs the architecture he sells is the whole point.

§
— Next step

Want this measured for you?

We measure and classify your AI traffic, build an evaluation gate from your real tasks, and show exactly what right-sized routing would save — quality proven before and after. The map is yours either way.