Every prompt is an export
Each API call sends your prompt — and every retrieved document behind it — across your boundary to a third party. Contracts soften that; they don’t undo it.
Open-weight models reached frontier quality in 2026, and with them the old trade — capability in exchange for control — expired. We select, fine-tune and operate models that run entirely on infrastructure you control: your data never leaves your network, your costs don’t scale with someone else’s pricing page, and every answer can be audited.
For three years, using frontier AI meant sending your data to someone else’s model. That was the price of capability. Open-weight models closed the quality gap — which turns four quiet costs of the API route into decisions you get to revisit.
Each API call sends your prompt — and every retrieved document behind it — across your boundary to a third party. Contracts soften that; they don’t undo it.
Per-token pricing feels cheap in a pilot and compounds in production. Past a steady volume, owned capacity beats metered inference — and the meter never sleeps.
Hosted models get updated, deprecated and retired on someone else’s schedule. Behaviour you validated in January answers differently in June — mid audit cycle.
Regulators ask where the data went, and why the model said what it said. When inference happens in another jurisdiction on a model you can’t inspect, both answers get hard.
Everything that touches your data lives inside your perimeter — not as a compliance patch at the end, but as the constraint the whole system is engineered around from day one.
Inference, learning and evidence all stay inside. Nothing crosses out.
Three layers, one system. Each is built to be operated by your team, not rented back to you.
Selection across the open-weight families — licence-vetted for your use — then fine-tuned on your corpus with adapter methods, preference-tuned to your standards, and distilled where a smaller model earns its keep. Most enterprise workloads are won by a mid-size model tuned well, not the largest checkpoint available.
Quantisation chosen against your accuracy bar, high-throughput serving with continuous batching and KV-cache management, and hardware sized to the measured workload — from a single GPU server to a multi-node cluster. Capacity you own, latency you control.
An evaluation harness that gates every change before it ships, drift monitoring, rollback, observability and access control — the MLOps discipline that makes the system yours to run. Enablement for your team is part of the build, not an upsell.
“On-prem” is a spectrum, not a single architecture. We deploy against the constraint you actually have.
No external egress at all. Model weights and updates arrive by controlled import; everything else — inference, retrieval, evaluation — runs disconnected. For classified, clinical and legally privileged environments.
The same stack in your VPC on dedicated GPU capacity: your keys, your network rules, your audit trail — and no third-party model API in the data path. Sovereignty without a server room.
Sensitive data and retrieval stay on-prem; heavy, non-sensitive compute — training runs, batch jobs — bursts to isolated cloud capacity in your tenancy. The boundary follows the data, not the fashion.
Bespoke on-prem deployment is an investment that pays back under specific conditions. If several of these describe you, the economics and the risk case are usually decisive.
Patient records, lending decisions, legal privilege, defence-adjacent work — data that cannot cross to a third party, whatever the DPA says.
Document processing, classification, internal assistants running all day — workloads where owned capacity beats the meter within the first year.
UK GDPR and the EU AI Act on one side of the Atlantic; SEC and FINRA recordkeeping and HIPAA on the other — plus contracts that require you to say, precisely, where processing happens.
Audit cycles, validated processes and regulated decisions need a model that changes when you decide — with an evaluation record for every change.
Proprietary research, designs, deal flow — knowledge that is the business, and that should never train anyone else’s model.
And the honest counterpart: if your volumes are modest and your data is not sensitive, a hosted API is probably the right answer — and in an assessment, we will tell you so. The point is to make it a decision, not a default.
Less than the headlines suggest. A well-tuned mid-size model quantised for inference runs on a single modern GPU server; larger checkpoints scale across a small cluster. Frontier-scale open models exist and are deployable, but they demand terabyte-class GPU memory — and most enterprise workloads don’t need them. We size against your measured workload in the assessment, not against a vendor’s reference architecture.
On scoped enterprise tasks — with fine-tuning, retrieval and a proper evaluation harness — routinely, yes. As broad general-purpose assistants, the hosted frontier still leads in places. That’s why we benchmark candidate models on your tasks and your data before anything is committed, and show you the numbers.
“Open weights” is not one licence. Some checkpoints are MIT or Apache licensed and commercially clean; others carry acceptable-use restrictions or revenue thresholds. Licence vetting is part of model selection — we shortlist only models whose terms fit your use, and document why.
Your team — by design. You own the weights, the adapters, the code, the evaluation harness and the runbooks, and enablement is built into the engagement. We stay available for support where you want it; a dependency on us is a design failure, not a revenue line.
No. Training, retrieval and inference all run inside your boundary. Nothing is sent to a third-party API, and nothing you own trains anyone else’s model — that is the point of the architecture.
We take one real workload and answer the questions that decide this properly: which open-weight models clear your quality bar on your tasks, what hardware the measured workload actually needs, what it costs against your current API spend, and what your regulators will ask to see.
What you get: a benchmark of shortlisted models on your data; a hardware sizing and capacity plan; a total-cost comparison against the API route with the break-even point; a target architecture for your boundary — air-gapped, VPC or hybrid; and a staged delivery plan. Led by a senior consultant — fixed scope, fixed fee.
Book a Deployment Assessment →A 30-minute conversation with a senior consultant. Bring the workload you’d most like to take off a third-party API — or the one your regulator keeps asking about. We’ll tell you what an on-prem deployment would take, what it would cost, and whether it’s worth it for you.
Book a Deployment Assessment →