Insights / Private AI
What it takes to run your own AI
Open-weight models have closed much of the gap. Here’s how to decide whether running AI on your own hardware makes sense for you.
Nothing leaves, nothing gets in
A few years ago, running a useful language model yourself meant a research budget. Today, open-weight model families from Meta, Mistral, Alibaba, Google, DeepSeek and OpenAI can be downloaded and run on hardware you own. OpenAI says its gpt-oss-120b model runs on a single 80 GB GPU, and the smaller gpt-oss-20b on machines with 16 GB of memory.
That doesn’t mean every business should do it. It means the question has changed from “can we?” to “is it worth it for this work?”
When it makes sense
- Your data can’t leave. Client-privileged documents, patient records, unreleased financials, source code under contract.
- The work is repetitive and well-defined. Summarizing, drafting from templates, extracting fields, searching your own documents. Mid-sized open models do this well.
- Usage is steady and high. Per-request fees add up; owned hardware is a fixed cost.
- You need it to keep working. Offline sites, or work that can’t stop when a vendor has an outage.
The question has changed from “can we?” to “is it worth it for this work?”
When it doesn’t
If you need the very best reasoning on hard, open-ended problems, frontier models from the large labs still lead. If your usage is light, a subscription is cheaper than hardware. And if nobody on your team can look after a server, factor in managed support.
What the stack looks like
Most private setups have four parts: the hardware (a workstation or a GPU server), a model server that loads the model and answers requests, a search layer that lets the model answer from your documents while respecting who can see what, and the interface your staff use. The part that decides success is rarely the model. It’s whether the search layer finds the right documents and whether permissions match the ones you already have.
- Interfacewhere your staff ask
- Search layerfinds the right documents, keeps your permissions
- Model serverloads the model and answers
- Hardwarea workstation or GPU server you own
Test before you buy
Public leaderboards won’t tell you how a model does on your contracts or your tickets. Build a small test set from real work, run candidate models against it, and compare cost per useful answer. That’s the first step of our private AI engagements.