Insights / Private AI
Running AI on your own infrastructure: a decision guide
Open-weight models can now handle much of everyday business work. This guide covers when running them in-house makes sense and what it requires.
Prompts and documents stay on your network
Running a capable language model in-house once required a research budget. Today, open-weight model families from Meta, Mistral, Alibaba, Google, DeepSeek and OpenAI can be downloaded and run on hardware an organization owns. OpenAI states that its gpt-oss-120b model runs on a single 80 GB GPU, and the smaller gpt-oss-20b on machines with 16 GB of memory.
Whether it is worthwhile depends on the workload, the data involved and the volume of use.
When in-house AI makes sense
- Your data must stay in your environment. Privileged client documents, patient records, unreleased financial results, or source code covered by contract.
- The work is repetitive and well defined. Summarization, template-based drafting, field extraction and internal document search are well within the ability of mid-sized open models.
- Usage is high and steady. Per-request fees grow with volume, while owned hardware is a fixed cost.
- Availability matters. Offline sites, or processes that must continue during a provider outage.
The question has changed from “can we?” to “is it worth it for this work?”
When it does not
Frontier models from the major labs still lead on complex, open-ended reasoning. For light or occasional use, a subscription costs less than hardware. Organizations without staff to maintain a server should budget for managed support.
The components
A private deployment has four parts: the hardware (a workstation or GPU server), a model server that loads the model and handles requests, a retrieval layer that lets the model answer from your documents while enforcing existing permissions, and the interface your staff use. Results depend most on the retrieval layer: whether it finds the right documents, and whether its permissions match the ones you already have.
- Interfacethe application your staff use
- Retrieval layerfinds relevant documents and enforces your permissions
- Model serverloads the model and handles requests
- Hardwarea workstation or GPU server you own
Evaluate before you buy
Public benchmarks do not show how a model performs on your contracts or support tickets. Build a test set from real work, run candidate models against it, and compare the cost of each acceptable answer. This evaluation is the first step in our Private & Local AI engagements.