ArchitectureEnterprise AI
A reference architecture for on-premise LLM deployment
Running LLMs inside your own data center or private cloud gives you control over where prompts are processed, which model versions run and what gets logged. It also makes you the operator of GPUs, serving software, model imports and an internal gateway. This reference architecture covers each component, a sizing method, updates for restricted networks and release practices that keep model changes reversible.
On this page
- When self-hosting earns its operating cost
- How one request moves through a private LLM stack
- Choosing open-weight models: capability, size and license
- Inference serving engines and what to compare
- A capacity planning method for GPU clusters
- What the AI gateway must enforce
- Updating models in restricted or air-gapped networks
- Operating risks in a private LLM cluster
- Running costs and the skills a private cluster needs
- Questions and answers
- Sources
When self-hosting earns its operating cost
- If
Contracts, your classification policy or a regulator forbid processing outside your environment.
ThenSelf-host for the affected workflows and keep the others on managed services.
The constraint is binary; no provider contract term can resolve it.
- If
A provider may process the data under a data-processing agreement with no training use, limited retention and private connectivity.
ThenUse a managed model service through private endpoints before building a cluster.
You control the data path without taking on GPU operations.
- If
The task needs capability that open-weight models do not yet match on your evaluation set.
ThenUse a managed model for that task and re-test open-weight releases as they appear.
A weaker private model does not help if it fails the task.
- If
The site has no dependable internet connection or must remain air-gapped.
ThenSelf-host, and design the offline update path before the first model goes in.
Retrofitting updates into a disconnected site is slow and error-prone.
How one request moves through a private LLM stack
Every call passes through the gateway on the way in and on the way out, which keeps identity, policy and audit consistent across models.
- Calling application
An assistant, agent or batch job holding gateway credentials, never direct model access.
- AI gateway
Authenticates, checks quota and policy, resolves the model alias and writes the audit record.
- Retrieval service
Searches indexes inside the perimeter and filters results by the end user's permissions.
- Inference cluster
A serving engine on GPU nodes running a pinned, registered model version.
- Telemetry and audit
Collects traces, token counts, latency and policy events under defined retention.
Choosing open-weight models: capability, size and license
Start from your evaluation set, not a public leaderboard. Shortlist open-weight models that handle your languages, document types and output formats, then test them on real tasks from the workflows you intend to serve. A smaller model that passes is cheaper, faster and easier to fit on existing hardware.
Size determines memory. Weights need roughly the parameter count times the bytes stored per parameter, so quantizing to lower precision cuts memory and can raise throughput, at a quality cost only your evaluation set can measure. Mixture-of-experts models compute with part of the network per token but must hold all of it in memory.
Read the license before the benchmark results. Some models ship under permissive terms such as the Apache License 2.01; others use custom licenses with acceptable-use policies, attribution requirements or conditions for very large deployments. Check whether the terms cover fine-tuned derivatives and sharing with subsidiaries, and record the answer in the model registry.
Verify provenance too: download from the publisher's official repository, compare published hashes and prefer the safetensors format, designed as a safe alternative to pickle-based checkpoints2.
Inference serving engines and what to compare
The serving engine sets throughput, hardware support and how much integration code you write. Compare candidates on your own model and traffic, not published benchmarks.
| Criterion | LLM-focused engines (vLLM, SGLang) | General model server (NVIDIA Dynamo-Triton) | Local engines (llama.cpp and similar) |
|---|---|---|---|
| Main strength | vLLM documents paged attention memory management, continuous batching and prefix caching for high-throughput serving3. | Serving models from many frameworks, including TensorRT, PyTorch, ONNX and OpenVINO, with multi-model ensembles5. | Running quantized models on modest hardware, such as workstations or field devices, for light loads. |
| Integration surface | vLLM ships an OpenAI-compatible API server, which eases migration from hosted APIs3. | NVIDIA points LLM-specific serving to its companion Dynamo framework, which adds features such as disaggregated serving5. | Lightweight local servers suited to single users or low concurrency. |
| Hardware breadth | vLLM lists NVIDIA and AMD GPUs and several CPU architectures3; SGLang lists GPUs, TPUs, NPUs, CPUs and Apple Silicon4. | Strongest on NVIDIA GPUs and the TensorRT toolchain. | CPUs and consumer GPUs, including Apple Silicon. |
| What to test | Latency and throughput at your concurrency, context length and quantization. | Whether LLM latency targets hold without adding Dynamo. | Whether answer quality survives heavy quantization. |
Hugging Face's Text Generation Inference is in maintenance mode, and its maintainers recommend vLLM and SGLang for new work6; existing TGI deployments should plan a migration.
A capacity planning method for GPU clusters
Sizing is a calculation followed by a test: the calculation decides what to trial, the test decides what to buy.
Describe the workload
Per workflow, record peak concurrent requests, input and output lengths and whether traffic is interactive or batch. Long retrieved contexts matter more than user counts.
Set latency targets
Agree time to first token and generation speed for interactive use, or a completion deadline for batch jobs. These decide how aggressively requests can be batched.
Estimate memory per replica
Add the weights to the key-value cache. For every token in flight, the cache holds a key and a value per layer and attention head group, times the bytes per value. Scale by peak tokens in flight and add runtime overhead.
Choose parallelism and replicas
Use tensor parallelism across GPUs in one node when a model does not fit on a single card; add whole replicas for throughput and resilience.
Load-test with replayed traffic
Run the chosen engine with your model, quantization and production-shaped synthetic prompts, raising concurrency until a latency target breaks.
Add headroom
Keep spare capacity to lose a node at peak and to run a candidate model beside the current one during an A/B comparison.
Re-measure after every change
A new model, quantization level, engine version or context limit invalidates earlier results; repeat the load test before shifting traffic.
What the AI gateway must enforce
Applications never call the inference cluster directly; the gateway is where enterprise rules become code.
Updating models in restricted or air-gapped networks
In a disconnected environment, every model, container image, driver and library arrives through a controlled import path. A connected staging environment downloads approved artifacts, verifies hashes or signatures, scans images, records a software bill of materials and exports a signed bundle. On the restricted side, an import station verifies it again before loading the internal registries. Pin driver, CUDA toolkit and engine versions as a set, and schedule imports like patch cycles so serving-stack security fixes do not wait for the next model release.
Operating risks in a private LLM cluster
ColdAI designs inference pipelines that run entirely within client infrastructure, with model versioning, A/B testing and audit logging7. These practices contain the risks below.
Silent quality regression after a model swap
Early signalUsers report worse answers days after an upgrade that passed its load test.
MitigationGate each new version on the workflows' evaluation sets, run it in shadow or on an A/B split, and keep the previous version deployed until sign-off, so rollback is a gateway alias change.
Memory exhaustion under long contexts
Early signalRequests fail or queue sharply when users paste long documents.
MitigationEnforce a maximum context per workflow at the gateway and size the cache for that limit, not for the average request.
Patch lag in the serving stack
Early signalEngine and container versions trail upstream security fixes by months.
MitigationTrack serving components in vulnerability management like any production service, with upgrades tested on a staging cluster.
Running costs and the skills a private cluster needs
Beyond hardware or reserved GPU capacity, budget for power and cooling, support contracts for the serving stack, storage for weights and indexes, and the evaluation runs every model change requires; the AI total cost of ownership model breaks these down. You also need engineers who can operate GPU nodes, a container platform, serving engines and the gateway, plus security staff who treat the AI stack as production infrastructure. Without them, start with a managed service over private connectivity.
Questions and answers
Can we mix self-hosted and cloud-hosted models?
Yes, and most enterprises do. The gateway routes each workflow to a permitted model by data class: restricted workflows to the private cluster, others to managed services over private networking. Applications call the same interface either way, so moving a workflow between tiers is a routing change plus a fresh run of its evaluation set.
Should several teams share one GPU cluster?
Sharing improves utilization, the main driver of cost. It works when the gateway enforces quotas and priorities, interactive and batch traffic are separated and each team's usage is metered. Separate clusters make sense when data classes or legal entities must be isolated, or a workload needs guaranteed capacity.
How do we keep the serving stack patched?
Treat the inference engine, containers, drivers and gateway as production software under vulnerability management. Load-test and evaluate upgrades on a staging cluster before release, pin compatible versions together and, in air-gapped sites, import on a regular patch cycle rather than waiting for model releases.
Can we fine-tune models on-premise as well?
Yes, if the license permits derivatives and you have training capacity, which is sized differently from inference. Retrieval or instruction changes often solve the problem first; the LLM fine-tuning practice covers when tuning pays off. Tuned weights then enter the same registry, evaluation and release process as any other model.
Sources
- Apache License, Version 2.0 — The Apache Software Foundation · checked 10 October 2026
- Safetensors documentation — Hugging Face · checked 10 October 2026
- vLLM documentation — vLLM project · checked 10 October 2026
- SGLang repository — SGLang project · checked 10 October 2026
- NVIDIA Dynamo-Triton (formerly Triton Inference Server) — NVIDIA · checked 10 October 2026
- Text Generation Inference documentation — Hugging Face · checked 10 October 2026
- Enterprise AI: private inference pipelines with versioning, A/B testing and audit logging — ColdAI