ArchitectureEnterprise AI

A reference architecture for on-premise LLM deployment

Running LLMs inside your own data center or private cloud gives you control over where prompts are processed, which model versions run and what gets logged. It also makes you the operator of GPUs, serving software, model imports and an internal gateway. This reference architecture covers each component, a sizing method, updates for restricted networks and release practices that keep model changes reversible.

Reviewed 8 min read

On this page
  1. When self-hosting earns its operating cost
  2. How one request moves through a private LLM stack
  3. Choosing open-weight models: capability, size and license
  4. Inference serving engines and what to compare
  5. A capacity planning method for GPU clusters
  6. What the AI gateway must enforce
  7. Updating models in restricted or air-gapped networks
  8. Operating risks in a private LLM cluster
  9. Running costs and the skills a private cluster needs
  10. Questions and answers
  11. Sources

When self-hosting earns its operating cost

  • If

    Contracts, your classification policy or a regulator forbid processing outside your environment.

    Then

    Self-host for the affected workflows and keep the others on managed services.

    The constraint is binary; no provider contract term can resolve it.

  • If

    A provider may process the data under a data-processing agreement with no training use, limited retention and private connectivity.

    Then

    Use a managed model service through private endpoints before building a cluster.

    You control the data path without taking on GPU operations.

  • If

    The task needs capability that open-weight models do not yet match on your evaluation set.

    Then

    Use a managed model for that task and re-test open-weight releases as they appear.

    A weaker private model does not help if it fails the task.

  • If

    The site has no dependable internet connection or must remain air-gapped.

    Then

    Self-host, and design the offline update path before the first model goes in.

    Retrofitting updates into a disconnected site is slow and error-prone.

How one request moves through a private LLM stack

Every call passes through the gateway on the way in and on the way out, which keeps identity, policy and audit consistent across models.

Request + user identityPermission-scoped queryAllowed passagesPrompt to model aliasStreamed tokensTrace and audit eventFiltered response01Callingapplication02AI gateway03Retrievalservice04Inferencecluster05Telemetry andaudit
  1. Calling application

    An assistant, agent or batch job holding gateway credentials, never direct model access.

  2. AI gateway

    Authenticates, checks quota and policy, resolves the model alias and writes the audit record.

  3. Retrieval service

    Searches indexes inside the perimeter and filters results by the end user's permissions.

  4. Inference cluster

    A serving engine on GPU nodes running a pinned, registered model version.

  5. Telemetry and audit

    Collects traces, token counts, latency and policy events under defined retention.

  1. Calling application to AI gatewayRequest + user identity
  2. AI gateway to Retrieval servicePermission-scoped query
  3. Retrieval service to AI gatewayAllowed passages
  4. AI gateway to Inference clusterPrompt to model alias
  5. Inference cluster to AI gatewayStreamed tokens
  6. AI gateway to Telemetry and auditTrace and audit event
  7. AI gateway to Calling applicationFiltered response
Conceptual request path in a self-hosted deployment; real systems add caching, retries and fallbacks.

Choosing open-weight models: capability, size and license

Start from your evaluation set, not a public leaderboard. Shortlist open-weight models that handle your languages, document types and output formats, then test them on real tasks from the workflows you intend to serve. A smaller model that passes is cheaper, faster and easier to fit on existing hardware.

Size determines memory. Weights need roughly the parameter count times the bytes stored per parameter, so quantizing to lower precision cuts memory and can raise throughput, at a quality cost only your evaluation set can measure. Mixture-of-experts models compute with part of the network per token but must hold all of it in memory.

Read the license before the benchmark results. Some models ship under permissive terms such as the Apache License 2.01; others use custom licenses with acceptable-use policies, attribution requirements or conditions for very large deployments. Check whether the terms cover fine-tuned derivatives and sharing with subsidiaries, and record the answer in the model registry.

Verify provenance too: download from the publisher's official repository, compare published hashes and prefer the safetensors format, designed as a safe alternative to pickle-based checkpoints2.

Inference serving engines and what to compare

The serving engine sets throughput, hardware support and how much integration code you write. Compare candidates on your own model and traffic, not published benchmarks.

CriterionLLM-focused engines (vLLM, SGLang)General model server (NVIDIA Dynamo-Triton)Local engines (llama.cpp and similar)
Main strengthvLLM documents paged attention memory management, continuous batching and prefix caching for high-throughput serving3.Serving models from many frameworks, including TensorRT, PyTorch, ONNX and OpenVINO, with multi-model ensembles5.Running quantized models on modest hardware, such as workstations or field devices, for light loads.
Integration surfacevLLM ships an OpenAI-compatible API server, which eases migration from hosted APIs3.NVIDIA points LLM-specific serving to its companion Dynamo framework, which adds features such as disaggregated serving5.Lightweight local servers suited to single users or low concurrency.
Hardware breadthvLLM lists NVIDIA and AMD GPUs and several CPU architectures3; SGLang lists GPUs, TPUs, NPUs, CPUs and Apple Silicon4.Strongest on NVIDIA GPUs and the TensorRT toolchain.CPUs and consumer GPUs, including Apple Silicon.
What to testLatency and throughput at your concurrency, context length and quantization.Whether LLM latency targets hold without adding Dynamo.Whether answer quality survives heavy quantization.

Hugging Face's Text Generation Inference is in maintenance mode, and its maintainers recommend vLLM and SGLang for new work6; existing TGI deployments should plan a migration.

A capacity planning method for GPU clusters

Sizing is a calculation followed by a test: the calculation decides what to trial, the test decides what to buy.

  1. Describe the workload

    Per workflow, record peak concurrent requests, input and output lengths and whether traffic is interactive or batch. Long retrieved contexts matter more than user counts.

    Output
    Workload profile
  2. Set latency targets

    Agree time to first token and generation speed for interactive use, or a completion deadline for batch jobs. These decide how aggressively requests can be batched.

    Output
    Service-level targets
  3. Estimate memory per replica

    Add the weights to the key-value cache. For every token in flight, the cache holds a key and a value per layer and attention head group, times the bytes per value. Scale by peak tokens in flight and add runtime overhead.

    Output
    Memory budget
  4. Choose parallelism and replicas

    Use tensor parallelism across GPUs in one node when a model does not fit on a single card; add whole replicas for throughput and resilience.

    Output
    Node layout
  5. Load-test with replayed traffic

    Run the chosen engine with your model, quantization and production-shaped synthetic prompts, raising concurrency until a latency target breaks.

    Output
    Measured capacity curve
  6. Add headroom

    Keep spare capacity to lose a node at peak and to run a candidate model beside the current one during an A/B comparison.

    Output
    Production sizing
  7. Re-measure after every change

    A new model, quantization level, engine version or context limit invalidates earlier results; repeat the load test before shifting traffic.

    Output
    Updated capacity record

What the AI gateway must enforce

Applications never call the inference cluster directly; the gateway is where enterprise rules become code.

0 of 7 checked

Updating models in restricted or air-gapped networks

In a disconnected environment, every model, container image, driver and library arrives through a controlled import path. A connected staging environment downloads approved artifacts, verifies hashes or signatures, scans images, records a software bill of materials and exports a signed bundle. On the restricted side, an import station verifies it again before loading the internal registries. Pin driver, CUDA toolkit and engine versions as a set, and schedule imports like patch cycles so serving-stack security fixes do not wait for the next model release.

Operating risks in a private LLM cluster

ColdAI designs inference pipelines that run entirely within client infrastructure, with model versioning, A/B testing and audit logging7. These practices contain the risks below.

Silent quality regression after a model swap

Early signalUsers report worse answers days after an upgrade that passed its load test.

MitigationGate each new version on the workflows' evaluation sets, run it in shadow or on an A/B split, and keep the previous version deployed until sign-off, so rollback is a gateway alias change.

Memory exhaustion under long contexts

Early signalRequests fail or queue sharply when users paste long documents.

MitigationEnforce a maximum context per workflow at the gateway and size the cache for that limit, not for the average request.

Patch lag in the serving stack

Early signalEngine and container versions trail upstream security fixes by months.

MitigationTrack serving components in vulnerability management like any production service, with upgrades tested on a staging cluster.

Running costs and the skills a private cluster needs

Beyond hardware or reserved GPU capacity, budget for power and cooling, support contracts for the serving stack, storage for weights and indexes, and the evaluation runs every model change requires; the AI total cost of ownership model breaks these down. You also need engineers who can operate GPU nodes, a container platform, serving engines and the gateway, plus security staff who treat the AI stack as production infrastructure. Without them, start with a managed service over private connectivity.

Questions and answers

Can we mix self-hosted and cloud-hosted models?

Yes, and most enterprises do. The gateway routes each workflow to a permitted model by data class: restricted workflows to the private cluster, others to managed services over private networking. Applications call the same interface either way, so moving a workflow between tiers is a routing change plus a fresh run of its evaluation set.

Should several teams share one GPU cluster?

Sharing improves utilization, the main driver of cost. It works when the gateway enforces quotas and priorities, interactive and batch traffic are separated and each team's usage is metered. Separate clusters make sense when data classes or legal entities must be isolated, or a workload needs guaranteed capacity.

How do we keep the serving stack patched?

Treat the inference engine, containers, drivers and gateway as production software under vulnerability management. Load-test and evaluate upgrades on a staging cluster before release, pin compatible versions together and, in air-gapped sites, import on a regular patch cycle rather than waiting for model releases.

Can we fine-tune models on-premise as well?

Yes, if the license permits derivatives and you have training capacity, which is sized differently from inference. Retrieval or instruction changes often solve the problem first; the LLM fine-tuning practice covers when tuning pays off. Tuned weights then enter the same registry, evaluation and release process as any other model.

Sources

  1. Apache License, Version 2.0 — The Apache Software Foundation · checked 10 October 2026
  2. Safetensors documentation — Hugging Face · checked 10 October 2026
  3. vLLM documentation — vLLM project · checked 10 October 2026
  4. SGLang repository — SGLang project · checked 10 October 2026
  5. NVIDIA Dynamo-Triton (formerly Triton Inference Server) — NVIDIA · checked 10 October 2026
  6. Text Generation Inference documentation — Hugging Face · checked 10 October 2026
  7. Enterprise AI: private inference pipelines with versioning, A/B testing and audit logging — ColdAI

More in Enterprise AI

Back to Enterprise AI

Next step

Share your constraints for a private inference design

Tell us which data classes must stay inside your environment, the workflows you want to serve and the hardware or cloud tenancy you have. We will reply with a view on whether self-hosting fits and which components to design first.

Discuss private inference