Home / Articles / Self-hosted LLM networking: one gateway instead of a host per model

This article is published in English.

Self-hosted LLM networking: one gateway instead of a host per model

Adding open-weight models should not mean new hostnames and peering for every app. A LiteLLM-style gateway keeps one endpoint while backends change behind it.

1044 words

Self-hosting open-weight models often starts simple: one model, one hostname, one happy path. The pain appears when a second and third model arrive. Each new weight file tends to drag along fresh networking, peering, and application config—even though the product only asked for another model id. Putting a gateway in front of the backends restores a single endpoint while models come and go behind it. The lesson is architectural, not product-specific: stop teaching every client the topology of every GPU box.

It started with one model

An open-weight Qwen-class checkpoint served locally and exposed through a hostname:

llm.my-app.com

The application called that endpoint with a model name and the rest of the request payload. Traffic worked; dashboards looked fine. Then a faster Gemma instance wanted its own listener for latency-sensitive paths:

flash.llm.my-app.com

That worked too. Image models and evaluation copies of DeepSeek-class checkpoints invited yet more hosts:

images.llm.my-app.com
deepseek-v4.llm.my-app.com

Production Qwen stayed untouched while development pointed at the experiment endpoint. Each step was individually reasonable. Collectively they created a mesh of base URLs that only the person who wired them could remember.

It worked, but something felt wrong

None of the steps were hard—and that was the problem. Every model repeated the same choreography: deploy weights, expose a service, configure networking, peer VPCs so private model hosts reach the app, then teach the app a new base URL. None of that is “the model.” It is undifferentiated plumbing. Teammates could not be told “use our LLM service and pick this model”; they needed the right host spelled out, and the conversation restarted whenever a new endpoint appeared. Onboarding a contractor meant shipping a private glossary of hostnames instead of a single catalog.

The DeepSeek experiment made this obvious

Evaluating a separate DeepSeek-class deployment without disturbing production Qwen meant yet another endpoint wired into development. Technically nothing was broken. The nagging question remained: why should adding a model require adding networking? The artifact that changed was the weights; the application’s infrastructure should not have to change in lockstep. If every experiment mutates DNS and peering, experimentation slows to the pace of network change tickets.

Then how large providers expose models

Major providers do not mint a distinct public base URL per model. Callers do not juggle families of hosts such as:

gpt-4.api.openai.com
gpt-5.api.openai.com
some-other-model.api.openai.com

They keep one API and select the model inside the request. If providers can front dozens of models behind one surface, self-hosted fleets can aim for the same shape. The client contract becomes stable; the fleet behind it can churn.

The idea: put a proxy in front of the models

Applications should always talk to one stable base:

llm.my-app.com

and choose the model in the body:

{
  model: "qwen-3.8",
  ...
}
{
  model: "gemma-3",
  ...
}
{
  model: "deepseek-v4",
  ...
}

Backends can move between machines, regions, or runtimes; the client contract stays put. That is the same abstraction cloud providers already sell—just aimed at open-weight fleets under your control.

Building a proxy versus adopting a gateway

A naive forwarder—read model, route, return—looks small until it must be a shared service. Then virtual keys, usage accounting, access control, model configuration, rate limits, logging, and an admin UI appear. That is no longer a toy proxy; it is an LLM gateway. Rebuilding those controls from scratch recreates years of edge cases (key rotation, per-team budgets, audit logs). LiteLLM targets that control-plane layer so teams do not have to.

Why LiteLLM fit

Applications keep one endpoint:

llm.my-app.com

LiteLLM routes to heterogeneous backends:

My Application
                        │
                        ▼
                 llm.my-app.com
                        │
                        ▼
                   LiteLLM
                  /   |    \
                 /    |     \
              Qwen  Gemma  DeepSeek

Adding an experimental model becomes: deploy weights, register them on the gateway, keep the app URL unchanged. Hostnames stop leaking into product code. Developers pick models the way they pick temperatures—inside the request—not by editing environment files per experiment.

Deployment was straightforward

LiteLLM ships as a Docker image and typically wants PostgreSQL for durable state such as keys and spend:

LiteLLM container
        │
        ├── PostgreSQL
        │
        └── Open-weight model servers

Model servers stay focused on inference; the gateway owns the control plane above them. GPU boxes no longer need to be directly addressable from every microservice. Peering concentrates on the gateway tier instead of an N-by-M mesh between apps and model hosts.

The part that actually changed

Before the gateway, adding a model felt like repeating the full networking playbook:

New model
   ↓
New deployment
   ↓
New networking
   ↓
New hostname
   ↓
Application changes

Afterwards the playbook shrinks to registration on a stable front door:

New model
   ↓
Deploy it
   ↓
Register it with the gateway
   ↓
Use the existing endpoint

The durable lesson: applications should not need to know where an LLM runs—only which model to request. LiteLLM is one practical way to enforce that boundary. The original pain was not the gateway product; it was the second model. Once that pattern is visible, every subsequent model either reinforces the mesh of hostnames or reinforces a single catalog. Choose the catalog.

Operational footnotes worth keeping

  • Treat model registration as a change-managed act: who may add backends, how keys map to teams, and how experiments expire.
  • Keep health checks on backends separate from the public endpoint so a cold model does not look like a total outage.
  • Document the one base URL in onboarding docs and refuse to publish additional hostnames to application teams unless there is a hard isolation requirement.

Those habits preserve the abstraction the gateway was introduced to create.