Home / Articles / What LiteLLM cannot see: monitoring GPU serving under the gateway

This article is published in English.

What LiteLLM cannot see: monitoring GPU serving under the gateway

LiteLLM tracks requests and spend. Queue wait, prefill versus decode, container restarts, and host pressure need Prometheus, vLLM metrics, cAdvisor, and traces.

975 words

Moving from a hand-rolled FastAPI gateway to LiteLLM centralized request logs, token and virtual-key accounting, and cut custom gateway code. A gateway still only sees traffic that crosses its boundary. Production serving raised an entire category of questions LiteLLM alone cannot answer—questions about queues, prefill versus decode, container restarts, and host saturation that decide whether an eleven-second call was healthy or stuck.

Background

LiteLLM sits at the request edge. It records that a call arrived, which model served it, token counts, which virtual key paid, and success or failure. For day-to-day cost tracking, that is most of what finance and product need.

Below that edge, the picture diverges. An eleven-second request might be queue wait on a saturated GPU or healthy generation of a long answer. From the gateway those look identical; the fixes do not. Treating them as one metric sends on-call in the wrong direction.

Three layers beneath the gateway each need their own truth:

  • The model runtime, where latency is decided. Time to first token mixes queue wait and prefill more than decode. Prefill and decode throughputs bottleneck differently. Those signals come from the serving engine’s own /metrics (for example vLLM), not from the gateway.
  • Containers, via cAdvisor: which process is eating memory, which container restarted overnight.
  • Hosts, via node-exporter: total CPU, memory, and disk on the box that hosts the GPUs.

Four sources matter because no single collector sees every layer. Gateway logs without runtime metrics are a partial story; runtime metrics without host and container context miss the noisy neighbor that restarted at three in the morning.

The full stack

On one on-premise GPU server, two Docker Compose projects separate concerns deliberately.

Gateway compose: LiteLLM, PostgreSQL for virtual keys and spend, Redis for rate-limit cache.

Monitoring compose: Prometheus, Grafana, node-exporter, cAdvisor. vLLM processes often run on the host outside both stacks, one process per served model. Prometheus scrapes the four targets that cover gateway-adjacent and model-local metrics. The split is not aesthetic—it is how optional infrastructure stays optional.

Key decisions

1. Two compose files, not one. LiteLLM was already serving other teams when monitoring was added. Folding Prometheus into the same compose would couple every scrape-config edit to the production gateway file. Separation gives fault isolation: rebuild monitoring freely; LiteLLM does not notice. Dependency flows one way. Networking between stacks is a one-time cost against ongoing blast-radius risk. If monitoring dies, models still answer; if the gateway dies, monitoring still records host health for the postmortem.

2. No postgres-exporter or redis-exporter by default. The host could run them comfortably. They were still deferred:

  • While LiteLLM behaves, database and cache internals rarely need a dashboard; failures surface at the gateway first.
  • node-exporter, cAdvisor, and the model /metrics endpoints already cover the critical layers.
  • Extra exporters are extra versions and operational memory.

Unused dashboards have maintenance cost without payoff. Revisit when connection errors to PostgreSQL become regular, virtual-key auth slows from query contention, or Redis memory pressure is a live theory—then add exporters that week. Write the reversal criteria down so the skip stays a decision rather than amnesia.

3. Keeping Langfuse. LiteLLM consolidates request observability, yet one product action is often many model calls: retrieve, summarize, follow up. Shared session ids can group calls in LiteLLM, but trace UX and retention differ. Gateway logs are not meant as indefinite archives. Older traces feed evaluation sets for model swaps and help debug reports from prior weeks. Metrics aggregate; traces reconstruct. No amount of the first fully covers the second—which is why both stay.

Running it in practice

1. Declaring models only in config.yaml was painful. File-owned models show a config badge in the UI and resist edit or delete without editing the mount and reloading a busy proxy. Models registered through the Admin API live in PostgreSQL and survive restarts:

curl -X POST <http://localhost:4000/model/new> \
  -H "Authorization: Bearer$LITELLM_MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model_name": "CHAT_MODEL",
    "litellm_params": {
      "model": "openai/CHAT_MODEL",
      "api_base": "<http://host.docker.internal:8001/v1>",
      "api_key": "dummy"
    }
  }'

Prefer config for settings and the database for the model catalog. That split lets operators add an experimental model without touching the gateway process other teams are calling.

2. The node-exporter dashboard saw little use. After workloads stabilized and deploys slowed, host totals stopped answering interesting questions. Keep the scrape; do not over-invest in unused panels. cAdvisor and vLLM panels earned more daily attention because they map to user-visible latency.

3. Useful Grafana starters

  • vLLM community dashboards (for example the public grafana.com dashboard numbered 23991)
  • cAdvisor (14282)
  • Node Exporter Full (1860)

Import them as baselines, then delete panels that never get opened.

Sum up

Alerting is the obvious gap: metrics exist, but no page fires when thresholds cross. Wiring application mail APIs onto GPU hosts mixes concerns; pick a safer notification path next. Conceptually, Prometheus and Grafana answer questions already known to ask; Langfuse reconstructs what one user action did across fan-out calls. Aggregation and reconstruction are complementary. A gateway migration without a monitoring plan only moves the blind spot one layer down.