Home / Articles / Two LLM workloads on one GPU: vLLM LoRA, agents, and Azure ML fine-tunes

This article is published in English.

Two LLM workloads on one GPU: vLLM LoRA, agents, and Azure ML fine-tunes

Serve interactive and batch agent traffic from one card with demand LoRA, ClickHouse analytics, Blob artifacts, and Unsloth training tracked in MLflow.

1917 words

The core challenge: two workloads, one model, one budget

Many teams need both interactive agent chat and batch analytics over the same fine-tuned model family—but GPU budget says “one card.” This architecture serves two LLM workloads from a single GPU on Azure: a low-latency inference path with demand-loaded LoRA adapters via vLLM, and an agentic analytics path that writes structured events into ClickHouse, with training/fine-tuning on Azure ML using Unsloth and MLflow.

The inference layer: vLLM + demand LoRA

vLLM hosts the base model. LoRA adapters load on demand so specialized behaviors (support tone, analytics extraction, tool-calling variants) do not each need a full replica. Keep adapter sizes small; thrash shows up as latency spikes when too many adapters rotate.

vllm serve your-org/base-vlm-7b \
  --enable-lora \
  --max-lora-rank 64 \
  --lora-modules domain-adapter=/data/lora-adapters/current/ \
  --host 0.0.0.0 \
  --port 8000 \
  --gpu-memory-utilization 0.90 \
  --max-model-len 8192
payload = {
    "model": "domain-adapter",
    "messages": [
        {"role": "user", "content": [
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_base64}"}},
            {"type": "text", "text": f"Analyse this for category: {category}, location: {label}"}
        ]}
    ],
    "max_tokens": 1024
}
response = requests.post(VLLM_API_URL, json=payload, timeout=30,
                          proxies={"http": None, "https": None})
from agents.extensions.models.litellm_model import LitellmModel

model = LitellmModel(
    model="openai/base-vlm-7b",
    base_url="http://10.x.x.x:8000/v1",
    api_key="not-needed"
)
from agents import Agent, Runner, function_tool as tool

@tool
def get_dashboard_snapshot_tool() -> Dict[str, Any]:
    """Get all KPIs in one call. Call this FIRST for any summary question."""
    with SessionLocal() as db:
        return jsonable_encoder(analytics_service.get_dashboard_snapshot(db))

@tool
def get_entity_ranking_tool(top: bool = True) -> List[Dict[str, Any]]:
    """Get ranked entity performance list. top=True for best, False for worst."""
    with SessionLocal() as db:
        return jsonable_encoder(analytics_service.get_performance_ranking(db, top))

overview_agent = Agent(
    name="OverviewAgent",
    instructions="Handle general dashboard and summary questions. Always call get_dashboard_snapshot_tool first.",
    tools=[get_dashboard_snapshot_tool, get_entity_ranking_tool],
    model=model,
)

status_agent = Agent(
    name="StatusAgent",
    instructions="Handle live status and stream/field health questions.",
    tools=[get_stream_status_tool, get_field_status_tool],
    model=model,
)

quality_agent = Agent(
    name="QualityAgent",
    instructions="Handle rejection-rate and data-quality questions.",
    tools=[get_rejection_stats_tool],
    model=model,
)

orchestrator = Agent(
    name="AnalyticsOrchestrator",
    instructions=(
        "Handle all user communication. Route each question to the right specialist tool "
        "and synthesize its result into a plain-text answer. Do not hallucinate metrics; "
        "only report what a specialist returns."
    ),
    tools=[
        overview_agent.as_tool(
            tool_name="overview_expert",
            tool_description="Answers general dashboard and summary questions.",
        ),
        status_agent.as_tool(
            tool_name="status_expert",
            tool_description="Answers live status and stream/field health questions.",
        ),
        quality_agent.as_tool(
            tool_name="quality_expert",
            tool_description="Answers rejection-rate and data-quality questions.",
        ),
    ],
)
CREATE TABLE events_queue (
    event_id UUID,
    entity_id UInt64,
    event_type LowCardinality(String),
    payload String,
    created_at DateTime64(3)
) ENGINE = Kafka
SETTINGS
    kafka_broker_list = 'your-namespace.servicebus.windows.net:9093',
    kafka_topic_list = 'app-events',
    kafka_group_name = 'clickhouse-consumer',
    kafka_format = 'JSONEachRow';

CREATE MATERIALIZED VIEW events_mv TO events AS
SELECT * FROM events_queue;
azureblob://
├── ml-artifacts/
│   ├── lora-adapters/{v1, v2, v3}
│   ├── base-models/base-vlm-7b/
│   └── eval-results/v3/
├── training-data/{raw, processed, annotations}
└── inference-inputs/{year}/{month}/{day}/{entity_id}/{request_id}.bin
import great_expectations as gx

context = gx.get_context()
batch = context.sources.pandas_default.read_json("annotations_batch.jsonl")
suite = context.get_expectation_suite("annotation_schema_v1")
results = context.run_validation_operator(
    "action_list_operator",
    assets_to_validate=[batch],
    run_id="training-v4-ingestion"
)
if not results["success"]:
    raise ValueError(f"Data validation failed: {results['statistics']}")
import mlflow
from unsloth import FastVisionModel

mlflow.set_experiment("domain-lora-finetuning")
with mlflow.start_run(run_name=f"lora-v{adapter_version}") as run:
    model, tokenizer = FastVisionModel.from_pretrained(
        "unsloth/base-vlm-7b",
        load_in_4bit=True,
    )
    model = FastVisionModel.get_peft_model(
        model,
        r=64,
        lora_alpha=128,
        finetune_vision_layers=True,
    )
    mlflow.log_params({
        "base_model": "base-vlm-7b",
        "lora_rank": 64,
        "lora_alpha": 128,
        "epochs": 3,
        "dataset_version": "v4",
    })
    for epoch in range(epochs):
        train_loss = train_one_epoch(model, dataloader)
        mlflow.log_metric("train_loss", train_loss, step=epoch)
    metrics = evaluate(model, eval_dataloader)
    mlflow.log_metrics({"eval_f1": metrics["f1"]})
    mlflow.log_artifacts(output_dir, artifact_path="lora-adapter")
    mlflow.set_tag("promotion_status", "candidate")
# azure-ml-lora-job.yml
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json
type: command
code: ./training/
command: >
  python finetune_lora_unsloth.py
  --dataset ${{inputs.training_data}}
  --output-dir ${{outputs.lora_adapter}}
  --lora-rank 64 --lora-alpha 128 --epochs 3
inputs:
  training_data:
    type: uri_folder
    path: azureml://datastores/training_blob/paths/domain/v4/
outputs:
  lora_adapter:
    type: uri_folder
    path: azureml://datastores/artifacts_blob/paths/lora-adapters/v4/
environment: azureml:vlm-unsloth-env:1
compute: azureml:gpu-cluster-nc24ads   # single A100, sufficient for Unsloth LoRA fine-tuning
experiment_name: domain-lora-finetuning
from azure.ai.ml.entities import Model

model = Model(
    name="domain-lora-adapter",
    version="4",
    path="azureml://datastores/artifacts_blob/paths/lora-adapters/v4/",
    tags={
        "base_model": "base-vlm-7b",
        "lora_rank": "64",
        "eval_f1": "0.93",   # illustrative
        "promotion_status": "candidate",
    }
)
ml_client.models.create_or_update(model)
def run_eval(adapter_version: str, eval_dataset_version: str):
    candidate = evaluate_adapter(adapter_path=..., eval_data=...)
    production = evaluate_adapter(adapter_path=get_production_adapter_path(), eval_data=...)
    passed = (
        candidate["f1"] >= production["f1"] - 0.01
        and candidate["f1"] >= MIN_F1_THRESHOLD
      )
    return passed, candidate
New annotation batch → Great Expectations validation
    → dataset versioned in Azure ML
    → Azure ML training job (Unsloth, single GPU, spot instance)
    → offline eval vs. production, per-category
    → [pass] tag adapter "production" in registry
    → sync adapter to local disk on the vLLM host
    → rolling restart of vLLM with new --lora-modules path
# .gitlab-ci.yml
stages:
  - validate
  - train
  - eval
  - deploy

validate-data:
  stage: validate
  script:
    - python eval/validate_dataset.py --version $DATASET_VERSION
submit-training:
  stage: train
  needs: [validate-data]
  script:
    - az ml job create --file azure-ml-lora-job.yml
      --set inputs.dataset_version=$DATASET_VERSION
      --workspace-name $AML_WORKSPACE --resource-group $RESOURCE_GROUP
run-eval:
  stage: eval
  needs: [submit-training]
  script:
    - python eval/run_eval.py --version $DATASET_VERSION
    - python scripts/promote_adapter.py --version $DATASET_VERSION
deploy:
  stage: deploy
  needs: [run-eval]
  when: manual
  script:
    - az containerapp update --name vllm-server --resource-group $RESOURCE_GROUP
      --set-env-vars LORA_ADAPTER_VERSION=$DATASET_VERSION

The agentic layer: three specialized agents, one orchestrator

Split responsibilities:

  1. Router / orchestrator — classifies intent and picks tools
  2. Retrieval specialist — fetches domain context
  3. Analytics specialist — emits structured facts for warehousing

Shared GPU inference means careful concurrency limits and queueing so interactive chat is not starved by batch agent runs.

The analytics backend: ClickHouse

Agent outputs that matter for product analytics land in ClickHouse: latency, tool choices, retrieved doc IDs, user outcomes. Columnar storage fits high-cardinality event streams better than forcing everything through the OLTP database.

Azure Blob Storage: the MLOps glue layer

Checkpoints, datasets, adapter artifacts, and evaluation reports live in Blob. Training jobs read/write here; inference pulls approved adapters by version. Treat blob paths as part of the API contract between training and serving.

The training pipeline: Unsloth on Azure ML

Fine-tune with Unsloth for speed/memory efficiency on Azure ML compute. Track runs in MLflow: hyperparameters, eval metrics, artifact URIs. Only promote adapters that beat a baseline on a frozen eval set.

Data validation before training

Validate schema, PII scrubbing, class balance, and leakage between train/eval before GPU minutes burn. Fail closed on validation errors.

Fine-tuning with Unsloth, tracked in MLflow

Log adapter weights to Blob with immutable version IDs. Inference configuration references those IDs explicitly—no “latest” in production without a pin.

Serving both workloads safely

  • Separate queues for interactive vs batch
  • Per-adapter concurrency caps
  • Warm a default adapter; cold-load specialists
  • Emit GPU utilization and queue depth to the same ClickHouse (or metrics stack)
  • Roll back adapters by flipping a config pointer, not by redeploying the whole VM

Takeaways

One GPU can host two product surfaces if LoRA demand-loading, queue isolation, and MLOps artifact discipline are first-class. vLLM + Unsloth + Azure ML + Blob + ClickHouse is a practical MLOps shape for teams that cannot afford duplicate inference fleets yet still need specialized behaviors and analytics.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.

Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.