This article is published in English.
Two LLM workloads on one GPU: vLLM LoRA, agents, and Azure ML fine-tunes
Serve interactive and batch agent traffic from one card with demand LoRA, ClickHouse analytics, Blob artifacts, and Unsloth training tracked in MLflow.
The core challenge: two workloads, one model, one budget
Many teams need both interactive agent chat and batch analytics over the same fine-tuned model family—but GPU budget says “one card.” This architecture serves two LLM workloads from a single GPU on Azure: a low-latency inference path with demand-loaded LoRA adapters via vLLM, and an agentic analytics path that writes structured events into ClickHouse, with training/fine-tuning on Azure ML using Unsloth and MLflow.
The inference layer: vLLM + demand LoRA
vLLM hosts the base model. LoRA adapters load on demand so specialized behaviors (support tone, analytics extraction, tool-calling variants) do not each need a full replica. Keep adapter sizes small; thrash shows up as latency spikes when too many adapters rotate.
vllm serve your-org/base-vlm-7b \
--enable-lora \
--max-lora-rank 64 \
--lora-modules domain-adapter=/data/lora-adapters/current/ \
--host 0.0.0.0 \
--port 8000 \
--gpu-memory-utilization 0.90 \
--max-model-len 8192
payload = {
"model": "domain-adapter",
"messages": [
{"role": "user", "content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{image_base64}"}},
{"type": "text", "text": f"Analyse this for category: {category}, location: {label}"}
]}
],
"max_tokens": 1024
}
response = requests.post(VLLM_API_URL, json=payload, timeout=30,
proxies={"http": None, "https": None})
from agents.extensions.models.litellm_model import LitellmModel
model = LitellmModel(
model="openai/base-vlm-7b",
base_url="http://10.x.x.x:8000/v1",
api_key="not-needed"
)
from agents import Agent, Runner, function_tool as tool
@tool
def get_dashboard_snapshot_tool() -> Dict[str, Any]:
"""Get all KPIs in one call. Call this FIRST for any summary question."""
with SessionLocal() as db:
return jsonable_encoder(analytics_service.get_dashboard_snapshot(db))
@tool
def get_entity_ranking_tool(top: bool = True) -> List[Dict[str, Any]]:
"""Get ranked entity performance list. top=True for best, False for worst."""
with SessionLocal() as db:
return jsonable_encoder(analytics_service.get_performance_ranking(db, top))
overview_agent = Agent(
name="OverviewAgent",
instructions="Handle general dashboard and summary questions. Always call get_dashboard_snapshot_tool first.",
tools=[get_dashboard_snapshot_tool, get_entity_ranking_tool],
model=model,
)
status_agent = Agent(
name="StatusAgent",
instructions="Handle live status and stream/field health questions.",
tools=[get_stream_status_tool, get_field_status_tool],
model=model,
)
quality_agent = Agent(
name="QualityAgent",
instructions="Handle rejection-rate and data-quality questions.",
tools=[get_rejection_stats_tool],
model=model,
)
orchestrator = Agent(
name="AnalyticsOrchestrator",
instructions=(
"Handle all user communication. Route each question to the right specialist tool "
"and synthesize its result into a plain-text answer. Do not hallucinate metrics; "
"only report what a specialist returns."
),
tools=[
overview_agent.as_tool(
tool_name="overview_expert",
tool_description="Answers general dashboard and summary questions.",
),
status_agent.as_tool(
tool_name="status_expert",
tool_description="Answers live status and stream/field health questions.",
),
quality_agent.as_tool(
tool_name="quality_expert",
tool_description="Answers rejection-rate and data-quality questions.",
),
],
)
CREATE TABLE events_queue (
event_id UUID,
entity_id UInt64,
event_type LowCardinality(String),
payload String,
created_at DateTime64(3)
) ENGINE = Kafka
SETTINGS
kafka_broker_list = 'your-namespace.servicebus.windows.net:9093',
kafka_topic_list = 'app-events',
kafka_group_name = 'clickhouse-consumer',
kafka_format = 'JSONEachRow';
CREATE MATERIALIZED VIEW events_mv TO events AS
SELECT * FROM events_queue;
azureblob://
├── ml-artifacts/
│ ├── lora-adapters/{v1, v2, v3}
│ ├── base-models/base-vlm-7b/
│ └── eval-results/v3/
├── training-data/{raw, processed, annotations}
└── inference-inputs/{year}/{month}/{day}/{entity_id}/{request_id}.bin
import great_expectations as gx
context = gx.get_context()
batch = context.sources.pandas_default.read_json("annotations_batch.jsonl")
suite = context.get_expectation_suite("annotation_schema_v1")
results = context.run_validation_operator(
"action_list_operator",
assets_to_validate=[batch],
run_id="training-v4-ingestion"
)
if not results["success"]:
raise ValueError(f"Data validation failed: {results['statistics']}")
import mlflow
from unsloth import FastVisionModel
mlflow.set_experiment("domain-lora-finetuning")
with mlflow.start_run(run_name=f"lora-v{adapter_version}") as run:
model, tokenizer = FastVisionModel.from_pretrained(
"unsloth/base-vlm-7b",
load_in_4bit=True,
)
model = FastVisionModel.get_peft_model(
model,
r=64,
lora_alpha=128,
finetune_vision_layers=True,
)
mlflow.log_params({
"base_model": "base-vlm-7b",
"lora_rank": 64,
"lora_alpha": 128,
"epochs": 3,
"dataset_version": "v4",
})
for epoch in range(epochs):
train_loss = train_one_epoch(model, dataloader)
mlflow.log_metric("train_loss", train_loss, step=epoch)
metrics = evaluate(model, eval_dataloader)
mlflow.log_metrics({"eval_f1": metrics["f1"]})
mlflow.log_artifacts(output_dir, artifact_path="lora-adapter")
mlflow.set_tag("promotion_status", "candidate")
# azure-ml-lora-job.yml
$schema: https://azuremlschemas.azureedge.net/latest/commandJob.schema.json
type: command
code: ./training/
command: >
python finetune_lora_unsloth.py
--dataset ${{inputs.training_data}}
--output-dir ${{outputs.lora_adapter}}
--lora-rank 64 --lora-alpha 128 --epochs 3
inputs:
training_data:
type: uri_folder
path: azureml://datastores/training_blob/paths/domain/v4/
outputs:
lora_adapter:
type: uri_folder
path: azureml://datastores/artifacts_blob/paths/lora-adapters/v4/
environment: azureml:vlm-unsloth-env:1
compute: azureml:gpu-cluster-nc24ads # single A100, sufficient for Unsloth LoRA fine-tuning
experiment_name: domain-lora-finetuning
from azure.ai.ml.entities import Model
model = Model(
name="domain-lora-adapter",
version="4",
path="azureml://datastores/artifacts_blob/paths/lora-adapters/v4/",
tags={
"base_model": "base-vlm-7b",
"lora_rank": "64",
"eval_f1": "0.93", # illustrative
"promotion_status": "candidate",
}
)
ml_client.models.create_or_update(model)
def run_eval(adapter_version: str, eval_dataset_version: str):
candidate = evaluate_adapter(adapter_path=..., eval_data=...)
production = evaluate_adapter(adapter_path=get_production_adapter_path(), eval_data=...)
passed = (
candidate["f1"] >= production["f1"] - 0.01
and candidate["f1"] >= MIN_F1_THRESHOLD
)
return passed, candidate
New annotation batch → Great Expectations validation
→ dataset versioned in Azure ML
→ Azure ML training job (Unsloth, single GPU, spot instance)
→ offline eval vs. production, per-category
→ [pass] tag adapter "production" in registry
→ sync adapter to local disk on the vLLM host
→ rolling restart of vLLM with new --lora-modules path
# .gitlab-ci.yml
stages:
- validate
- train
- eval
- deploy
validate-data:
stage: validate
script:
- python eval/validate_dataset.py --version $DATASET_VERSION
submit-training:
stage: train
needs: [validate-data]
script:
- az ml job create --file azure-ml-lora-job.yml
--set inputs.dataset_version=$DATASET_VERSION
--workspace-name $AML_WORKSPACE --resource-group $RESOURCE_GROUP
run-eval:
stage: eval
needs: [submit-training]
script:
- python eval/run_eval.py --version $DATASET_VERSION
- python scripts/promote_adapter.py --version $DATASET_VERSION
deploy:
stage: deploy
needs: [run-eval]
when: manual
script:
- az containerapp update --name vllm-server --resource-group $RESOURCE_GROUP
--set-env-vars LORA_ADAPTER_VERSION=$DATASET_VERSION
The agentic layer: three specialized agents, one orchestrator
Split responsibilities:
- Router / orchestrator — classifies intent and picks tools
- Retrieval specialist — fetches domain context
- Analytics specialist — emits structured facts for warehousing
Shared GPU inference means careful concurrency limits and queueing so interactive chat is not starved by batch agent runs.
The analytics backend: ClickHouse
Agent outputs that matter for product analytics land in ClickHouse: latency, tool choices, retrieved doc IDs, user outcomes. Columnar storage fits high-cardinality event streams better than forcing everything through the OLTP database.
Azure Blob Storage: the MLOps glue layer
Checkpoints, datasets, adapter artifacts, and evaluation reports live in Blob. Training jobs read/write here; inference pulls approved adapters by version. Treat blob paths as part of the API contract between training and serving.
The training pipeline: Unsloth on Azure ML
Fine-tune with Unsloth for speed/memory efficiency on Azure ML compute. Track runs in MLflow: hyperparameters, eval metrics, artifact URIs. Only promote adapters that beat a baseline on a frozen eval set.
Data validation before training
Validate schema, PII scrubbing, class balance, and leakage between train/eval before GPU minutes burn. Fail closed on validation errors.
Fine-tuning with Unsloth, tracked in MLflow
Log adapter weights to Blob with immutable version IDs. Inference configuration references those IDs explicitly—no “latest” in production without a pin.
Serving both workloads safely
- Separate queues for interactive vs batch
- Per-adapter concurrency caps
- Warm a default adapter; cold-load specialists
- Emit GPU utilization and queue depth to the same ClickHouse (or metrics stack)
- Roll back adapters by flipping a config pointer, not by redeploying the whole VM
Takeaways
One GPU can host two product surfaces if LoRA demand-loading, queue isolation, and MLOps artifact discipline are first-class. vLLM + Unsloth + Azure ML + Blob + ClickHouse is a practical MLOps shape for teams that cannot afford duplicate inference fleets yet still need specialized behaviors and analytics.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.
Capacity planning note: measure tokens/second with one adapter versus N adapters under concurrent interactive load. Demand LoRA is not free—loading cost and memory fragmentation can erase the savings of “one GPU” if product traffic ignores queue classes. Put a hard limit on simultaneous adapters and shed batch work first when interactive SLOs slip.