This article is published in English.
Practical notes: Building and Serving a Custom Model with Azure ML, Then Wiring
Operable walkthrough of Practical notes: Building and Serving a Custom Model with Azure ML, Then Wiring: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Building and Serving a Custom Model with Azure ML, Then Wiring It Into a Foundry Agent. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The mental model first
When working through the The mental model first stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Prerequisites
When working through the Prerequisites stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
pip install azure-ai-ml azure-ai-projects azure-identity
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential
ml_client = MLClient(
DefaultAzureCredential(),
subscription_id="<subscription-id>",
resource_group_name="<resource-group>",
workspace_name="<aml-workspace-name>",
)
Step 1: train the model
When working through the Step 1 train the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Step 1 train the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
from azure.ai.ml import command, Input, Output
train_job = command(
code="./src",
command="python train.py --data ${{inputs.training_data}} --model_output ${{outputs.model_output}}",
inputs={"training_data": Input(type="uri_folder", path="azureml://datastores/workspaceblobstore/paths/churn-training/")},
outputs={"model_output": Output(type="uri_folder")},
environment="azureml://registries/azureml/environments/sklearn-1.5/labels/latest",
compute="cpu-cluster",
display_name="churn-model-training",
)
returned_job = ml_client.jobs.create_or_update(train_job)
ml_client.jobs.stream(returned_job.name)
Step 2: register the trained model
The Step 2 register the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
from azure.ai.ml.entities import Model
from azure.ai.ml.constants import AssetTypes
model = ml_client.models.create_or_update(
Model(
path=f"azureml://jobs/{returned_job.name}/outputs/model_output",
name="churn-classifier",
type=AssetTypes.MLFLOW_MODEL,
description="Customer churn classifier, trained on 18 months of account history.",
)
)
Step 3: deploy it behind a managed endpoint
The Step 3 deploy it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment
endpoint = ManagedOnlineEndpoint(name="churn-endpoint", auth_mode="key")
ml_client.online_endpoints.begin_create_or_update(endpoint).result()
deployment = ManagedOnlineDeployment(
name="blue",
endpoint_name="churn-endpoint",
model=model,
instance_type="Standard_DS3_v2",
instance_count=1,
)
ml_client.online_deployments.begin_create_or_update(deployment).result()
endpoint.traffic = {"blue": 100}
ml_client.online_endpoints.begin_create_or_update(endpoint).result()
Step 4: confirm it works before anything else touches it
The Step 4 confirm it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
import json
test_input = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[14, 89.50, 3]]}}
response = ml_client.online_endpoints.invoke(
endpoint_name="churn-endpoint",
request_file=None,
deployment_name="blue",
input_data=json.dumps(test_input),
)
print(response)
The Step 4 confirm it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Step 5: wrap the endpoint as a Foundry agent function tool
For the Step 5 wrap the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
import os
import requests
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import PromptAgentDefinition, Tool, FunctionTool
from azure.identity import DefaultAzureCredential
def get_churn_score(tenure_months: int, monthly_spend: float, support_tickets: int) -> dict:
payload = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[tenure_months, monthly_spend, support_tickets]]}}
resp = requests.post(
"https://churn-endpoint.<region>.inference.ml.azure.com/score",
headers={"Authorization": f"Bearer {os.environ['AML_ENDPOINT_KEY']}", "Content-Type": "application/json"},
json=payload,
timeout=10,
)
resp.raise_for_status()
return {"churn_probability": resp.json()[0]}
func_tool = FunctionTool(
name="get_churn_score",
description="Predict churn probability for a customer given tenure, spend, and support ticket history.",
parameters={
"type": "object",
"properties": {
"tenure_months": {"type": "integer", "description": "How many months the customer has been active."},
"monthly_spend": {"type": "number", "description": "Average monthly spend in dollars."},
"support_tickets": {"type": "integer", "description": "Number of support tickets in the last 90 days."},
},
"required": ["tenure_months", "monthly_spend", "support_tickets"],
"additionalProperties": False,
},
strict=True,
)
project = AIProjectClient(endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"], credential=DefaultAzureCredential())
tools: list[Tool] = [func_tool]
agent = project.agents.create_version(
agent_name="retention-agent",
definition=PromptAgentDefinition(
model="gpt-4.1-mini",
instructions="Help the team assess churn risk. Call get_churn_score whenever specific customer numbers are provided.",
tools=tools,
),
)
Step 6: run it end to end
For the Step 6 run it stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
import json
from openai.types.responses.response_input_param import FunctionCallOutput
openai_client = project.get_openai_client()
conversation = openai_client.conversations.create()
response = openai_client.responses.create(
input="A customer's been with us 14 months, spends about $90/month, and filed 3 tickets recently. Churn risk?",
conversation=conversation.id,
extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}},
)
for item in response.output:
if item.type == "function_call" and item.name == "get_churn_score":
result = get_churn_score(**json.loads(item.arguments))
follow_up = openai_client.responses.create(
input=[FunctionCallOutput(type="function_call_output", call_id=item.call_id, output=json.dumps(result))],
conversation=conversation.id,
extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}},
)
print(follow_up.output_text)
Where this fits in the bigger picture
For the Where this fits in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Where this fits in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Production considerations before you commit
When working through the Production considerations before you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Where this leaves you
When working through the Where this leaves you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
References
When working through the References stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the References stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 2117731bf26b: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.