Home / Articles / Practical notes: Building and Serving a Custom Model with Azure ML, Then Wiring

This article is published in English.

Practical notes: Building and Serving a Custom Model with Azure ML, Then Wiring

Operable walkthrough of Practical notes: Building and Serving a Custom Model with Azure ML, Then Wiring: contracts, checks, and drop-in code slots for teams shipping this pattern.

1928 words

This walkthrough rebuilds the path from raw materials to a working system for: Building and Serving a Custom Model with Azure ML, Then Wiring It Into a Foundry Agent. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The mental model first

When working through the The mental model first stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Prerequisites

When working through the Prerequisites stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

pip install azure-ai-ml azure-ai-projects azure-identity
from azure.ai.ml import MLClient
from azure.identity import DefaultAzureCredential

ml_client = MLClient(
    DefaultAzureCredential(),
    subscription_id="<subscription-id>",
    resource_group_name="<resource-group>",
    workspace_name="<aml-workspace-name>",
)

Step 1: train the model

When working through the Step 1 train the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Step 1 train the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

from azure.ai.ml import command, Input, Output

train_job = command(
    code="./src",
    command="python train.py --data ${{inputs.training_data}} --model_output ${{outputs.model_output}}",
    inputs={"training_data": Input(type="uri_folder", path="azureml://datastores/workspaceblobstore/paths/churn-training/")},
    outputs={"model_output": Output(type="uri_folder")},
    environment="azureml://registries/azureml/environments/sklearn-1.5/labels/latest",
    compute="cpu-cluster",
    display_name="churn-model-training",
)
returned_job = ml_client.jobs.create_or_update(train_job)
ml_client.jobs.stream(returned_job.name)

Step 2: register the trained model

The Step 2 register the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

from azure.ai.ml.entities import Model
from azure.ai.ml.constants import AssetTypes

model = ml_client.models.create_or_update(
    Model(
        path=f"azureml://jobs/{returned_job.name}/outputs/model_output",
        name="churn-classifier",
        type=AssetTypes.MLFLOW_MODEL,
        description="Customer churn classifier, trained on 18 months of account history.",
    )
)

Step 3: deploy it behind a managed endpoint

The Step 3 deploy it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment

endpoint = ManagedOnlineEndpoint(name="churn-endpoint", auth_mode="key")
ml_client.online_endpoints.begin_create_or_update(endpoint).result()
deployment = ManagedOnlineDeployment(
    name="blue",
    endpoint_name="churn-endpoint",
    model=model,
    instance_type="Standard_DS3_v2",
    instance_count=1,
)
ml_client.online_deployments.begin_create_or_update(deployment).result()
endpoint.traffic = {"blue": 100}
ml_client.online_endpoints.begin_create_or_update(endpoint).result()

Step 4: confirm it works before anything else touches it

The Step 4 confirm it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

import json

test_input = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[14, 89.50, 3]]}}
response = ml_client.online_endpoints.invoke(
    endpoint_name="churn-endpoint",
    request_file=None,
    deployment_name="blue",
    input_data=json.dumps(test_input),
)
print(response)

The Step 4 confirm it stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Step 5: wrap the endpoint as a Foundry agent function tool

For the Step 5 wrap the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

import os
import requests
from azure.ai.projects import AIProjectClient
from azure.ai.projects.models import PromptAgentDefinition, Tool, FunctionTool
from azure.identity import DefaultAzureCredential

def get_churn_score(tenure_months: int, monthly_spend: float, support_tickets: int) -> dict:
    payload = {"input_data": {"columns": ["tenure_months", "monthly_spend", "support_tickets"], "data": [[tenure_months, monthly_spend, support_tickets]]}}
    resp = requests.post(
        "https://churn-endpoint.<region>.inference.ml.azure.com/score",
        headers={"Authorization": f"Bearer {os.environ['AML_ENDPOINT_KEY']}", "Content-Type": "application/json"},
        json=payload,
        timeout=10,
    )
    resp.raise_for_status()
    return {"churn_probability": resp.json()[0]}
func_tool = FunctionTool(
    name="get_churn_score",
    description="Predict churn probability for a customer given tenure, spend, and support ticket history.",
    parameters={
        "type": "object",
        "properties": {
            "tenure_months": {"type": "integer", "description": "How many months the customer has been active."},
            "monthly_spend": {"type": "number", "description": "Average monthly spend in dollars."},
            "support_tickets": {"type": "integer", "description": "Number of support tickets in the last 90 days."},
        },
        "required": ["tenure_months", "monthly_spend", "support_tickets"],
        "additionalProperties": False,
    },
    strict=True,
)
project = AIProjectClient(endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"], credential=DefaultAzureCredential())
tools: list[Tool] = [func_tool]
agent = project.agents.create_version(
    agent_name="retention-agent",
    definition=PromptAgentDefinition(
        model="gpt-4.1-mini",
        instructions="Help the team assess churn risk. Call get_churn_score whenever specific customer numbers are provided.",
        tools=tools,
    ),
)

Step 6: run it end to end

For the Step 6 run it stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

import json
from openai.types.responses.response_input_param import FunctionCallOutput

openai_client = project.get_openai_client()
conversation = openai_client.conversations.create()
response = openai_client.responses.create(
    input="A customer's been with us 14 months, spends about $90/month, and filed 3 tickets recently. Churn risk?",
    conversation=conversation.id,
    extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}},
)
for item in response.output:
    if item.type == "function_call" and item.name == "get_churn_score":
        result = get_churn_score(**json.loads(item.arguments))
        follow_up = openai_client.responses.create(
            input=[FunctionCallOutput(type="function_call_output", call_id=item.call_id, output=json.dumps(result))],
            conversation=conversation.id,
            extra_body={"agent_reference": {"name": agent.name, "type": "agent_reference"}},
        )
        print(follow_up.output_text)

Where this fits in the bigger picture

For the Where this fits in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Where this fits in stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Production considerations before you commit

When working through the Production considerations before you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Where this leaves you

When working through the Where this leaves you stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

References

When working through the References stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the References stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 2117731bf26b: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.