Home / Articles / Practical notes: Inside ARD: How the Agentic Resource Discovery Spec Actually

This article is published in English.

Practical notes: Inside ARD: How the Agentic Resource Discovery Spec Actually

Operable walkthrough of Practical notes: Inside ARD: How the Agentic Resource Discovery Spec Actually: contracts, checks, and drop-in code slots for teams shipping this pattern.

4954 words

The following notes reconstruct a practical path around “Inside ARD: How the Agentic Resource Discovery Spec Actually Works”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing. When working through the Overview stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The problem ARD is solving

The The problem ARD is stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

The mental model: describe, crawl, search, invoke

The The mental model describe stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Describing a resource: the ai-catalog.json manifest

The Describing a resource the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Describing a resource the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

https://yourdomain.com/.well-known/ai-catalog.json
{
  "specVersion": "1.0",
  "host": {
    "displayName": "Northwind Labs",
    "identifier": "northwindlabs.dev"
  },
  "entries": [
    {
      "identifier": "urn:ai:northwindlabs.dev:tools:pdf-table-extractor",
      "displayName": "PDF Table Extractor",
      "type": "application/mcp-server+json",
      "url": "https://tools.northwindlabs.dev/pdf-extractor/mcp.json",
      "description": "Extracts structured tables from scanned or digital
                      PDFs into CSV or JSON.",
      "representativeQueries": [
        "pull the line-item table out of this invoice PDF",
        "convert the tables in this scanned report into a spreadsheet"
      ]
    }
  ]
}

Identity: why the identifier looks like a URN

For the Identity why the identifier stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

The API: search, explore, and a plain list

For the The API search explore stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

{
  "query": {
    "text": "I need to digitize an invoice's line items",
    "filter": {
      "type": ["application/mcp-server+json"]
    }
  },
  "pageSize": 5
}

Federation: registries talking to registries

For the Federation registries talking to stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Federation registries talking to stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Where this actually plugs into a chatbot

When working through the Where this actually plugs stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Building it for real: a production ARD implementation on Snowflake

When working through the Building it for real stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

┌─────────────────────────────────────────────────────────────┐
│                    Streamlit UI Layer                        │
│   (Serves /.well-known/ai-catalog.json + search interface)  │
├─────────────────────────────────────────────────────────────┤
│                    API Procedures Layer                      │
│   ARD_SEARCH │ ARD_LIST_AGENTS │ ARD_EXPLORE │ ARD_GATE     │
├─────────────────────────────────────────────────────────────┤
│                 Semantic Ranking Layer                       │
│   Python UDF: TF-IDF + Cosine Similarity (scikit-learn)     │
├─────────────────────────────────────────────────────────────┤
│                    Registry Layer                            │
│   ARD_REGISTRY_ENTRIES table + ARD_AUDIT_LOG                │
├─────────────────────────────────────────────────────────────┤
│                    Ingestion Layer                           │
│   ARD_INGEST_MANIFEST (parse JSON → populate registry)      │
├─────────────────────────────────────────────────────────────┤
│                    Generation Layer                          │
│   ARD_MANIFEST_GENERATOR (DESCRIBE AGENT → ai-catalog.json) │
└─────────────────────────────────────────────────────────────┘

Layer 1: Auto-generating the manifest from live agents

When working through the Layer 1 Auto-generating the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Layer 1 Auto-generating the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

SHOW AGENTS IN SCHEMA ANALYTICS.AGENTS;
{
  "specVersion": "1.0",
  "host": {
    "displayName": "Snowflake Analytics Platform",
    "identifier": "analytics.snowflake-demo.com"
  },
  "entries": [
    {
      "identifier": "urn:ai:analytics.snowflake-demo.com:analytics:finance-agent",
      "displayName": "Finance Agent",
      "type": "application/vnd.snowflake.cortex-agent+json",
      "url": "https://zkumjrw-uib48895.snowflakecomputing.com/api/v2/cortex/agents/...",
      "description": "Finance AI analyst with expertise in ASC 606...",
      "tags": ["finance", "revenue", "ASC-606", "ARR", "bookings"],
      "capabilities": ["text-to-sql", "metric-disambiguation"],
      "representativeQueries": [
        "What was our recognized revenue last quarter?",
        "Show me ARR trend over the past 12 months"
      ],
      "trustManifest": {
        "identity": {"type": "domain-verified", "domain": "analytics.snowflake-demo.com"},
        "attestations": [
          {"type": "RBAC-governed", "detail": "FINANCE_AGENT_ROLE required"}
        ]
      }
    }
  ]
}

Layer 2: Ingesting into a searchable registry

The Layer 2 Ingesting into stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

ARD_REGISTRY_ENTRIES
├── IDENTIFIER (URN, unique)
├── DISPLAY_NAME
├── TYPE (IANA media type)
├── URL
├── DESCRIPTION
├── TAGS (ARRAY)
├── CAPABILITIES (ARRAY)
├── REPRESENTATIVE_QUERIES (ARRAY)
├── TRUST_MANIFEST (VARIANT)
├── SEARCH_TEXT (lower-cased concatenation of description + queries + tags)
├── STATUS ('ACTIVE' | 'STALE' | 'REMOVED')
└── Timestamps (INGESTED_AT, LAST_VERIFIED_AT, UPDATED_AT)

Layer 3: Semantic search — the Python UDF approach

The Layer 3 Semantic search stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.

CREATE OR REPLACE FUNCTION ANALYTICS.AGENTS.ARD_SEMANTIC_RANK(
    query_text VARCHAR,
    candidates ARRAY
)
RETURNS ARRAY
LANGUAGE PYTHON
RUNTIME_VERSION = '3.11'
PACKAGES = ('scikit-learn', 'numpy')
HANDLER = 'rank_candidates'
AS
$
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
def rank_candidates(query_text, candidates):
    if not candidates or not query_text:
        return []
    identifiers = [c['identifier'] for c in candidates]
    texts = [c.get('search_text', '') for c in candidates]
    all_texts = [query_text.lower()] + [t.lower() for t in texts]
    vectorizer = TfidfVectorizer(
        ngram_range=(1, 3),
        max_features=5000,
        stop_words='english',
        sublinear_tf=True
    )
    try:
        tfidf_matrix = vectorizer.fit_transform(all_texts)
    except ValueError:
        return [{'identifier': id, 'score': 0} for id in identifiers]
    similarities = cosine_similarity(tfidf_matrix[0:1], tfidf_matrix[1:])[0]
    results = [
        {'identifier': id, 'score': round(float(sim) * 100, 1)}
        for id, sim in zip(identifiers, similarities)
    ]
    results.sort(key=lambda x: x['score'], reverse=True)
    return results
$;

Layer 4: The invocation gate — RBAC before execution

The Layer 4 The invocation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Layer 4 The invocation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

CALL ARD_INVOCATION_GATE(
    'urn:ai:analytics.snowflake-demo.com:analytics:finance-agent',
    'ACCOUNTADMIN'
)
-- Returns: {"authorized": true, "agentFqn": "ANALYTICS.AGENTS.FINANCE_AGENT", ...}

CALL ARD_INVOCATION_GATE(
    'urn:ai:analytics.snowflake-demo.com:analytics:finance-agent',
    'PUBLIC'
)
-- Returns: {"authorized": false, "reason": "Role PUBLIC lacks FINANCE_AGENT_ROLE grant."}

Layer 5: The Streamlit manifest server

For the Layer 5 The Streamlit stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

manifest = get_manifest()
st.code(json.dumps(manifest, indent=2), language="json")
st.download_button("Download", json.dumps(manifest, indent=2), "ai-catalog.json")
query = st.text_input("Query", placeholder="I need to analyze quarterly revenue")
cap_filter = st.selectbox("Capability", [None, "text-to-sql", "multi-tool-routing"])
if st.button("Search"):
    results = search_registry(query, filters)
    for entry in results["results"]:
        st.expander(f"{entry['displayName']} — Score: {entry['score']}")
stats = get_registry_stats()
# Shows: 4 entries, 18 tags across 4 agents, 3 capability types

Layer 6: The end-to-end test harness

For the Layer 6 The end-to-end stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

Test 1: MANIFEST_GENERATION
  → Calls ARD_MANIFEST_GENERATOR(), asserts specVersion = "1.0"
     and entries array is non-empty

Test 2: MANIFEST_INGESTION
  → Calls ARD_INGEST_MANIFEST(manifest), asserts status = "SUCCESS"
     and entries_ingested > 0
Test 3: SEARCH_FINANCE_QUERY
  → Searches "What was our revenue last quarter?"
  → Asserts top result identifier contains "finance"
Test 4: SEARCH_CHURN_QUERY
  → Searches "Which customers are likely to churn?"
  → Asserts top result identifier contains "cs"
Test 5: SEARCH_WITH_FILTER
  → Searches "pipeline forecast" with capabilities filter ["text-to-sql"]
  → Asserts results > 0 (filter applied correctly)
Test 6: LIST_AGENTS
  → Calls ARD_LIST_AGENTS(1, 10)
  → Asserts pagination.totalEntries > 0
Test 7: EXPLORE_FACETS
  → Calls ARD_EXPLORE()
  → Asserts facets.tags is not null and totalEntries > 0
Test 8: GATE_AUTHORIZED
  → Calls ARD_INVOCATION_GATE(finance URN, "ACCOUNTADMIN")
  → Asserts authorized = true
Test 9: GATE_UNAUTHORIZED
  → Calls ARD_INVOCATION_GATE(finance URN, "PUBLIC")
  → Asserts authorized = false
Test 10: HEALTH_CHECK
  → Calls ARD_HEALTH_CHECK()
  → Asserts status = "COMPLETE"
{
  "summary": {
    "total_tests": 10,
    "passed": 10,
    "failed": 0,
    "success_rate": "100.0%"
  },
  "tests": [...],
  "timestamp": "2026-06-18T..."
}

Production hardening: what breaks and how we fixed it

For the Production hardening what breaks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Production hardening what breaks stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

The Streamlit manifest server — serving ARD over HTTP

When working through the The Streamlit manifest server stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

Deployment

When working through the Deployment stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

CREATE STAGE IF NOT EXISTS ANALYTICS.AGENTS.STREAMLIT_STAGE
    ENCRYPTION = (TYPE = 'SNOWFLAKE_SSE');

-- Upload source (via COPY INTO from temp table)
COPY INTO @ANALYTICS.AGENTS.STREAMLIT_STAGE/ard_manifest_app/streamlit_app.py
FROM (SELECT content FROM _STREAMLIT_SRC)
FILE_FORMAT = (TYPE = CSV COMPRESSION = NONE ...)
SINGLE = TRUE OVERWRITE = TRUE;
CREATE OR REPLACE STREAMLIT ANALYTICS.AGENTS.ARD_MANIFEST_SERVER
    ROOT_LOCATION = '@ANALYTICS.AGENTS.STREAMLIT_STAGE/ard_manifest_app'
    MAIN_FILE = '/streamlit_app.py'
    QUERY_WAREHOUSE = COMPUTE_WH;

The full Streamlit source

When working through the The full Streamlit source stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

import streamlit as st
import json
from snowflake.snowpark.context import get_active_session
st.set_page_config(page_title="ARD Manifest Server", layout="wide")
session = get_active_session()
@st.cache_data(ttl=300)
def get_manifest():
    result = session.sql("CALL ANALYTICS.AGENTS.ARD_MANIFEST_GENERATOR()").collect()
    return json.loads(result[0][0])
@st.cache_data(ttl=300)
def search_registry(query, filters=None):
    safe_query = query.replace("'", "''")
    if filters:
        filter_json = json.dumps(filters).replace("'", "''")
        sql = f"CALL ANALYTICS.AGENTS.ARD_SEARCH('{safe_query}', PARSE_JSON('{filter_json}'))"
    else:
        sql = f"CALL ANALYTICS.AGENTS.ARD_SEARCH('{safe_query}')"
    result = session.sql(sql).collect()
    return json.loads(result[0][0])
@st.cache_data(ttl=300)
def get_registry_stats():
    result = session.sql("CALL ANALYTICS.AGENTS.ARD_EXPLORE()").collect()
    return json.loads(result[0][0])
tab1, tab2, tab3, tab4 = st.tabs([
    "ai-catalog.json", "Search", "Explorer", "API Docs"
])

Tab 1: The raw manifest

When working through the Tab 1 The raw stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

with tab1:
    st.markdown("## /.well-known/ai-catalog.json")
    manifest = get_manifest()
    c1, c2, c3 = st.columns(3)
    c1.metric("Spec Version", manifest.get("specVersion", "?"))
    c2.metric("Host", manifest.get("host", {}).get("identifier", "?"))
    c3.metric("Entries", len(manifest.get("entries", [])))
    st.code(json.dumps(manifest, indent=2), language="json")
    st.download_button(
        "Download ai-catalog.json",
        json.dumps(manifest, indent=2),
        "ai-catalog.json",
        "application/json"
    )

Tab 2: Interactive semantic search

When working through the Tab 2 Interactive semantic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

with tab2:
    st.markdown("## POST /search")
    query = st.text_input("Query", placeholder="e.g., I need to analyze quarterly revenue")
    cap_filter = st.selectbox("Capability", [None, "text-to-sql", "multi-tool-routing"])
    if st.button("Search", type="primary") and query:
        filters = {"capabilities": [cap_filter]} if cap_filter else None
        results = search_registry(query, filters)
        st.markdown(f"### {results['resultCount']} results")
        st.caption(f"Method: {results.get('method', 'keyword')}")
        for i, entry in enumerate(results.get("results", [])):
            with st.expander(f"#{i+1} {entry['displayName']} — Score: {entry['score']}"):
                st.markdown(f"**ID:** `{entry['identifier']}`")
                st.markdown(f"**URL:** `{entry.get('url', 'N/A')}`")
                st.markdown(f"**Tags:** {', '.join(entry.get('tags', []))}")
                st.markdown(f"**Capabilities:** {', '.join(entry.get('capabilities', []))}")
                if entry.get("representativeQueries"):
                    for q in entry["representativeQueries"]:
                        st.markdown(f"- _{q}_")

Tab 3: Faceted exploration

When working through the Tab 3 Faceted exploration stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

with tab3:
    st.markdown("## POST /explore")
    stats = get_registry_stats()
    st.metric("Active Entries", stats.get("totalEntries", 0))
    e1, e2, e3 = st.columns(3)
    with e1:
        st.markdown("### Types")
        for f in stats.get("facets", {}).get("type", []):
            st.markdown(f"- `{f['value']}` ({f['count']})")
    with e2:
        st.markdown("### Tags")
        for f in stats.get("facets", {}).get("tags", []):
            st.markdown(f"- `{f['value']}` ({f['count']})")
    with e3:
        st.markdown("### Capabilities")
        for f in stats.get("facets", {}).get("capabilities", []):
            st.markdown(f"- `{f['value']}` ({f['count']})")

Tab 4: API reference

When working through the Tab 4 API reference stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.

with tab4:
    st.markdown("""
    | ARD Endpoint | Procedure | Description |
    |---|---|---|
    | `GET /.well-known/ai-catalog.json` | `ARD_MANIFEST_GENERATOR()` | Live manifest |
    | `POST /search` | `ARD_SEARCH(query, filters)` | Semantic search |
    | `POST /explore` | `ARD_EXPLORE()` | Faceted browse |
    | `GET /agents` | `ARD_LIST_AGENTS(page, size)` | Paginated list |
    | Gate | `ARD_INVOCATION_GATE(urn, role)` | RBAC check |
Scoring: TF-IDF + cosine similarity (scikit-learn), 0-100 scale.
    Identity: urn:ai:<domain>:<namespace>:<agent-name>
    """)

Accessing the app

When working through the Accessing the app stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Accessing the app stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Live testing results

The Live testing results stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Search: “you need to analyze our quarterly revenue”

The Search you need to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Results: 2 found | Method: tfidf-cosine-similarity
#1 Finance Agent — Score: 3.5
   ID: urn:ai:analytics.snowflake-demo.com:analytics:finance-agent
   Tags: finance, revenue, ASC-606, ARR, bookings
   Capabilities: text-to-sql, metric-disambiguation#2 Executive Agent — Score: 1.5
   ID: urn:ai:analytics.snowflake-demo.com:analytics:executive-agent
   Tags: executive, cross-domain, orchestrator, KPI
   Capabilities: text-to-sql, metric-disambiguation, multi-tool-routing

Search: “Which customers are likely to churn?”

The Search Which customers are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Results: 1 found | Method: tfidf-cosine-similarity
#1 CS Agent — Score: 10.5
   ID: urn:ai:analytics.snowflake-demo.com:analytics:cs-agent
   Tags: customer-success, health-score, churn, NPS, CSAT

Search: “pipeline forecast” with filter capabilities=[“text-to-sql”]

The Search pipeline forecast with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Results: 2 found (filtered from 4 total)
#1 Sales Agent — Score: 8.2
#2 Finance Agent — Score: 2.1

Explorer facets

The Explorer facets stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

Total Active Entries: 4
Types:
  - application/vnd.snowflake.cortex-agent+json (4)
Tags (18 total):
  - bookings (2), finance (1), revenue (1), ASC-606 (1), ARR (1),
    sales (1), pipeline (1), forecast (1), win-rate (1),
    customer-success (1), health-score (1), churn (1), NPS (1),
    CSAT (1), executive (1), cross-domain (1), orchestrator (1), KPI (1)
Capabilities:
  - text-to-sql (7), metric-disambiguation (7), multi-tool-routing (1)

Invocation gate test

The Invocation gate test stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

CALL ARD_INVOCATION_GATE('urn:ai:...finance-agent', 'ACCOUNTADMIN')
→ {"authorized": true, "reason": "Role ACCOUNTADMIN is authorized..."}
CALL ARD_INVOCATION_GATE('urn:ai:...finance-agent', 'PUBLIC')
→ {"authorized": false, "reason": "Role PUBLIC lacks FINANCE_AGENT_ROLE grant."}

The monitoring layer

The The monitoring layer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.

What this means practically

The What this means practically stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The What this means practically stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

"I need to analyze our quarterly revenue figures"
Finance Agent — Score: 15.8
Executive Agent — Score: 3.5
Sales Agent — Score: 3.2

Tools for implementers

For the Tools for implementers stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary.

What’s next

For the What s next stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Getting started

For the Getting started stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Getting started stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

git clone https://github.com/satish/ard-registry.git
cd ard-registry
-- In Snowsight, execute these SQL files in order:
sql/01_infrastructure.sql        -- Creates stage, tables, audit log
sql/02_manifest_generator.sql    -- Reads agent metadata → ARD manifest
sql/03_ingest.sql                -- Parses manifest → searchable registry
sql/04_semantic_rank.sql         -- Python UDF (TF-IDF + cosine similarity)
sql/05_search.sql                -- Semantic search endpoint
sql/06_list_and_explore.sql      -- List + explore endpoints
sql/07_invocation_gate.sql       -- RBAC authorization gate
sql/08_monitoring.sql            -- Scheduled refresh + health check
sql/10_e2e_test.sql              -- Test harness-- Then ingest and verify:
EXECUTE IMMEDIATE $
DECLARE v_manifest VARIANT; v_result VARIANT;
BEGIN
    CALL ANALYTICS.AGENTS.ARD_MANIFEST_GENERATOR() INTO v_manifest;
    CALL ANALYTICS.AGENTS.ARD_INGEST_MANIFEST(:v_manifest) INTO v_result;
    RETURN :v_result;
END;
$;CALL ANALYTICS.AGENTS.ARD_END_TO_END_TEST();
-- Expected: 10/10 PASS (100%)

Operational checklist

For the Operational checklist stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Write a short runbook: how to rotate keys, how to drain the queue, how to roll back the last ingest.

Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for ba61be007942: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.