Home / Articles / Practical notes: What If Your Agent’s Embedding Model Is About to Sunset?

This article is published in English.

Practical notes: What If Your Agent’s Embedding Model Is About to Sunset?

Operable walkthrough of Practical notes: What If Your Agent’s Embedding Model Is About to Sunset?: contracts, checks, and drop-in code slots for teams shipping this pattern.

4214 words

Use this as an operator-facing rebuild of the ideas in “What If Your Agent’s Embedding Model Is About to Sunset?”: clear stages, ordered code slots, and recovery notes that survive a handoff. The Overview stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Backfill
   ↓
Validate
   ↓
Canary
   ↓
Cut over
   ↓
Soak
   ↓
Clean up

The Problem

For the The Problem stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

1. Historical data

For the 1 Historical data stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

2. Live data

For the 2 Live data stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the 2 Live data stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

3. Production traffic

When working through the 3 Production traffic stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Historical data      → distributed backfill
Live changes         → async dual-write
Production traffic   → canary + guardrails

The First Design Rule: Version the Embeddings

When working through the The First Design Rule stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

record
 ├── namespace
 ├── knowledge-base version
 ├── embed_version
 └── embedding
CREATE TABLE semantic_cache (
    id            BIGSERIAL,
    embed_version TEXT NOT NULL,
    namespace     TEXT NOT NULL,
    query_hash    BYTEA NOT NULL,
    embedding     vector(...) NOT NULL,
    response      JSONB NOT NULL,
    created_at    TIMESTAMPTZ NOT NULL DEFAULT now(),
    last_hit_at   TIMESTAMPTZ NOT NULL DEFAULT now(),
    hit_count     INT NOT NULL DEFAULT 0,
    expires_at    TIMESTAMPTZ NOT NULL,
    PRIMARY KEY (embed_version, id)
) PARTITION BY LIST (embed_version);
CREATE TABLE semantic_cache_v1
    PARTITION OF semantic_cache
    FOR VALUES IN ('gemini-embedding-001');
CREATE TABLE semantic_cache_v2
    PARTITION OF semantic_cache
    FOR VALUES IN ('text-embedding-3-large');
                 semantic_cache
                      │
          ┌───────────┴───────────┐
          │                       │
     embed_version=V1        embed_version=V2
          │                       │
      V1 vectors              V2 vectors

Phase 0 — Backfill the V2 Representation

When working through the Phase 0 Backfill the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Phase 0 Backfill the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

SELECT * FROM semantic_cache;

0.1 Create the V2 partition

The 0 1 Create the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

V1 partition
   │
   │ still serving
   ▼
V2 partition
   │
   │ being populated
   ▼
migration control table

0.2 Split the corpus into claimable ranges

The 0 2 Split the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

┌─────────────────────────────────────────┐
│ Migration control table                 │
├──────────────┬─────────────┬────────────┤
│ Range        │ Status      │ Worker     │
├──────────────┼─────────────┼────────────┤
│ 1 - 50K      │ complete    │ worker-1   │
│ 50K - 100K   │ processing  │ worker-2   │
│ 100K - 150K  │ pending     │ worker-3   │
│ 150K - 200K  │ pending     │ worker-4   │
└──────────────┴─────────────┴────────────┘
SELECT ...
FOR UPDATE SKIP LOCKED;

0.3 Fetch with keyset pagination

The 0 3 Fetch with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 0 3 Fetch with stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

SELECT ...
FROM semantic_cache_v1
WHERE id > :last_id
  AND id <= :range_end
ORDER BY id
LIMIT :batch_size;
Migration range
    ≈ scheduling/checkpoint boundary
Embedding batch
    ≈ model/provider throughput boundary

0.4 Generate embeddings through a dedicated pool

For the 0 4 Generate embeddings stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Embedding capacity
                           │
            ┌──────────────┴──────────────┐
            │                             │
     online traffic                 migration traffic
            │                             │
            ▼                             ▼
       production path              dedicated pool

0.5 Buffer and throttle writes

For the 0 5 Buffer and stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

existing records
      │
      ▼
embedding batches
      │
      ▼
buffer
      │
      ▼
throttled writes
      │
      ▼
V2 partition

0.6 Validate each completed range

For the 0 6 Validate each stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the 0 6 Validate each stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

write range
    │
    ▼
validate
    │
 ┌──┴────┐
PASS    FAIL
 │        │
 ▼        ▼
checkpoint retry
complete   │
           ▼
      persistent failure
           │
           ▼
        halt + alert

0.7 Build the V2 indexes

When working through the 0 7 Build the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

V1
├── existing data
└── existing index
V2
├── migrated data
└── new index

Phase 1 — Validate V2 Before Production

When working through the Phase 1 Validate V2 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

row_count(V1) == row_count(V2)
retrieval_quality(V2) < retrieval_quality(V1)
Golden-set query
       │
       ├──────────────► V1 retrieval
       │
       └──────────────► V2 retrieval
                              │
                              ▼
                       quality comparison
Golden-set evaluation
          │
      ┌───┴───┐
     PASS    FAIL
      │        │
      ▼        ▼
   Canary     Stop
              tune

Phase 2 — Canary the New Path

When working through the Phase 2 Canary the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Phase 2 Canary the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

1% → 10% → 50% → 100%

2.1 The request path

The 2 1 The request stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Incoming query
      │
      ▼
   L1 cache
      │
   ┌──┴───┐
  HIT    MISS
   │       │
   ▼       ▼
return   V2 embedding
           │
           ▼
       V2 retrieval
           │
           ▼
      quality check
        ┌──┴───┐
      strong  weak
        │       │
        ▼       ▼
       RAG   fallback V1
        │       │
        └──┬────┘
           ▼
        response

Version-Aware Retrieval

The Version-Aware Retrieval stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

namespace
+
knowledge-base version
+
embedding version
namespace   = customer-A
kb_version  = 42
embed_version = V2

2.2 Quality Guardrails

The 2 2 Quality Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 2 2 Quality Guardrails stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

System health

For the System health stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Retrieval quality

For the Retrieval quality stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Migration health

For the Migration health stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the Migration health stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

1%
 │
 ▼
guardrail
 │ PASS
 ▼
10%
 │
 ▼
guardrail
 │ PASS
 ▼
50%
 │
 ▼
guardrail
 │ PASS
 ▼
100%

2.3 Rollback Must Be a Configuration Change

When working through the 2 3 Rollback Must stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

stop traffic
restore data
redeploy services
rebuild indexes
hope
active embed_version = V2
             │
             ▼
        config change
             │
             ▼
active embed_version = V1

The Race During Backfill: What About Live Updates?

When working through the The Race During Backfill stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

10:00  worker reads record A
10:01  record A is updated
10:02  worker writes V2 generated from the older content
                application write
                        │
                ┌───────┴───────┐
                │               │
                ▼               ▼
             V1 write       async queue
                                │
                                ▼
                             V2 write

Phase 3 — Flip the Read Path

When working through the Phase 3 Flip the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Phase 3 Flip the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Before:
reads → V1
After:reads → V2

Phase 4 — Post-Flip Soak

The Phase 4 Post-Flip Soak stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Phase 5 — Final Validation and Cleanup

The Phase 5 Final Validation stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

1. Stop V1 dual-write

The 1 Stop V1 dual-write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 1 Stop V1 dual-write stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

2. Run final validation

For the 2 Run final validation stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

V2
 │
 ▼
final validation
 │
 ├── FAIL → stop cleanup
 │
 └── PASS
        │
        ▼
    continue cleanup

3. Delete the V1 representation

For the 3 Delete the V1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Backfill
  ↓
Validate
  ↓
Canary
  ↓
100% V2
  ↓
Soak
  ↓
Final validation
  ↓
Delete V1

The Complete Lifecycle

For the The Complete Lifecycle stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For the The Complete Lifecycle stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

EMBEDDING MIGRATION
                           │
                           ▼
                    Create V2 storage
                           │
                           ▼
                      Backfill V2
                 work stealing + retries
                           │
                           ▼
                    Range validation
                           │
                           ▼
                    Build V2 indexes
                           │
                           ▼
                  Golden-set evaluation
                           │
                    ┌──────┴──────┐
                   FAIL          PASS
                    │              │
                    ▼              ▼
                stop/tune       Canary
                              1% → 10% → 50%
                                       │
                                       ▼
                                  Guardrails
                                       │
                                       ▼
                                  100% V2
                                       │
                                       ▼
                                  Read flip
                                       │
                                       ▼
                                   V2 soak
                                       │
                                       ▼
                                Final validation
                                       │
                                       ▼
                                  Drop V1
Live production writes
                           │
                           ▼
                       V1 write
                           │
                           └──── async dual-write → V2
Migration safety
                       │
        ┌──────────────┼──────────────┐
        ▼              ▼              ▼
       Data          Quality         Traffic
       safety         safety         safety
        │              │              │
   checkpoints      golden set      canary
   retries          guardrails      fallback
   validation       evaluation      rollback

The Lessons That Generalize

When working through the The Lessons That Generalize stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

1. Make embedding version a first-class concept

When working through the 1 Make embedding version stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

2. Backfill like a database engineer

When working through the 2 Backfill like a stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

3. Separate data migration from traffic migration

V2 exists
   ≠
V2 is trusted
   ≠
V2 is serving production

4. Treat retrieval quality as part of correctness

Data correctness
      +
Retrieval quality
      +
Production health

5. Keep the old path as the safety net

6. Make rollback boring

change configuration
change code + redeploy + restore state

7. Delete last

The Bottom Line

model identity
      +
vector storage
      +
traffic routing
Partition
   ↓
Backfill
   ↓
Validate
   ↓
Canary
   ↓
Cut over
   ↓
Soak
   ↓
Clean up

Operational checklist