This article is published in English.
Practical notes: PostgreSQL + pgvector and SQL Server 2025 as Vector Stores
Operable walkthrough of Practical notes: PostgreSQL + pgvector and SQL Server 2025 as Vector Stores: contracts, checks, and drop-in code slots for teams shipping this pattern.
The following notes reconstruct a practical path around “ PostgreSQL + pgvector and SQL Server 2025 as Vector Stores for RAG — A Practitioner’s Guide”. Emphasis stays on contracts, checks, and drop-in code placeholders rather than motivational framing.
The Vector Database Landscape in 2025
When working through the The Vector Database Landscape stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
-- Find the most relevant chunks, but only from documents
-- belonging to enterprise-tier customers — a single SQL query
SELECT c.content, 1 - (c.embedding <=> query_vec) AS score
FROM rag_chunks c
JOIN documents d ON d.filename = c.source
JOIN customers cu ON cu.id = d.customer_id
WHERE cu.tier = 'enterprise'
ORDER BY score DESC
LIMIT 5;
The Full Stack
When working through the The Full Stack stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Part 1 — PostgreSQL 18 + pgvector
When working through the Part 1 PostgreSQL 18 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Installing pgvector on Windows
When working through the Installing pgvector on Windows stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
set "PGROOT=C:\Program Files\PostgreSQL\18"
cd %TEMP%
git clone --branch v0.8.0 https://github.com/pgvector/pgvector.git
cd pgvector
nmake /F Makefile.win
nmake /F Makefile.win install
docker run -d -p 5432:5432 -e POSTGRES_PASSWORD=postgres --name pgvector pgvector/pgvector:pg18
conda install -c conda-forge pgvector
Creating Tables in pgAdmin
When working through the Creating Tables in pgAdmin stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Creating Tables in pgAdmin stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
-- Enable the pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Main chunks table with VECTOR(768) column
CREATE TABLE IF NOT EXISTS rag_chunks (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
source TEXT NOT NULL,
chunk_index INTEGER NOT NULL,
content TEXT NOT NULL,
file_hash TEXT,
ingested_at TIMESTAMPTZ DEFAULT NOW(),
embedding vector(768) -- pgvector native type
);
-- HNSW index for approximate cosine similarity search
CREATE INDEX IF NOT EXISTS idx_rag_chunks_embedding
ON rag_chunks USING hnsw (embedding vector_cosine_ops);
-- Source filter index
CREATE INDEX IF NOT EXISTS idx_rag_chunks_source
ON rag_chunks (source);
-- Staleness registry
CREATE TABLE IF NOT EXISTS rag_staleness (
doc_name TEXT PRIMARY KEY,
file_hash TEXT NOT NULL,
chunk_count INTEGER,
ingested_at TIMESTAMPTZ DEFAULT NOW(),
version INTEGER DEFAULT 1
);
-- CDC chunk registry
CREATE TABLE IF NOT EXISTS rag_chunk_registry (
doc_name TEXT NOT NULL,
chunk_hash TEXT NOT NULL,
chunk_id TEXT NOT NULL,
PRIMARY KEY (doc_name, chunk_hash)
);
-- Conversation sessions
CREATE TABLE IF NOT EXISTS rag_sessions (
session_id TEXT PRIMARY KEY,
created_at TIMESTAMPTZ DEFAULT NOW(),
updated_at TIMESTAMPTZ DEFAULT NOW(),
model TEXT,
embed_model TEXT,
turn_count INTEGER DEFAULT 0
);
-- Conversation turns
CREATE TABLE IF NOT EXISTS rag_turns (
id SERIAL PRIMARY KEY,
session_id TEXT REFERENCES rag_sessions(session_id) ON DELETE CASCADE,
role TEXT NOT NULL CHECK (role IN ('user','assistant')),
content TEXT NOT NULL,
sources TEXT[],
created_at TIMESTAMPTZ DEFAULT NOW()
);
CREATE INDEX IF NOT EXISTS idx_rag_turns_session
ON rag_turns (session_id, created_at);
Python Dependencies
The Python Dependencies stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Pin the interpreter and dependency lockfile before teaching the loop. Drift between laptop and CI is the most common silent break for API demos.
uv add google-genai pypdf pgvector psycopg2-binary python-dotenv huggingface_hub
Connection Setup (Cell 2)
The Connection Setup Cell 2 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
import psycopg2
from pgvector.psycopg2 import register_vector
PG_HOST = "localhost"
PG_PORT = 5432
PG_DB = "postgres"
PG_USER = "postgres"
PG_PASSWORD = os.environ.get("PG_PASSWORD", "postgres")
def get_pg_conn():
"""Returns a fresh PostgreSQL connection with pgvector registered."""
conn = psycopg2.connect(
host=PG_HOST, port=PG_PORT,
dbname=PG_DB, user=PG_USER, password=PG_PASSWORD
)
register_vector(conn) # tells psycopg2 how to handle vector type
return conn
Storing Embeddings (Cell 6)
The Storing Embeddings Cell 6 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
import numpy as np
def store_in_postgres(chunks, embeddings, doc_name) -> int:
conn = get_pg_conn()
cur = conn.cursor()
for i, (chunk, emb) in enumerate(zip(chunks, embeddings)):
cur.execute("""
INSERT INTO rag_chunks (source, chunk_index, content, embedding)
VALUES (%s, %s, %s, %s)
""", (doc_name, i, chunk, np.array(emb))) # np.array → pgvector handles serialization
conn.commit()
conn.close()
return len(chunks)
The Storing Embeddings Cell 6 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Retrieval with Cosine Similarity (Cell 6 continued)
For the Retrieval with Cosine Similarity stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
def retrieve_context(query: str) -> list[dict]:
query_embedding = embed_query(query)
conn = get_pg_conn()
cur = conn.cursor()
cur.execute("""
SELECT
content,
source,
chunk_index,
ingested_at,
1 - (embedding <=> %s) AS cosine_score -- <=> is cosine distance
FROM rag_chunks
ORDER BY embedding <=> %s -- sort ascending (smallest distance first)
LIMIT %s
""", (np.array(query_embedding), np.array(query_embedding), TOP_K))
rows = cur.fetchall()
conn.close()
return [{"text": r[0], "source": r[1], "chunk_index": r[2],
"ingested_at": str(r[3]) if r[3] else "",
"score": round(float(r[4]), 4)} for r in rows]
CDC with pgvector — Upsert pattern
For the CDC with pgvector Upsert stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
cur.execute("""
INSERT INTO rag_staleness (doc_name, file_hash, chunk_count, version)
VALUES (%s, %s, %s, 1)
ON CONFLICT (doc_name) DO UPDATE SET
file_hash = EXCLUDED.file_hash,
chunk_count = EXCLUDED.chunk_count,
ingested_at = NOW(),
version = rag_staleness.version + 1;
""", (doc_name, file_hash, chunk_count))
Part 2 — SQL Server 2025 (Native Vector)
For the Part 2 SQL Server stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Why SQL Server 2025 Needs No Extension
For the Why SQL Server 2025 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Setup in SSMS
For the Setup in SSMS stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
-- Batch 1: Run this first — must commit before vector objects are recognized
ALTER DATABASE SCOPED CONFIGURATION SET PREVIEW_FEATURES = ON;
-- Batch 2: Create all tables
CREATE TABLE rag_chunks (
id INT IDENTITY(1,1) PRIMARY KEY CLUSTERED,
source NVARCHAR(500) NOT NULL,
chunk_index INT NOT NULL,
content NVARCHAR(MAX) NOT NULL,
file_hash NVARCHAR(64),
ingested_at DATETIME2 DEFAULT GETUTCDATE(),
embedding VECTOR(768) -- native SQL Server 2025 type
);
CREATE INDEX idx_rag_source ON rag_chunks(source);
CREATE TABLE rag_staleness (
doc_name NVARCHAR(500) PRIMARY KEY,
file_hash NVARCHAR(64) NOT NULL,
chunk_count INT,
ingested_at DATETIME2 DEFAULT GETUTCDATE(),
version INT DEFAULT 1
);
CREATE TABLE rag_chunk_registry (
doc_name NVARCHAR(500) NOT NULL,
chunk_hash NVARCHAR(64) NOT NULL,
chunk_id NVARCHAR(64) NOT NULL,
PRIMARY KEY (doc_name, chunk_hash)
);
CREATE TABLE rag_sessions (
session_id NVARCHAR(100) PRIMARY KEY,
created_at DATETIME2 DEFAULT GETUTCDATE(),
updated_at DATETIME2 DEFAULT GETUTCDATE(),
model NVARCHAR(200),
embed_model NVARCHAR(200),
turn_count INT DEFAULT 0
);
CREATE TABLE rag_turns (
id INT IDENTITY(1,1) PRIMARY KEY,
session_id NVARCHAR(100) NOT NULL REFERENCES rag_sessions(session_id),
role NVARCHAR(20) NOT NULL CHECK (role IN ('user','assistant')),
content NVARCHAR(MAX) NOT NULL,
sources NVARCHAR(MAX),
created_at DATETIME2 DEFAULT GETUTCDATE()
);
CREATE INDEX idx_rag_turns_session ON rag_turns(session_id, created_at);
-- Batch 3: Must run AFTER Batch 2 commits
-- Cannot run inside a transaction — this is a known SQL Server 2025 preview constraint
CREATE VECTOR INDEX idx_rag_embedding
ON rag_chunks(embedding)
WITH (METRIC = 'COSINE');
Python Dependencies
For the Python Dependencies stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
uv add google-genai pypdf pyodbc python-dotenv huggingface_hub fpdf2
Connection Setup (Cell 2)
For the Connection Setup Cell 2 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
import pyodbc
SQL_SERVER = "YOUR_SERVER_NAME" # from SSMS title bar
SQL_DATABASE = "local_rag"
SQL_CONN_STR = (
f"DRIVER={{ODBC Driver 17 for SQL Server}};"
f"SERVER={SQL_SERVER};"
f"DATABASE={SQL_DATABASE};"
f"Trusted_Connection=yes;" # Windows Authentication - no password needed
)
def get_conn():
return pyodbc.connect(SQL_CONN_STR)
def get_conn_autocommit():
"""Required for CREATE/DROP VECTOR INDEX - cannot run inside a transaction."""
return pyodbc.connect(SQL_CONN_STR, autocommit=True)
Storing Embeddings (Cell 6)
For the Storing Embeddings Cell 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Storing Embeddings Cell 6 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
import json
def vec_to_json(embedding: list[float]) -> str:
return json.dumps(embedding) # '[0.12, 0.34, ...]'
def store_in_sqlserver(chunks, embeddings, doc_name) -> int:
drop_vector_index() # must drop before any INSERT
conn = get_conn()
cur = conn.cursor()
sql = """
INSERT INTO rag_chunks (source, chunk_index, content, embedding)
VALUES (?, ?, ?, CAST(? AS VECTOR(768)))
"""
for i, (chunk, emb) in enumerate(zip(chunks, embeddings)):
cur.setinputsizes([
(_pyodbc.SQL_WVARCHAR, 500, 0),
_pyodbc.SQL_INTEGER,
(_pyodbc.SQL_WVARCHAR, 0, 0),
(_pyodbc.SQL_VARCHAR, 0, 0), # ← must be VARCHAR, not NTEXT
])
cur.execute(sql, (doc_name, i, chunk, vec_to_json(emb)))
conn.commit()
conn.close()
create_vector_index() # recreate after all inserts
return len(chunks)
Retrieval with VECTOR_DISTANCE (Cell 6 continued)
When working through the Retrieval with VECTORDISTANCE Cell stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
def retrieve_context(query: str) -> list[dict]:
query_embedding = embed_query(query)
query_vec_json = vec_to_json(query_embedding)
conn = get_conn()
cur = conn.cursor()
cur.setinputsizes([(_pyodbc.SQL_VARCHAR, 0, 0)]) # force VARCHAR for vector param
cur.execute(f"""
SELECT TOP ({TOP_K})
content, source, chunk_index, ingested_at,
VECTOR_DISTANCE('cosine', embedding, CAST(? AS VECTOR(768))) AS distance
FROM rag_chunks
ORDER BY distance ASC;
""", (query_vec_json,))
rows = cur.fetchall()
conn.close()
return [{"text": r[0], "source": r[1], "chunk_index": r[2],
"ingested_at": str(r[3]) if r[3] else "",
"score": round(1 - float(r[4]), 4)} for r in rows]
SQL Server 2025 Vector Gotchas — Every One We Hit
When working through the SQL Server 2025 Vector stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Gotcha 1 — Unknown object type ‘VECTOR’ in CREATE statement
When working through the Gotcha 1 Unknown object stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Gotcha 1 Unknown object stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Msg 343, Level 15: Unknown object type 'VECTOR' used in CREATE, DROP, or ALTER statement.
Gotcha 2 — Primary key must be a single 4-byte INT column
The Gotcha 2 Primary key stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Msg 42217: Table must have a clustered primary key on a single 4 byte INT column to create a vector index.
-- ❌ Does not work with vector index
id UNIQUEIDENTIFIER PRIMARY KEY DEFAULT NEWID()
-- ✅ Required
id INT IDENTITY(1,1) PRIMARY KEY CLUSTERED
Gotcha 3 — Cannot INSERT/DELETE/UPDATE while vector index exists
The Gotcha 3 Cannot INSERT stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Msg 42231: Data modification statement failed because table 'rag_chunks' has a vector index on it.
def drop_vector_index():
conn = get_conn_autocommit() # autocommit required
conn.cursor().execute("""
IF EXISTS (
SELECT 1 FROM sys.indexes
WHERE name = 'idx_rag_embedding'
AND object_id = OBJECT_ID('rag_chunks')
)
DROP INDEX idx_rag_embedding ON rag_chunks;
""")
conn.close()
def create_vector_index():
conn = get_conn_autocommit() # autocommit required
conn.cursor().execute("""
CREATE VECTOR INDEX idx_rag_embedding
ON rag_chunks(embedding)
WITH (METRIC = 'COSINE');
""")
conn.close()
Gotcha 4 — CREATE VECTOR INDEX cannot run inside a transaction
The Gotcha 4 CREATE VECTOR stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Gotcha 4 CREATE VECTOR stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Msg 574: CREATE VECTOR INDEX statement cannot be used inside a user transaction.
# ❌ Fails — implicit transaction
conn = pyodbc.connect(SQL_CONN_STR)
conn.cursor().execute("CREATE VECTOR INDEX ...")
# ✅ Works - no transaction wrapper
conn = pyodbc.connect(SQL_CONN_STR, autocommit=True)
conn.cursor().execute("CREATE VECTOR INDEX ...")
Gotcha 5 — Explicit conversion from ntext to vector is not allowed
For the Gotcha 5 Explicit conversion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Msg 529: Explicit conversion from data type ntext to vector is not allowed.
cur.setinputsizes([
(_pyodbc.SQL_WVARCHAR, 500, 0), # source — Unicode fine
_pyodbc.SQL_INTEGER, # chunk_index
(_pyodbc.SQL_WVARCHAR, 0, 0), # content — Unicode fine
(_pyodbc.SQL_VARCHAR, 0, 0), # embedding ← must be ASCII VARCHAR
])
cur.execute(sql, (doc_name, i, chunk, vec_to_json(emb)))
Gotcha 6 — VECTOR_SEARCH does not accept ? parameter placeholders
For the Gotcha 6 VECTORSEARCH does stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Msg 102: Incorrect syntax near '('.
-- ❌ VECTOR_SEARCH with parameter placeholder — fails
FROM VECTOR_SEARCH(
TABLE = rag_chunks USING VECTOR INDEX idx_rag_embedding,
SIMILAR_TO = CAST(? AS VECTOR(768)), -- pyodbc cannot pass ? here
...
)
-- ✅ VECTOR_DISTANCE - fully parameterized, GA, works perfectly
SELECT TOP (5)
content,
VECTOR_DISTANCE('cosine', embedding, CAST(? AS VECTOR(768))) AS distance
FROM rag_chunks
ORDER BY distance ASC;
PostgreSQL vs SQL Server — Side by Side
For the PostgreSQL vs SQL Server stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
What Both Databases Add That ChromaDB Cannot
For the What Both Databases Add stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
-- PostgreSQL: Find chunks from documents belonging to a specific customer
SELECT c.content, 1 - (c.embedding <=> query_vec) AS score
FROM rag_chunks c
JOIN documents d ON d.filename = c.source
JOIN customers cu ON cu.id = d.customer_id
WHERE cu.tier = 'enterprise'
ORDER BY score DESC
LIMIT 5;
-- SQL Server: Same query, T-SQL syntax
SELECT TOP 5
c.content,
1 - VECTOR_DISTANCE('cosine', c.embedding, CAST(? AS VECTOR(768))) AS score
FROM rag_chunks c
JOIN documents d ON d.filename = c.source
JOIN customers cu ON cu.id = d.customer_id
WHERE cu.tier = 'enterprise'
ORDER BY score DESC;
Conclusion
For the Conclusion stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.