Home / Articles / Practical notes: Multi-Tenant RAG on Amazon S3 Vectors - Part 3: The Pipeline

This article is published in English.

Practical notes: Multi-Tenant RAG on Amazon S3 Vectors - Part 3: The Pipeline

Operable walkthrough of Practical notes: Multi-Tenant RAG on Amazon S3 Vectors - Part 3: The Pipeline: contracts, checks, and drop-in code slots for teams shipping this pattern.

3107 words

This walkthrough rebuilds the path from raw materials to a working system for: Multi-Tenant RAG on Amazon S3 Vectors - Part 3: The Pipeline. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Architecture

When working through the Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Ingest:  S3 (tenant prefix) ──▶ Lambda: extract + chunk ──▶ SQS ──▶ Lambda: embed (Bedrock Titan V2)
                                                                     └──▶ PutVectors → index/org-<tenant>
Query:   eID provider (OIDC) ──▶ API authorizer (tenant, role) ──▶ STS session policy (one index ARN)
                        ──▶ embed question ──▶ QueryVectors(index/org-<tenant>, filter=classification)
                        ──▶ LLM, grounded on returned chunk_text, citations = document_id + version
Audit:   CloudTrail data events, resource type AWS::S3Vectors::Index, queried by resources.ARN

Index Creation

When working through the Index Creation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

import { S3VectorsClient, CreateIndexCommand } from "@aws-sdk/client-s3vectors";
const s3v = new S3VectorsClient({ region: "eu-west-1" });
await s3v.send(new CreateIndexCommand({
  vectorBucketName: "kb-eu-west-1",      // the bucket is regional → residency
  indexName: "org-b",                       // the index is the tenant → isolation
  dataType: "float32",
  dimension: 1024,                          // Titan Text Embeddings V2, 1024-d
  distanceMetric: "cosine",
  metadataConfiguration: {
    nonFilterableMetadataKeys: ["chunk_text"],   // returned, never scanned, never filterable
  },
}));

Ingestion: Choose the Keys

When working through the Ingestion Choose the Keys stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Ingestion Choose the Keys stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

import { PutVectorsCommand } from "@aws-sdk/client-s3vectors";
async function ingestDocument(tenant: string, doc: Document, chunks: string[]) {
  const vectors = await Promise.all(chunks.map(async (text, i) => ({
    key: `${doc.id}-k${String(i + 1).padStart(3, "0")}`,      // chosen by us
    data: { float32: await embed(text) },
    metadata: {
      document_id: doc.id,                 // filterable
      version: doc.version,                // filterable → appears on the citation
      classification: doc.classification, // filterable → role filter at query time
      chunk_text: text,                    // non-filterable → rides with the vector
    },
  })));  for (const batch of chunk(vectors, 500)) {                    // ≤ 500 per call
    await s3v.send(new PutVectorsCommand({
      vectorBucketName: "kb-eu-west-1", indexName: `org-${tenant}`, vectors: batch,
    }));
  }
  await manifest.put(tenant, doc.id, doc.version, vectors.map(v => v.key));   // written down
}
import { BedrockRuntimeClient, InvokeModelCommand } from "@aws-sdk/client-bedrock-runtime";
const bedrock = new BedrockRuntimeClient({ region: "eu-west-1" });
async function embed(text: string): Promise<number[]> {
  const res = await bedrock.send(new InvokeModelCommand({
    modelId: "amazon.titan-embed-text-v2:0",
    contentType: "application/json",
    body: JSON.stringify({ inputText: text, dimensions: 1024, normalize: true }),
  }));
  return JSON.parse(new TextDecoder().decode(res.body)).embedding;
}

Query: Credentials that Name one Index

The Query Credentials that Name stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

import { STSClient, AssumeRoleCommand } from "@aws-sdk/client-sts";
function indexArn(tenant: string) {
  return `arn:aws:s3vectors:eu-west-1:${ACCOUNT}:bucket/kb-eu-west-1/index/org-${tenant}`;
}async function tenantCredentials(tenant: string) {
  const res = await sts.send(new AssumeRoleCommand({
    RoleArn: QUERY_ROLE_ARN,                       // broad role
    RoleSessionName: `q-${tenant}`,
    DurationSeconds: 900,
    Policy: JSON.stringify({                       // session policy: intersection, never a widening
      Version: "2012-10-17",
      Statement: [{
        Effect: "Allow",
        Action: ["s3vectors:QueryVectors", "s3vectors:GetVectors"],
        Resource: indexArn(tenant),                // exactly one ARN
      }],
    }),
  }));
  return res.Credentials!;
}
import { QueryVectorsCommand } from "@aws-sdk/client-s3vectors";
async function retrieve(tenant: string, role: Role, question: string) {
  const client = new S3VectorsClient({ region: "eu-west-1", credentials: await tenantCredentials(tenant) });
  const res = await client.send(new QueryVectorsCommand({
    indexArn: indexArn(tenant),
    queryVector: { float32: await embed(question) },
    topK: 10,
    filter: { classification: { $in: allowedClassifications(role) } },   // evaluated during search
    returnMetadata: true,
    returnDistance: true,
  }));
  return res.vectors ?? [];    // each: key, distance, metadata incl. chunk_text
}

The Deliberate Bug

The The Deliberate Bug stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

const client = new S3VectorsClient({ region: "eu-west-1", credentials: await tenantCredentials("b") });
await client.send(new QueryVectorsCommand({ indexArn: indexArn("a"), queryVector: { float32: q }, topK: 10 }));
// → AccessDeniedException: User ... is not authorized to perform: s3vectors:QueryVectors on resource: .../index/org-a

Erasure: Delete by Key, Verify by Key

The Erasure Delete by Key stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Erasure Delete by Key stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

import { DeleteVectorsCommand, GetVectorsCommand } from "@aws-sdk/client-s3vectors";
async function eraseDocument(tenant: string, docId: string, version: number) {
  const keys = await manifest.get(tenant, docId, version);
  for (const batch of chunk(keys, 500)) {
    await s3v.send(new DeleteVectorsCommand({ vectorBucketName: "kb-eu-west-1", indexName: `org-${tenant}`, keys: batch }));
  }
  // verify — strongly consistent, so this is valid immediately
  for (const batch of chunk(keys, 100)) {
    const res = await s3v.send(new GetVectorsCommand({ vectorBucketName: "kb-eu-west-1", indexName: `org-${tenant}`, keys: batch }));
    if ((res.vectors ?? []).length) throw new Error(`erasure incomplete: ${res.vectors!.length} keys remain`);
  }
  await audit.record({ tenant, docId, version, keyCount: keys.length, verifiedAt: new Date() });
}

Audit: CloudTrail Data Events by ARN

For the Audit CloudTrail Data Events stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

// CDK
new cloudtrail.Trail(this, "Trail", { sendToCloudWatchLogs: false })
  .addEventSelector(cloudtrail.DataResourceType.S3_VECTORS_INDEX, [
    `arn:aws:s3vectors:eu-west-1:${account}:bucket/kb-eu-west-1/index/*`,
  ]);
// Check the CDK enum/name for the S3 Vectors index resource type; the underlying resource type is AWS::S3Vectors::Index.
{
  "eventSource": "s3vectors.amazonaws.com",
  "eventName": "QueryVectors",
  "eventTime": "2026-09-09T10:41:07Z",
  "userIdentity": { "type": "AssumedRole", "arn": "arn:aws:sts::123456789012:assumed-role/kb-query/q-b" },
  "resources": [{ "type": "AWS::S3Vectors::Index",
                  "ARN": "arn:aws:s3vectors:eu-west-1:123456789012:bucket/kb-eu-west-1/index/org-b" }],
  "errorCode": null
}
SELECT eventTime, eventName, userIdentity.arn, errorCode
FROM   cloudtrail_events
WHERE  eventSource = 's3vectors.amazonaws.com'
  AND  element_at(resources, 1).arn LIKE '%/index/org-a'
ORDER  BY eventTime DESC;

Four Additions that Fit the Model

For the Four Additions that Fit stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

A KMS key per index.

For the A KMS key per stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the A KMS key per stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Export.

When working through the Export stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Private network path.

When working through the Private network path stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Cost per tenant.

When working through the Cost per tenant stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Cost per tenant stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

What the Index Boundary Does Not Solve

The What the Index Boundary stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Roles inside a tenant.

The Roles inside a tenant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Existence leakage.

The Existence leakage stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Existence leakage stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Cross-language similarity.

For the Cross-language similarity stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

No lexical search.

For the No lexical search stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

Chunk boundaries.

For the Chunk boundaries stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the Chunk boundaries stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Model migration.

When working through the Model migration stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Bedrock Knowledge Bases.

When working through the Bedrock Knowledge Bases stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Requirements, revisited

When working through the Requirements revisited stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Requirements revisited stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Operational checklist

The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for 0b50bb1ebd01: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.