This article is published in English.
VPC-Only RAG on One EC2: Ollama, Qdrant, and LangChain
Replace tribal knowledge with a cited runbook assistant that never lets documents leave your AWS account—three containers on a GPU EC2.
The unofficial knowledge oracle problem
Every engineering org has people who hold the answers: how deploys work, what leave policy says, why Kafka won over SQS. That works until those people are on PTO, change teams, or a new hire needs truth at 11 PM during an incident.
A better shape: anyone types “Can I deploy on Friday?” and gets a documented answer in seconds — grounded in real runbooks, with a citation. Not a guess. Not a fluent hallucination.
One constraint drove the design: no internal document leaves the AWS account. No OpenAI keys. No hosted vector DB. Everything had to run in-VPC on a single EC2 box.
What you get
A simple Streamlit chat. Ask a question; expand sources to see which file and section grounded the reply. Example shape: dependencies of the payment service answered from deployment-runbook.md, not inventing hostnames.
Architecture — three containers
One docker compose up brings:
version: "3.9"
services:
ollama:
image: ollama/ollama:latest
container_name: ollama
ports:
- "11434:11434"
volumes:
- ollama_data:/root/.ollama
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
restart: unless-stopped
qdrant:
image: qdrant/qdrant:v1.18.2
container_name: qdrant
ports:
- "6333:6333"
- "6334:6334"
volumes:
- qdrant_data:/qdrant/storage
environment:
- QDRANT__SERVICE__ENABLE_CORS=true
restart: unless-stopped
app:
build:
context: .
dockerfile: Dockerfile
container_name: rag-app
ports:
- "8501:8501"
environment:
- OLLAMA_BASE_URL=http://ollama:11434
- QDRANT_URL=http://qdrant:6333
- COLLECTION=company_docs
- EMBED_MODEL=nomic-embed-text
- LLM_MODEL=llama3.1:8b
- S3_CORPUS_BUCKET=${S3_CORPUS_BUCKET}
- S3_CORPUS_PREFIX=${S3_CORPUS_PREFIX:-corpus/}
- AWS_REGION=${AWS_REGION:-us-east-1}
- AWS_ACCESS_KEY_ID=${AWS_ACCESS_KEY_ID:-}
- AWS_SECRET_ACCESS_KEY=${AWS_SECRET_ACCESS_KEY:-}
volumes:
- corpus_data:/app/corpus
- ./scripts:/app/scripts
depends_on:
- ollama
- qdrant
restart: unless-stopped
volumes:
ollama_data:
qdrant_data:
corpus_data:
- Ollama — local embedding model plus a chat model (for example llama3.1). No API keys, no per-token bill.
- Qdrant — stores chunk vectors and retrieves nearest neighbors fast.
- LangChain — retrieval → context packing → generation, with citations.
All of it sits on a GPU EC2 instance inside the VPC so corpora never cross the boundary.
Wiring the ask path
The LCEL-style chain retrieves, formats, and prompts the local chat model:
from operator import itemgetter
from langchain_core.runnables import RunnableLambda, RunnableParallel
chain = RunnableParallel(
question=itemgetter("question"), # pass question through
docs=itemgetter("question") | retriever, # embed + search Qdrant
).assign(
answer=(
{"context": itemgetter("docs") | RunnableLambda(format_docs),
"question": itemgetter("question")}
| prompt_template # inject into system/human message
| llm # send to Ollama
| str_parser # extract text from response
)
)
Retrieval stays lexical+dense as you configure; generation stays on Ollama. Citations come from metadata on the chunks Qdrant returns.
Ingestion and ops notes
Drop markdown/runbooks into the ingest path, chunk them, embed with the local embedding model, upsert into Qdrant. Re-run when docs change. Watch disk for models, GPU memory for concurrent asks, and security groups so only the VPN/bastion reaches Streamlit and Qdrant.
Why this pattern wins for internal knowledge
You trade peak model quality for data gravity and cost control. For “what does our runbook say?” that trade is correct. When Aaron is out, the runbook still answers — and the citation shows where to verify.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Keep the deployment boring: compose up, health checks, and a single place to rotate models. Boring is how internal tools survive contact with real on-call schedules.
Constraint: data never leaves the VPC
The non-negotiable rule shaped every vendor choice. Managed OpenAI-compatible APIs were out. Hosted vector databases were out. Even “temporary” staging uploads were out. That left a single EC2 GPU instance, a compose file, and IAM/security-group discipline.
If your legal team cannot draw a line around the EC2 ENI, this architecture is not for you. If they can, you get citations without shipping PDFs to a third party.
Ingest pipeline details
Runbooks and ADRs land as markdown in a watched directory. A job splits them into overlapping chunks, embeds with the local embedding model via Ollama, and upserts points into Qdrant with path and section metadata. Failures should abort the upsert for that file rather than silently dropping half a policy document.
Re-embed when a file’s hash changes. Avoid re-embedding the entire corpus on every tiny edit if you can diff by path.
Query-time behavior
The Streamlit app sends the user question to a retriever configured with a small k. Retrieved texts are formatted into a prompt that tells the chat model to answer only from context and to refuse when context is thin. Citations render from metadata so a human can open the exact runbook section.
That refuse-when-thin behavior matters more than temperature knobs. Fluent nonsense with a fake citation is worse than “I don’t see that in the runbooks.”
Failure modes to plan for
GPU OOM when someone pulls a larger chat model without raising instance size. Qdrant disk growth if you ingest binary junk. Stale answers after a process change that nobody re-ingested. Compose stacks that bind ports to 0.0.0.0 on a public subnet — keep them private.
On-call should know how to bounce Ollama, rebuild the Qdrant volume from source markdown, and verify a golden question still cites the right file.
What “done” looks like
A new hire on day two can ask deploy questions without paging Aaron. Answers include paths. The VPC boundary holds under review. Cost is one EC2 plus storage — predictable. That is enough success for an internal knowledge assistant; chase frontier model quality only after citations are trustworthy.