This article is published in English.
Ten ways to cut hallucinations without fine-tuning
Grounding, constraints, retrieval, and eval loops that reduce invented facts in production answers.
This walkthrough rebuilds the path from raw materials to a working system for: 10 Ways to Reduce AI Hallucinations Without Fine-Tuning Your Model. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For Overview, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
The Basic Problem
When working through The Basic Problem, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
const response = await client.responses.create({
model: "gpt-5",
input: `
Where is order #48291?
`
});
console.log(response.output_text);
Customer
↓
AI Agent
↓
Order System
↓
Actual Order State
↓
AI Agent
↓
Response
1. Give the Model Better Context
When working through 1. Give the Model Better Context, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
{
"orderId": "48291",
"status": "SHIPPED",
"carrier": "FedEx",
"trackingNumber": "784512963",
"estimatedDelivery": "2026-08-25"
}
const context = {
order: {
id: "48291",
status: "SHIPPED",
carrier: "FedEx",
trackingNumber: "784512963",
estimatedDelivery: "2026-08-25"
}
};
const response = await client.responses.create({
model: "gpt-5",
input: `
Answer the customer using only the supplied order information.
Context:
${JSON.stringify(context, null, 2)}
Customer:
Where is my order #48291?
`
});
The pattern
When working through The pattern, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Bad:
Customer → LLM → Answer
Better:
Customer
↓
Retrieve state
↓
Build context
↓
LLM
↓
Answer
When working through The pattern, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
2. Retrieve Before You Generate
- Retrieve Before You Generate works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Customer Question
↓
Retrieve
↓
Rank
↓
Build Context
↓
LLM
↓
Answer
const results = await vectorStore.search({
query: customerQuestion,
topK: 10
});
const relevant = results
.filter(item => item.score > 0.8)
.slice(0, 5);
const context = relevant
.map(item => item.content)
.join("\n\n");
const answer = await generateAnswer(
customerQuestion,
context
);
3. Reduce Context Noise
- Reduce Context Noise works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
20 retrieved documents
+
15 previous messages
+
10 previous tool responses
+
customer profile
+
product catalog
+
order history
+
promotion metadata
Available Information
↓
Relevance Filtering
↓
Metadata Filtering
↓
Ranking
↓
Deduplication
↓
Context Compression
↓
LLM
const context = results
.filter(x => x.score >= 0.82)
.filter(x => x.metadata.category === "returns")
.filter(x => x.metadata.region === customer.region)
.sort((a, b) => b.score - a.score)
.slice(0, 5)
.map(x => x.content);
4. Use a Knowledge Graph for Structured Facts
- Use a Knowledge Graph for Structured Facts works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
- Use a Knowledge Graph for Structured Facts works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Customer
│
└── PLACED → Order
│
├── CONTAINS → Product
│
├── PAID_BY → Payment
│
├── FULFILLED_BY → Warehouse
│
└── SHIPPED_BY → Carrier
Order #48291
↓
FULFILLED_BY
↓
Warehouse #17
MATCH (o:Order {id: "48291"})
-[:FULFILLED_BY]->
(w:Warehouse)
RETURN w.id, w.name, w.location;
{
"orderId": "48291",
"warehouse": {
"id": "WH-17",
"name": "Delhi Fulfillment Center",
"location": "Delhi"
}
}
Without structured knowledge:
User → LLM
↓
Guess
With Knowledge Graph:
User
↓
Entity Identification
↓
Graph Traversal
↓
Verified Relationship
↓
Context
↓
LLM
5. Ground Answers With Evidence
For 5. Ground Answers With Evidence, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
{
"answer": "Your order is being fulfilled by Warehouse WH-17.",
"confidence": 0.98,
"evidence": [
{
"type": "order_record",
"source": "orders_db",
"orderId": "48291"
},
{
"type": "warehouse_relationship",
"source": "knowledge_graph",
"warehouseId": "WH-17"
}
]
}
For every factual claim:
1. Identify supporting evidence.
2. Use only available evidence.
3. Never invent a source.
4. If evidence is unavailable, say so.
5. Clearly distinguish facts from inference.
if (result.confidence < 0.7) {
return escalateToHuman(result);
}
LLM → Answer
LLM
↓
Answer
↓
Evidence
↓
Confidence
↓
Decision
6. Give the Agent Tools Instead of Making It Guess
For 6. Give the Agent Tools Instead of Making It Guess, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
const tools = [{
name: "get_refund_status",
description: "Retrieve the current refund status for an order",
parameters: {
type: "object",
properties: {
orderId: {
type: "string"
}
},
required: ["orderId"]
}
}];
Customer
↓
LLM
↓
get_refund_status()
↓
Payment System
↓
Actual Refund State
↓
LLM
↓
Customer
7. Validate Both Tool Inputs and Outputs
For 7. Validate Both Tool Inputs and Outputs, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For 7. Validate Both Tool Inputs and Outputs, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
{
"orderId": "48291",
"amount": "one hundred",
"currency": "dollars"
}
{
"orderId": "48291",
"amount": 100,
"currency": "USD"
}
import { z } from "zod";
const RefundRequest = z.object({
orderId: z.string(),
amount: z.number().positive(),
currency: z.enum(["USD", "EUR", "GBP", "INR"])
});
const refundRequest = RefundRequest.parse(
modelOutput
);
if (refundRequest.amount > order.total) {
throw new Error(
"Refund amount exceeds order total"
);
}
LLM
↓
Schema Validation
↓
Business Validation
↓
Permission Check
↓
Execute
↓
Response Validation
↓
Accept / Retry / Escalate
8. Separate Facts From Reasoning
When working through 8. Separate Facts From Reasoning, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
{
"facts": [
"Order 48291 is shipped",
"Order 48291 is fulfilled by Warehouse WH-17",
"Warehouse WH-17 is currently operating"
],
"reasoning": [
"The order is likely to remain on schedule"
],
"conclusion": "The order is currently expected to arrive on time."
}
Was the fact wrong?
OR
Was the reasoning wrong?
FACT
→ Order shipped
FACT
→ Estimated delivery: Aug 25
INFERENCE
→ Delivery is currently expected on schedule
9. Learn From Successful and Failed Executions
When working through 9. Learn From Successful and Failed Executions, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
{
"orderId": "48291",
"status": "PROCESSING",
"cancelable": true
}
cancel_order(48291)
{
"success": true,
"cancellationId": "CAN-83921"
}
{
"task": "Cancel order",
"orderState": "PROCESSING",
"action": "cancel_order",
"result": "SUCCESS",
"cancellationId": "CAN-83921"
}
New Request
↓
Find Similar Successful Execution
↓
Retrieve Relevant Pattern
↓
Check Current Order State
↓
Generate Action
↓
Validate
↓
Execute
Attempt 1
↓
cancel_order()
↓
Rejected: Order already shipped
↓
Agent retrieves shipping information
↓
Explains cancellation is unavailable
Execute
↓
Observe
↓
Evaluate
↓
Store Experience
↓
Improve Future Context
10. Evaluate Every Change
When working through 10. Evaluate Every Change, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through 10. Evaluate Every Change, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
[
{
"question": "Where is order 48291?",
"expected": "SHIPPED"
},
{
"question": "Which warehouse fulfills order 48291?",
"expected": "WH-17"
},
{
"question": "Can order 48291 be cancelled?",
"expected": false
}
]
Answer Accuracy
Groundedness
Retrieval Precision
Tool Selection Accuracy
Tool Success Rate
Task Success Rate
Recovery Rate
Hallucination Rate
Latency
Cost
Before After
Answer Accuracy 72% 95%
Groundedness 69% 97%
Tool Success 81% 98%
Hallucination Rate 17% 3%
Change
↓
Test
↓
Measure
↓
Compare
↓
Improve
Putting It All Together
Putting It All Together works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
┌──────────────────┐
│ Knowledge Graph │
└────────┬─────────┘
│
┌────────▼─────────┐
│ RAG / Search │
└────────┬─────────┘
│
Customer → Intent → Context Engine → LLM
↑ │
│ ▼
Memory Tool Selection
↑ │
│ ▼
│ Validation
│ │
│ ▼
│ Real Systems
│ │
│ ▼
└─────── Feedback
│
▼
Evaluation
The Bigger Lesson
The Bigger Lesson works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
The LLM Is Only One Part of the System
The LLM Is Only One Part of the System works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The LLM Is Only One Part of the System works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
┌───────────────────┐
│ Knowledge + RAG │
└─────────┬─────────┘
↓
Customer → Context → LLM → Tools → Real World
↑ ↓ ↓
Memory Reasoning Validation
↑ ↓
└──── Feedback
↓
Evaluation
Final Thought
For Final Thought, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Operational checklist
Operational checklist works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for d3bfab080c19: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.