Home / Articles / Ten ways to cut hallucinations without fine-tuning

This article is published in English.

Ten ways to cut hallucinations without fine-tuning

Grounding, constraints, retrieval, and eval loops that reduce invented facts in production answers.

2855 words

This walkthrough rebuilds the path from raw materials to a working system for: 10 Ways to Reduce AI Hallucinations Without Fine-Tuning Your Model. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For Overview, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

The Basic Problem

When working through The Basic Problem, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

const response = await client.responses.create({
  model: "gpt-5",
  input: `
    Where is order #48291?
  `
});

console.log(response.output_text);
Customer
   ↓
AI Agent
   ↓
Order System
   ↓
Actual Order State
   ↓
AI Agent
   ↓
Response

1. Give the Model Better Context

When working through 1. Give the Model Better Context, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

{
  "orderId": "48291",
  "status": "SHIPPED",
  "carrier": "FedEx",
  "trackingNumber": "784512963",
  "estimatedDelivery": "2026-08-25"
}
const context = {
  order: {
    id: "48291",
    status: "SHIPPED",
    carrier: "FedEx",
    trackingNumber: "784512963",
    estimatedDelivery: "2026-08-25"
  }
};

const response = await client.responses.create({
  model: "gpt-5",
  input: `
    Answer the customer using only the supplied order information.
    Context:
    ${JSON.stringify(context, null, 2)}
    Customer:
    Where is my order #48291?
  `
});

The pattern

When working through The pattern, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

Bad:

Customer → LLM → Answer

Better:
Customer
   ↓
Retrieve state
   ↓
Build context
   ↓
LLM
   ↓
Answer

When working through The pattern, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

2. Retrieve Before You Generate

  1. Retrieve Before You Generate works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Customer Question
       ↓
    Retrieve
       ↓
     Rank
       ↓
Build Context
       ↓
      LLM
       ↓
    Answer
const results = await vectorStore.search({
  query: customerQuestion,
  topK: 10
});

const relevant = results
  .filter(item => item.score > 0.8)
  .slice(0, 5);
const context = relevant
  .map(item => item.content)
  .join("\n\n");
const answer = await generateAnswer(
  customerQuestion,
  context
);

3. Reduce Context Noise

  1. Reduce Context Noise works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
20 retrieved documents
+
15 previous messages
+
10 previous tool responses
+
customer profile
+
product catalog
+
order history
+
promotion metadata
Available Information
        ↓
Relevance Filtering
        ↓
Metadata Filtering
        ↓
Ranking
        ↓
Deduplication
        ↓
Context Compression
        ↓
LLM
const context = results
  .filter(x => x.score >= 0.82)
  .filter(x => x.metadata.category === "returns")
  .filter(x => x.metadata.region === customer.region)
  .sort((a, b) => b.score - a.score)
  .slice(0, 5)
  .map(x => x.content);

4. Use a Knowledge Graph for Structured Facts

  1. Use a Knowledge Graph for Structured Facts works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
  2. Use a Knowledge Graph for Structured Facts works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Customer
   │
   └── PLACED → Order
                  │
                  ├── CONTAINS → Product
                  │
                  ├── PAID_BY → Payment
                  │
                  ├── FULFILLED_BY → Warehouse
                  │
                  └── SHIPPED_BY → Carrier
Order #48291
      ↓
FULFILLED_BY
      ↓
Warehouse #17
MATCH (o:Order {id: "48291"})
      -[:FULFILLED_BY]->
      (w:Warehouse)
RETURN w.id, w.name, w.location;
{
  "orderId": "48291",
  "warehouse": {
    "id": "WH-17",
    "name": "Delhi Fulfillment Center",
    "location": "Delhi"
  }
}
Without structured knowledge:

User → LLM
         ↓
      Guess
With Knowledge Graph:

User
 ↓
Entity Identification
 ↓
Graph Traversal
 ↓
Verified Relationship
 ↓
Context
 ↓
LLM

5. Ground Answers With Evidence

For 5. Ground Answers With Evidence, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

{
  "answer": "Your order is being fulfilled by Warehouse WH-17.",
  "confidence": 0.98,
  "evidence": [
    {
      "type": "order_record",
      "source": "orders_db",
      "orderId": "48291"
    },
    {
      "type": "warehouse_relationship",
      "source": "knowledge_graph",
      "warehouseId": "WH-17"
    }
  ]
}
For every factual claim:
1. Identify supporting evidence.
2. Use only available evidence.
3. Never invent a source.
4. If evidence is unavailable, say so.
5. Clearly distinguish facts from inference.
if (result.confidence < 0.7) {
  return escalateToHuman(result);
}
LLM → Answer
LLM
 ↓
Answer
 ↓
Evidence
 ↓
Confidence
 ↓
Decision

6. Give the Agent Tools Instead of Making It Guess

For 6. Give the Agent Tools Instead of Making It Guess, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

const tools = [{
  name: "get_refund_status",
  description: "Retrieve the current refund status for an order",
  parameters: {
    type: "object",
    properties: {
      orderId: {
        type: "string"
      }
    },
    required: ["orderId"]
  }
}];
Customer
   ↓
LLM
   ↓
get_refund_status()
   ↓
Payment System
   ↓
Actual Refund State
   ↓
LLM
   ↓
Customer

7. Validate Both Tool Inputs and Outputs

For 7. Validate Both Tool Inputs and Outputs, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call. For 7. Validate Both Tool Inputs and Outputs, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

{
  "orderId": "48291",
  "amount": "one hundred",
  "currency": "dollars"
}
{
  "orderId": "48291",
  "amount": 100,
  "currency": "USD"
}
import { z } from "zod";

const RefundRequest = z.object({
  orderId: z.string(),
  amount: z.number().positive(),
  currency: z.enum(["USD", "EUR", "GBP", "INR"])
});
const refundRequest = RefundRequest.parse(
  modelOutput
);
if (refundRequest.amount > order.total) {
  throw new Error(
    "Refund amount exceeds order total"
  );
}
LLM
 ↓
Schema Validation
 ↓
Business Validation
 ↓
Permission Check
 ↓
Execute
 ↓
Response Validation
 ↓
Accept / Retry / Escalate

8. Separate Facts From Reasoning

When working through 8. Separate Facts From Reasoning, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

{
  "facts": [
    "Order 48291 is shipped",
    "Order 48291 is fulfilled by Warehouse WH-17",
    "Warehouse WH-17 is currently operating"
  ],
"reasoning": [
    "The order is likely to remain on schedule"
  ],
  "conclusion": "The order is currently expected to arrive on time."
}
Was the fact wrong?

OR
Was the reasoning wrong?
FACT
→ Order shipped

FACT
→ Estimated delivery: Aug 25
INFERENCE
→ Delivery is currently expected on schedule

9. Learn From Successful and Failed Executions

When working through 9. Learn From Successful and Failed Executions, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

{
  "orderId": "48291",
  "status": "PROCESSING",
  "cancelable": true
}
cancel_order(48291)
{
  "success": true,
  "cancellationId": "CAN-83921"
}
{
  "task": "Cancel order",
  "orderState": "PROCESSING",
  "action": "cancel_order",
  "result": "SUCCESS",
  "cancellationId": "CAN-83921"
}
New Request
    ↓
Find Similar Successful Execution
    ↓
Retrieve Relevant Pattern
    ↓
Check Current Order State
    ↓
Generate Action
    ↓
Validate
    ↓
Execute
Attempt 1
    ↓
cancel_order()
    ↓
Rejected: Order already shipped
    ↓
Agent retrieves shipping information
    ↓
Explains cancellation is unavailable
Execute
   ↓
Observe
   ↓
Evaluate
   ↓
Store Experience
   ↓
Improve Future Context

10. Evaluate Every Change

When working through 10. Evaluate Every Change, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through 10. Evaluate Every Change, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

[
  {
    "question": "Where is order 48291?",
    "expected": "SHIPPED"
  },
  {
    "question": "Which warehouse fulfills order 48291?",
    "expected": "WH-17"
  },
  {
    "question": "Can order 48291 be cancelled?",
    "expected": false
  }
]
Answer Accuracy
Groundedness
Retrieval Precision
Tool Selection Accuracy
Tool Success Rate
Task Success Rate
Recovery Rate
Hallucination Rate
Latency
Cost
                      Before    After
Answer Accuracy        72%      95%
Groundedness           69%      97%
Tool Success           81%      98%
Hallucination Rate     17%       3%
Change
  ↓
Test
  ↓
Measure
  ↓
Compare
  ↓
Improve

Putting It All Together

Putting It All Together works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

                         ┌──────────────────┐
                         │ Knowledge Graph  │
                         └────────┬─────────┘
                                  │
                         ┌────────▼─────────┐
                         │   RAG / Search   │
                         └────────┬─────────┘
                                  │
       Customer → Intent → Context Engine → LLM
                     ↑             │
                     │             ▼
                   Memory      Tool Selection
                     ↑             │
                     │             ▼
                     │          Validation
                     │             │
                     │             ▼
                     │        Real Systems
                     │             │
                     │             ▼
                     └─────── Feedback
                                   │
                                   ▼
                               Evaluation

The Bigger Lesson

The Bigger Lesson works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

The LLM Is Only One Part of the System

The LLM Is Only One Part of the System works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The LLM Is Only One Part of the System works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

              ┌───────────────────┐
              │ Knowledge + RAG   │
              └─────────┬─────────┘
                        ↓
Customer → Context → LLM → Tools → Real World
            ↑         ↓      ↓
          Memory   Reasoning Validation
            ↑         ↓
            └──── Feedback
                    ↓
                Evaluation

Final Thought

For Final Thought, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

Operational checklist

Operational checklist works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.

Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.

Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.

Batch note for d3bfab080c19: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.