Home / Articles / Practical notes: From Fine-Tuning to Precision: Significantly Reducing

This article is published in English.

Practical notes: From Fine-Tuning to Precision: Significantly Reducing

Operable walkthrough of Practical notes: From Fine-Tuning to Precision: Significantly Reducing: contracts, checks, and drop-in code slots for teams shipping this pattern.

2822 words

This walkthrough rebuilds the path from raw materials to a working system for: From Fine-Tuning to Precision: Significantly Reducing Hallucinations in Your RAG Pipeline. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent.

1. Introduction

For the 1 Introduction stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

You: What was the revenue from contracts with customers in 2024?
Assistant: The revenue was €179,058,000.  ← Wrong! Correct answer is €159,088,000.

2. Why Fine-Tuning?

For the 2 Why Fine-Tuning stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

The Solution: Fine-Tuning with MLX

For the The Solution Fine-Tuning with stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

3. Architecture Diagram

For the 3 Architecture Diagram stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.

┌─────────────────────────────────────────────────────────────────────────────┐
│                       PART 3: FINE-TUNING WORKFLOW (M1)                     │
├─────────────────────────────────────────────────────────────────────────────┤
│                                                                             │
│   1. Dataset                2. MLX Fine-Tuning                              │
│  ┌──────────────────┐      ┌────────────────────────────────────────────┐   │
│  │ Chart Images     │─────▶│ Base Model (Qwen2-VL-2B)                   │   │
│  │ Q&A Pairs        │      │  + LoRA Adapters (mlx_vlm.lora)            │   │
│  │ (train.jsonl)    │      │  + Unified Memory Training on Apple Silicon│   │
│  └──────────────────┘      └────────────────┬───────────────────────────┘   │
│                                             │                               │
│                                             ▼                               │
│                            ┌──────────────────────────────────────────┐     │
│                            │ Fine-Tuned LoRA Adapters                 │     │
│                            │ (./fine_tuned_adapters/)                 │     │
│                            └────────────────┬─────────────────────────┘     │
│                                             │                               │
│                                             ▼                               │
│                            ┌──────────────────────────────────────────┐     │
│                            │ 3. Merge & Export to GGUF                │     │
│                            │ (mlx_vlm.fuse + convert_hf_to_gguf.py)   │     │
│                            └────────────────┬─────────────────────────┘     │
│                                             │                               │
│                                             ▼                               │
│                            ┌──────────────────────────────────────────┐     │
│                            │ 4. Deploy with Ollama                    │     │
│                            │ ollama create my-chart-model             │     │
│                            │ (text.gguf + mmproj.gguf)                │     │
│                            └────────────────┬─────────────────────────┘     │
│                                             │                               │
│                                             ▼                               │
│  ┌──────────────────────────────────────────────────────────────────────┐   │
│  │ 5. Update rag_engine.py                                              │   │
│  │                                                                      │   │
│  │   VISION_MODEL = "my-chart-model"  # ← Change ONE line               │   │
│  │   TEXT_MODEL   = "llama3.2:3b"     # Unchanged                       │   │
│  │                                                                      │   │
│  │   ✔ Existing app.py (from Part 2) automatically uses the new model!  │   │
│  └──────────────────────────────────────────────────────────────────────┘   │
└─────────────────────────────────────────────────────────────────────────────┘

4. Preparing the Dataset

For the 4 Preparing the Dataset stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap. For the 4 Preparing the Dataset stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

Model sees:    (Chart image)
Model reads:   (Question)
Model learns:  (Correct answer)

Dataset Sources

When working through the Dataset Sources stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Step-by-Step Data Preparation

When working through the Step-by-Step Data Preparation stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.

Step 1: Create Raw Dataset       →  chart_dataset/raw_train.jsonl
Step 2: Convert to MLX Format    →  chart_dataset/train.jsonl
Step 3: Verify Dataset           →  Check that train.jsonl exists

Step 1: Create the Raw Dataset

When working through the Step 1 Create the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Step 1 Create the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.

# download_chartqa.py
from datasets import load_dataset
import json
import os

os.makedirs("chart_dataset", exist_ok=True)

# Download ChartQA dataset
dataset = load_dataset("ahmed-masry/ChartQA", split="train")

# Convert to flat format with standard keys
converted = []
for item in dataset:
    converted.append({
        "image": item["imgname"],      # Path to chart image
        "question": item["query"],
        "answer": item["label"]
    })

# Save as raw dataset
with open("chart_dataset/raw_train.jsonl", "w") as f:
    for entry in converted:
        f.write(json.dumps(entry) + "\n")

print(f"✅ Converted {len(converted)} ChartQA samples to chart_dataset/raw_train.jsonl")
README.md: 100%|█████████████████████████████████████████████| 2.12k/2.12k [00:00<00:00, 3.72MB/s]
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
data/train-00000-of-00003.parquet: downloading bytes: ████████████████████████|  213MB, 17.3MB/s
......
Generating test split: 100%|████████████████████████| 2500/2500 [00:00<00:00, 20517.87 examples/s]
✅ Converted 28299 ChartQA samples to raw_train.jsonl

Option B: Generate Synthetic Data

The Option B Generate Synthetic stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

# generate_synthetic_data.py
import json
import matplotlib.pyplot as plt
import numpy as np
import os

os.makedirs("chart_dataset/images", exist_ok=True)

dataset = []
for i in range(100):
    categories = ['Q1', 'Q2', 'Q3', 'Q4']
    values = np.random.randint(100, 500, 4)

    plt.figure()
    plt.bar(categories, values)
    plt.title(f"Quarterly Revenue {i}")
    plt.savefig(f"chart_dataset/images/chart_{i:03d}.png")
    plt.close()

    dataset.append({
        "image": f"images/chart_{i:03d}.png",
        "question": "Which quarter had the highest revenue?",
        "answer": f"Q{np.argmax(values) + 1} with ${max(values)} million"
    })

with open("chart_dataset/raw_train.jsonl", "w") as f:
    for entry in dataset:
        f.write(json.dumps(entry) + "\n")

print("✅ Generated 100 synthetic samples in chart_dataset/raw_train.jsonl")

Step 2: Convert to MLX Format

The Step 2 Convert to stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

# convert_to_mlx_format.py
import json
import os

def convert_raw_to_mlx_format(data_dir="chart_dataset"):
    """
    Convert raw dataset to MLX format.

    Input:  chart_dataset/raw_train.jsonl
    Output: chart_dataset/train.jsonl (MLX-compatible)
    """
    raw_file = os.path.join(data_dir, "raw_train.jsonl")
    out_file = os.path.join(data_dir, "train.jsonl")

    if not os.path.exists(raw_file):
        print(f"❌ {raw_file} not found. Run download_chartqa.py or generate_synthetic_data.py first.")
        return None

    dataset = []
    with open(raw_file, "r") as f:
        for line in f:
            item = json.loads(line)

            # Build full image path
            img_str = item["image"]
            image_path = img_str if img_str.startswith(data_dir) else os.path.join(data_dir, img_str)

            dataset.append({
                "images": [image_path],  # MUST be a list
                "messages": [
                    {"role": "user", "content": item["question"]},
                    {"role": "assistant", "content": item["answer"]}
                ]
            })

    with open(out_file, "w") as f:
        for entry in dataset:
            f.write(json.dumps(entry) + "\n")

    print(f"✅ Converted {len(dataset)} samples to {out_file}")
    return dataset

if __name__ == "__main__":
    convert_raw_to_mlx_format()
python convert_to_mlx_format.py

Step 3: Verify the Dataset

The Step 3 Verify the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

ls -la chart_dataset/
chart_dataset/
├── raw_train.jsonl     # Raw Q&A pairs (flat keys)
├── train.jsonl         # MLX-formatted (overwritten by prepare_mlx_dataset)
└── images/             # Chart images

📌 Important: Before running the fine-tuning command, ensure train.jsonl exists in your dataset directory.
If you see a FileNotFoundError, you haven't run the conversion step yet.

5. Fine-Tuning with MLX

The 5 Fine-Tuning with MLX stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Prerequisites

The Prerequisites stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

# For fine-tuning vision models
pip install mlx-vlm

# For merging adapters (needed after training)
pip install mlx-lm

Run Fine-Tuning

The Run Fine-Tuning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

python -m mlx_vlm.lora \
    --model Qwen/Qwen2-VL-2B-Instruct \
    --dataset ./chart_dataset/train.jsonl \
    --iters 1000 \
    --batch-size 1 \
    --lora-rank 8 \
    --gradient-accumulation-steps 4 \
    --max-seq-length 512

Understanding Training Output

The Understanding Training Output stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

INFO:__main__:Loading model from Qwen/Qwen2-VL-2B-Instruct
Fetching 11 files: 100%|████████████████████████████████████████| 11/11 [00:00<00:00, 1465.47it/s]
Download complete: :                                                          |  0.00B
Reconstruction complete: |                                           |  0.00B /  0.00B
INFO:__main__:Loading dataset from ./chart_dataset/train.jsonl
INFO:__main__:Setting up LoRA
#trainable params: 9.232384 M || all params: 2208.9856 M || trainable%: 0.418%
INFO:__main__:Setting up optimizer
INFO:__main__:Training model (sft)
Starting training..., iterations: 1000
No validation dataset provided — training will run without validation.
......

Iter 120: Train loss 7.77592850, Learning Rate 2.000e-05,
It/sec 0.242, Tokens/sec 103.862, Trained Tokens 51480, Peak mem 8.069 GB
.....

Saved final adapter weights to adapters.safetensors.
INFO:__main__:Training completed! Model saved to adapters.safetensors

6. Troubleshooting

The 6 Troubleshooting stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.

Dataset Not Found Error

The Dataset Not Found Error stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.

FileNotFoundError: Couldn't find any data file at .../chart_dataset/train.jsonl
python convert_to_mlx_format.py

Out of Memory (OOM) Error

RuntimeError: [METAL] Command buffer execution failed: Insufficient Memory
python -m mlx_vlm.lora \
    --model Qwen/Qwen2-VL-2B-Instruct \
    --dataset ./chart_dataset/train.jsonl \
    --iters 1000 \
    --batch-size 1 \              # ← Reduced from 4 to 1
    --lora-rank 4 \               # ← Reduced from 8 to 4
    --gradient-accumulation-steps 8 \  # ← Added
    --max-seq-length 256               # ← Added

Adapter Path Not Found

FileNotFoundError: The adapter path does not exist: fine_tuned_adapters
ls -la adapters.safetensors adapter_config.json
python -m mlx_lm fuse \
    --model Qwen/Qwen2-VL-2B-Instruct \
    --adapter-path . \
    --save-path ./fine_tuned_model_merged

Incomplete Cache Error

IncompleteSnapshotError: The cached snapshot for 'Qwen/Qwen2-VL-2B-Instruct' is incomplete
rm -rf ~/.cache/huggingface/hub/models--Qwen--Qwen2-VL-2B-Instruct

7. How Long Will Training Take?

Total Time (seconds) = Total Iterations ÷ It/sec
Total Time (minutes) = Total Time (seconds) ÷ 60
1000 ÷ 0.242 = 4,132 seconds
4,132 ÷ 60 = ~69 minutes

8. Merging LoRA Adapters

python -m mlx_lm fuse \
    --model Qwen/Qwen2-VL-2B-Instruct \
    --adapter-path . \
    --save-path ./fine_tuned_model_merged

9. Export to GGUF for Ollama

Step 1: Export MLX to Hugging Face Format

#!/usr/bin/env python3
import mlx_lm
from mlx_lm import load, save

model, tokenizer, config = load("./fine_tuned_model_merged")
save(model, tokenizer, config, "./hf_export")
print("✅ Exported to ./hf_export")
python export_to_hf.py

Step 2: Convert to GGUF

# Clone llama.cpp (if not already done)
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp

# Convert text model
python convert_hf_to_gguf.py ../hf_export \
    --outfile chart_model-text.gguf \
    --outtype f16

Step 3: Convert the Vision Projector

python convert_hf_to_gguf.py ../hf_export \
    --outfile chart_model-mmproj.gguf \
    --outtype f16 \
    --mmproj
ls ~/.cache/huggingface/hub/models--Qwen--Qwen2-VL-2B-Instruct/snapshots/

# sample output . You'll see a hash directory (e.g., 895c3a49...)
# 895c3a49bc3fa70a340399125c650a463535e71c
cd llama.cpp
# replace <hash> with the output above (e.g., 895c3a49...)

python convert_hf_to_gguf.py ~/.cache/huggingface/hub/models--Qwen--Qwen2-VL-2B-Instruct/snapshots/<hash>/ \
    --outfile base-mmproj.gguf \
    --outtype f16 \
    --mmproj

10. Deploying with Ollama

FROM ./chart_model-text.gguf
FROM ./chart_model-mmproj.gguf
PARAMETER temperature 0.2

Create the Model

ollama create my-chart-model -f Modelfile

Verify the Model

ollama list

# output :

# NAME                     ID              SIZE      MODIFIED
# my-chart-model:latest    b3fd6d8fd742    4.4 GB    7 hours ago
# qwen2.5vl:3b             fb90415cde1e    3.2 GB    9 days ago
# llama3.2:3b              a80c4f17acd5    2.0 GB    10 days ago
# llama3:latest            365c0bd3c000    4.7 GB    3 months ago

# You should see my-chart-model in the list.

Test the Model

ollama run my-chart-model "What is 2+2?"

11. Integrate into Your Existing RAG App

# In rag_engine.py (from Parts 1 & 2)

# Before:
VISION_MODEL = "qwen2.5vl:3b"

# After:
VISION_MODEL = "my-chart-model"  # Your fine-tuned model
TEXT_MODEL = "llama3.2:3b"       # Unchanged
python app.py

12. Evaluation: Base vs. Fine-Tuned

Comparison

13. Conclusion

Get the Complete Code

git clone https://github.com/froilan-sia/m1_multimodal_rag.git
cd m1_multimodal_rag

Connect with Me

Operational checklist