This article is published in English.
Practical notes: Zero-Cost Local Multimodal RAG: Selective Vision Processing
Operable walkthrough of Practical notes: Zero-Cost Local Multimodal RAG: Selective Vision Processing: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: Zero-Cost Local Multimodal RAG: Selective Vision Processing with ChromaDB. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
The Core Challenges (On Apple Silicon + Beyond)
When working through the The Core Challenges On stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
The Selective Hybrid Pipeline (Optimized for Apple Silicon)
When working through the The Selective Hybrid Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
System Architecture
When working through the System Architecture stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
┌───────────────────────────────────────┐
│ Local PDF Document Store (M1 Mac) │
└───────────────────┬───────────────────┘
│
┌───────────────┴───────────────┐
│ Fast Layout-Aware Parser │
└───────┬───────────────┬───────┘
│ │
[Text & Tables] │ │ [Embedded Images]
▼ ▼
┌───────────────────┐ ┌───────────────────┐
│ Markdown Stream │ │ Cropped Images │
└─────────┬─────────┘ └─────────┬─────────┘
│ │
│ ▼
│ ┌───────────────────┐
│ │ Base64 Scaling & │
│ │ Native BBox OCR │
│ └─────────┬─────────┘
│ │
│ ▼
│ ┌───────────────────┐
│ │ Local Metal VLM │
│ │ (Ollama via UMA) │
│ └─────────┬─────────┘
│ │
│ [Text Summaries]
│ │
▼ ▼
┌───────────────────────────────────┐
│ Unified Chunking & Context Engine │
└─────────────────┬─────────────────┘
│
▼
┌───────────────────────────────────┐
│ Disk-Persisted Vector DB & Parent │
│ Context Stores (ChromaDB SQLite) │
└───────────────────────────────────┘
Hardening the Pipeline
When working through the Hardening the Pipeline stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Vector Database Choice
When working through the Vector Database Choice stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Parent-Child Retrieval (The Secret Sauce)
When working through the Parent-Child Retrieval The Secret stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Enforcing Accuracy: Structured Outputs & Verification
When working through the Enforcing Accuracy Structured Outputs stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Complete Project Setup & Code (Copy-Paste Ready)
When working through the Complete Project Setup Code stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Complete Project Setup Code stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
1. Project Structure
The 1 Project Structure stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
mkdir ~/m1_multimodal_rag && cd ~/m1_multimodal_rag
mkdir data chroma_db images_cache
touch main.py requirements.txt
2. requirements.txt
The 2 requirements txt stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
pymupdf
chromadb
ollama
pillow
3. The Complete main.py (Unified Ingestion + Query)
The 3 The Complete main stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The 3 The Complete main stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
#!/usr/bin/env python3
"""
Multimodal RAG for Apple Silicon (M1/M2/M3)
Usage:
python main.py ingest --pdf data/report.pdf
python main.py ingest --pdf data/new_report.pdf --clear
python main.py query --question "What was the Q3 revenue?"
"""
import argparse
import base64
import os
import sys
import uuid
from io import BytesIO
from pathlib import Path
import chromadb
import pymupdf as fitz
import ollama
from PIL import Image
# ---------- CONFIG ----------
TEXT_MODEL = "llama3.2:3b" # For final RAG answers (8GB friendly)
VISION_MODEL = "qwen2.5vl:3b" # For charts (8GB friendly)
CHROMA_PATH = "./chroma_db"
IMAGE_CACHE = "./images_cache"
# Initialize persistent Chroma client (SQLite, not RAM)
chroma_client = chromadb.PersistentClient(path=CHROMA_PATH)
child_collection = chroma_client.get_or_create_collection(name="child_chunks")
parent_collection = chroma_client.get_or_create_collection(name="parent_chunks")
Path(IMAGE_CACHE).mkdir(exist_ok=True)
# ---------- HELPER: Encode Image for Ollama ----------
def encode_image_for_ollama(image_bytes: bytes, max_size=800) -> str:
"""Convert PDF image bytes to Base64 data URI with size limiting."""
img = Image.open(BytesIO(image_bytes))
# Convert RGBA/P to RGB to avoid JPEG alpha errors
if img.mode in ('RGBA', 'LA', 'P'):
img = img.convert('RGB')
# Downscale massive images to save VRAM on M1
img.thumbnail((max_size, max_size))
buffered = BytesIO()
img.save(buffered, format="JPEG", quality=85)
img_base64 = base64.b64encode(buffered.getvalue()).decode('utf-8')
return img_base64
# ---------- PHASE 1: INGESTION ----------
def ingest_pdf(pdf_path: str):
"""Parse PDF, extract text, crop images, run VLM, and store in Chroma."""
print(f" Processing: {pdf_path}")
doc = fitz.open(pdf_path)
for page_num in range(len(doc)):
page = doc[page_num]
print(f" Page {page_num + 1}/{len(doc)}")
# 1. Extract main text
page_text = page.get_text("text").strip()
if not page_text:
page_text = "[No extractable text on this page]"
# 2. Find and process images
image_list = page.get_images(full=True)
visual_summaries = []
for img_idx, img in enumerate(image_list):
xref = img[0]
try:
base_image = doc.extract_image(xref)
image_bytes = base_image["image"]
# Encode for Ollama
encoded_img = encode_image_for_ollama(image_bytes)
# Prompt designed for financial charts with structured output
prompt = """
Extract key insights from this chart and return valid JSON.
Use this schema: {"chart_type": "", "x_axis": [], "y_axis": [], "key_trend": "", "data_points": []}
If it's not a chart, describe it briefly in text.
"""
response = ollama.chat(
model=VISION_MODEL,
messages=[{
"role": "user",
"content": prompt,
"images": [encoded_img]
}]
)
summary = response["message"]["content"]
visual_summaries.append(f"[Chart on page {page_num+1}]: {summary}")
except Exception as e:
print(f" Skipped image {img_idx} (Error: {e})")
continue
# 3. Merge text and summaries
full_page_content = page_text + "\n" + "\n".join(visual_summaries)
if not full_page_content.strip():
continue # Skip completely empty pages
# 4. Split into Parent (big) and Child (small) for retrieval
parent_text = full_page_content # Full page is the "Parent"
# Split into ~200 token chunks for children (roughly 800 chars)
child_chunks = []
chunk_size = 800
for i in range(0, len(parent_text), chunk_size):
child_chunks.append(parent_text[i:i+chunk_size])
if not child_chunks:
child_chunks = [parent_text] # Fallback
# 5. Store in Chroma (Parent-Child)
parent_id = str(uuid.uuid4())
metadata = {
"source": os.path.basename(pdf_path),
"page": page_num + 1,
"type": "hybrid"
}
# Store Parent (full context) - Persisted to disk, not RAM
parent_collection.add(
ids=[parent_id],
documents=[parent_text],
metadatas=[metadata]
)
# Store Children (granular search)
child_ids = []
child_metadatas = []
for idx, chunk in enumerate(child_chunks):
child_id = f"{parent_id}_child_{idx}"
child_ids.append(child_id)
child_metadatas.append({
**metadata,
"parent_ref": parent_id
})
child_collection.add(
ids=child_ids,
documents=child_chunks,
metadatas=child_metadatas
)
doc.close()
print(" Ingestion complete!")
# ---------- PHASE 2: QUERY ----------
def query_rag(question: str):
"""Retrieve relevant context using Child chunks, fetch Parent, ask LLM."""
print(f"❓ Query: {question}")
# 1. Retrieve top matching child chunks
results = child_collection.query(
query_texts=[question],
n_results=3
)
if not results["ids"] or not results["ids"][0]:
print(" No relevant documents found in the database.")
return
# 2. Fetch the full Parent contexts
parent_ids = list(set([m["parent_ref"] for m in results["metadatas"][0]]))
parent_results = parent_collection.get(ids=parent_ids)
full_context = "\n\n---\n\n".join(parent_results["documents"])
# 3. Build prompt for the text-only LLM
prompt = f"""
You are a financial research assistant. Answer the question based strictly on the context below.
If the context contains chart summaries or tables, use those numbers specifically.
If you cannot answer from the context, say "I don't have that information."
Context:
{full_context}
Question: {question}
Answer:
"""
# 4. Generate answer locally
response = ollama.chat(
model=TEXT_MODEL,
messages=[{"role": "user", "content": prompt}]
)
print("\n Answer:")
print(response["message"]["content"])
print("\n Sources:", ", ".join(parent_ids))
# ---------- DATABASE CLEAR FUNCTION ----------
def clear_database():
"""Delete all collections to reset the database."""
try:
chroma_client.delete_collection("child_chunks")
chroma_client.delete_collection("parent_chunks")
print(" Database cleared successfully!")
except ValueError:
print(" Database was already empty. Nothing to clear.")
except Exception as e:
print(f" Could not clear database: {e}")
# ---------- CLI ENTRY POINT ----------
def main():
parser = argparse.ArgumentParser(description="M1 Multimodal RAG Pipeline")
subparsers = parser.add_subparsers(dest="command", required=True)
# Ingest command with --clear flag
ingest_parser = subparsers.add_parser("ingest", help="Ingest a PDF")
ingest_parser.add_argument("--pdf", required=True, help="Path to PDF file")
ingest_parser.add_argument("--clear", action="store_true", help="Clear the database before ingesting")
# Query command
query_parser = subparsers.add_parser("query", help="Ask a question")
query_parser.add_argument("--question", required=True, help="Your question")
args = parser.parse_args()
if args.command == "ingest":
if not os.path.exists(args.pdf):
print(f" File not found: {args.pdf}")
sys.exit(1)
# Clear the database if the flag is set
if args.clear:
clear_database()
# Re-initialize collections after clearing
global child_collection, parent_collection
child_collection = chroma_client.get_or_create_collection(name="child_chunks")
parent_collection = chroma_client.get_or_create_collection(name="parent_chunks")
ingest_pdf(args.pdf)
elif args.command == "query":
query_rag(args.question)
if __name__ == "__main__":
main()
Step-by-Step Execution Commands
For the Step-by-Step Execution Commands stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
Step 1: Install Ollama & Pull Models
For the Step 1 Install Ollama stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
# Install Ollama via Homebrew
brew install ollama
# Verify version (requires >= 0.7.0 for qwen2.5vl models)
ollama --version
# Start the Ollama service (keep this running in a separate terminal tab)
ollama serve
# Pull the recommended models (For 8GB M1 Mac)
ollama pull qwen2.5vl:3b
ollama pull llama3.2:3b
# (For 16GB+ M1/M2/M3, optionally pull larger models)
# ollama pull llama3.2-vision:11b
# ollama pull qwen2.5:7b
# displays a list of all AI models stored locally on your machine
ollama list
Step 2: Set Up the Python Environment
For the Step 2 Set Up stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
# Ensure you are in the project root
cd ~/m1_multimodal_rag
# Create a virtual environment
python3 -m venv venv
# Activate the environment
source venv/bin/activate
Step 3: Install Python Dependencies
For the Step 3 Install Python stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate client construction from the message loop so providers can be swapped without rewriting the conversation state machine.
# Upgrade pip
pip install --upgrade pip
# Install requirements
pip install -r requirements.txt
# Verify installation
python -c "import chromadb, fitz, ollama, PIL; print(' All dependencies ready!')"
Step 4: Download a Sample PDF
For the Step 4 Download a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
# Download a sample financial report (EY IFRS Illustrative)
curl -L -o data/sample_financials.pdf \
"https://drive.google.com/uc?export=download&id=1OOE1vPBwPP31cB0_MNgooTrhw6KpY6n_"
Step 5: Ingest the PDF
For the Step 5 Ingest the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
python main.py ingest --pdf data/sample_financials.pdf
# Ingest a new PDF and clear the database first
python main.py ingest --pdf data/<your financial data file>.pdf --clear
Step 6: Ask a Question
For the Step 6 Ask a stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cite the passages that actually grounded the answer. Without citations, operators cannot tell hallucination from an indexing gap.
python main.py query --question "What was the total revenue shown in the financial statements?"
git clone https://github.com/froilan-sia/m1_multimodal_rag.git
cd m1_multimodal_rag
./setup.sh
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 313930800633: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.