Home / Articles / Wrapping Multimodal RAG in a Gradio Chat UI

This article is published in English.

Wrapping Multimodal RAG in a Gradio Chat UI

Refactor a CLI multimodal RAG engine into a drag-and-drop Gradio app with progress, citations, and streaming answers.

1866 words

From CLI multimodal RAG to a Gradio chat UI

Part 1 delivered a command-line multimodal RAG stack: ingest PDFs, extract text and charts, answer locally. Powerful — and awkward for non-technical colleagues. This follow-up wraps the same engine in a web UI: drag-and-drop a PDF, watch progress, chat with citations.

Recap of the CLI core

The engine already knew how to ingest, index, retrieve parent/child chunks, and answer with sources. The UI should not fork that logic — it should call it.

Refactor: rag_engine.py

Pull ingest and query into clear functions the UI can call:

# rag_engine.py - Core functions

def ingest_pdf(pdf_path: str, progress=None, status=None):
    """
    Parse PDF, extract text, process images, and store in Chroma.

    Args:
        pdf_path: Path to the PDF file
        progress: Gradio progress object (optional)
        status: Callback for status updates (optional)
    """
    # ... same logic as Part 1, but with progress callbacks

def query_rag(question: str) -> dict:
    """
    Retrieve relevant context and generate an answer.

    Args:
        question: User's question

    Returns:
        dict: {"answer": str, "sources": list[str]}
    """
    # ... returns answer with source parent IDs

def clear_database():
    """Reset the database and reinitialize collections."""
    # ... handles collection re-creation

def get_ingested_documents():
    """Return list of ingested PDFs."""
    # ... utility for status display

Progress callbacks

Long PDF ingest needs status updates in the browser:

# In rag_engine.py
def ingest_pdf(pdf_path: str, progress=None, status=None):
    if status:
        status(f"📄 Processing page {page_num + 1}/{total_pages}...")
    if progress:
        progress((page_num + 1) / total_pages, desc=f"Processing page {page_num + 1}/{total_pages}")
    # ... rest of the logic

Pass progress / status callables from Gradio so users see page counts instead of a frozen spinner.

Metadata for citations

Store readable source fields while chunking:

# During ingestion in rag_engine.py
metadata = {
    "source": os.path.basename(pdf_path),
    "page": page_num + 1,
    "parent_id": f"Page_{page_num + 1}",
    "section": "Unknown",  # Optional: extract from heading
    "chunk_index": chunk_idx,
    "type": "hybrid"
}

Parent-child retrieval (short recap)

Child chunks retrieve precisely; parent windows supply context for generation. The UI should surface the same sources the CLI printed.

Query path

def query_rag(question: str) -> dict:
    # 1. Retrieve top matching child chunks
    results = child_collection.query(
        query_texts=[question],
        n_results=3
    )

    # 2. Fetch the full Parent contexts with metadata
    parent_ids = list(set([m["parent_ref"] for m in results["metadatas"][0]]))
    parent_results = parent_collection.get(ids=parent_ids)

    # 3. Build full context from parent documents
    full_context = "\n\n---\n\n".join(parent_results["documents"])

    # 4. Get metadata for source display
    source_metadata = []
    for meta in parent_results["metadatas"]:
        page = meta.get("page", "Unknown")
        section = meta.get("section", "Section")
        source_metadata.append({"page": page, "section": section})

    # 5. Build prompt and generate answer
    prompt = f"""
    You are a financial research assistant. Answer the question based strictly on the context below.
    If the context contains chart summaries or tables, use those numbers specifically.
    If you cannot answer from the context, say "I don't have that information."

    Context:
    {full_context}

    Question: {question}
    Answer:
    """

    response = ollama.chat(
        model=TEXT_MODEL,
        messages=[{"role": "user", "content": prompt}]
    )

    return {
        "answer": response["message"]["content"],
        "sources": source_metadata  # Now contains page and section info
    }

Retrieve children, expand parents, build the prompt, return answer + sources.

Rendering sources in the UI

# In app.py
if result["sources"]:
    source_text = "\n\n📚 **Sources:** "
    sources_list = []
    for source in result["sources"][:3]:
        # Assuming source contains metadata
        page = source.get("page", "Unknown")
        section = source.get("section", "Section")
        sources_list.append(f"Page {page} ({section})")
    answer += source_text + ", ".join(sources_list)
if result["sources"]:
    source_text = "\n\n **Sources:** "
    sources_list = []
    for source in result["sources"][:3]:
        page = source.get("page", "Unknown")
        section = source.get("section", "Section")
        sources_list.append(f"Page {page} ({section})")
    answer += source_text + ", ".join(sources_list)

Example exchange shape:

User: What was the total revenue shown in the financial statements?

The total revenue shown in the financial statements is €180,462 for the year ended 31 December 2025 and €160,465 for the year ended 31 December 2024.

📚 Sources: Page 35 (Section 1), Page 79 (Section 2), Page 159 (Section 3)
User: What was the total revenue shown in the financial statements?

The total revenue shown in the financial statements is €180,462 for the year ended 31 December 2025 and €160,465 for the year ended 31 December 2024.

📚 Sources: Page 35 (Section 1), Page 79 (Section 2), Page 159 (Section 3)

Building app.py with Gradio

Gradio gives upload widgets, chat history, and streaming with little boilerplate.

Streaming answers

Prefer yielding tokens for perceived speed:

def chat_response(message, history):
    # ... retrieve context ...

    # Instead of returning, use yield to stream tokens
    full_response = ""
    for chunk in ollama.chat(model=TEXT_MODEL, messages=[...], stream=True):
        full_response += chunk["message"]["content"]
        yield history + [("user", message), ("assistant", full_response)]

Full app sketch

#!/usr/bin/env python3
"""
Multimodal RAG Gradio UI
Run with: python app.py
"""

import os
import shutil
from pathlib import Path

import gradio as gr

# Import the core engine
from rag_engine import ingest_pdf, query_rag, clear_database, get_ingested_documents

# ---------- CONFIG ----------
UPLOAD_DIR = Path("./data")
UPLOAD_DIR.mkdir(exist_ok=True)

# ---------- UI FUNCTIONS ----------
def process_upload(file_obj, progress=gr.Progress()):
    """Handle PDF upload and ingestion with progress bar."""
    if file_obj is None:
        return " Please upload a PDF file first."

    pdf_path = UPLOAD_DIR / os.path.basename(file_obj.name)
    shutil.copy(file_obj.name, pdf_path)

    try:
        result = ingest_pdf(str(pdf_path), progress=progress)
        return f"{result}\n\n📄 File saved to: {pdf_path}"
    except Exception as e:
        return f" Error during ingestion: {str(e)}"

def chat_response(message, history):
    """
    Handle user questions and return responses.

    Note: This uses synchronous return. For streaming responses,
    consider using yield with ollama's stream=True parameter.
    """
    if not message or not message.strip():
        return history

    docs = get_ingested_documents()
    if not docs:
        history.append({"role": "user", "content": message})
        history.append({"role": "assistant", "content": " No documents ingested. Please upload and ingest a PDF first."})
        return history

    result = query_rag(message)
    answer = result["answer"]

    # Display readable sources with page and section info
    if result["sources"]:
        source_text = "\n\n **Sources:** "
        sources_list = []
        for source in result["sources"][:3]:
            page = source.get("page", "Unknown")
            section = source.get("section", "Section")
            sources_list.append(f"Page {page} ({section})")
        answer += source_text + ", ".join(sources_list)

    history.append({"role": "user", "content": message})
    history.append({"role": "assistant", "content": answer})
    return history

def reset_database():
    """Clear the database, chat history, and reset file upload."""
    result = clear_database()
    return result, [], gr.update(value=None)

def get_status():
    """Get current system status."""
    docs = get_ingested_documents()
    if docs:
        return f" {len(docs)} document(s) ingested: {', '.join(docs)}"
    return " No documents ingested. Upload and ingest a PDF to get started."

# ---------- BUILD UI ----------
def create_ui():
    with gr.Blocks(title="Multimodal RAG Assistant") as demo:
        gr.Markdown("""
        # Zero-Cost Local Multimodal RAG

        A privacy-first assistant that answers questions from complex PDFs containing text, tables, and charts.

        **How it works:**
        1. Upload a PDF and click **Ingest**
        2. Wait for processing (charts will be analyzed by the VLM)
        3. Ask questions about the document content

        **100% local** — No data ever leaves your machine.

        **Sources** shown in answers include page numbers and sections from your PDF.
        """)

        # Status Bar
        with gr.Row():
            status_bar = gr.Textbox(value=get_status(), label=" Status", interactive=False, scale=3)
            refresh_btn = gr.Button(" Refresh", size="sm", scale=0)
            clear_db_btn = gr.Button(" Clear Database", size="sm", variant="stop", scale=0)

        # Upload & Chat
        with gr.Row():
            with gr.Column(scale=1):
                gr.Markdown("### Upload & Ingest")
                file_upload = gr.File(label="Upload PDF", file_types=[".pdf"], height=100)
                ingest_btn = gr.Button(" Ingest PDF", variant="primary", size="lg")
                upload_status = gr.Textbox(label="Upload Status", interactive=False, lines=3)

            with gr.Column(scale=2):
                gr.Markdown("### 💬 Ask Questions")
                chatbot = gr.Chatbot(label="Chat", height=400, avatar_images=(None, "🤖"))
                with gr.Row():
                    msg = gr.Textbox(label="Your question", placeholder="e.g., What was the total revenue?", scale=4, container=False)
                    send_btn = gr.Button("Send", variant="primary", scale=0)
                clear_chat_btn = gr.Button(" Clear Chat", size="sm")

        # Event Handlers
        ingest_btn.click(fn=process_upload, inputs=[file_upload], outputs=[upload_status]).then(
            fn=get_status, inputs=[], outputs=[status_bar])
        msg.submit(fn=chat_response, inputs=[msg, chatbot], outputs=[chatbot]).then(fn=lambda: "", outputs=[msg])
        send_btn.click(fn=chat_response, inputs=[msg, chatbot], outputs=[chatbot]).then(fn=lambda: "", outputs=[msg])
        clear_chat_btn.click(fn=lambda: [], outputs=[chatbot])
        clear_db_btn.click(fn=reset_database, inputs=[], outputs=[upload_status, chatbot, file_upload]).then(
            fn=get_status, inputs=[], outputs=[status_bar])
        refresh_btn.click(fn=get_status, inputs=[], outputs=[status_bar])
        demo.load(fn=get_status, inputs=[], outputs=[status_bar])

    return demo

if __name__ == "__main__":
    print(" Starting Gradio UI...")
    print(" Opening at: http://127.0.0.1:7860")
    demo = create_ui()
    demo.launch(server_name="127.0.0.1", server_port=7860, share=False, theme=gr.themes.Soft(), css="footer {visibility: hidden}")
def chat_response(message, history):
    """
    Note: This uses synchronous return. For streaming responses,
    consider using yield with ollama's stream=True parameter.
    """

Run it

python app.py
git clone https://github.com/froilan-sia/m1_multimodal_rag.git
cd m1_multimodal_rag
python app.py

Clone, install deps from Part 1, launch app.py, upload a PDF, ask a grounded question, expand sources.

Closing

The CLI proved the pipeline; Gradio makes it shareable. Keep rag_engine.py as the single source of truth so UI work never forks retrieval quality.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.

Keep demos reproducible: pin model tags, document ports, and prefer mock modes when CI has no GPU.