Home / Articles / Practical notes: Build a Simple RAG Application in Google Colab with LlamaIndex

This article is published in English.

Practical notes: Build a Simple RAG Application in Google Colab with LlamaIndex

Operable walkthrough of Practical notes: Build a Simple RAG Application in Google Colab with LlamaIndex: contracts, checks, and drop-in code slots for teams shipping this pattern.

1221 words

This walkthrough rebuilds the path from raw materials to a working system for: Build a Simple RAG Application in Google Colab with LlamaIndex and an Open-Source LLM. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

1. Install the required libraries

When working through the 1 Install the required stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

!pip install -q llama-index llama-index-readers-web html2text llama-index-llms-groq llama-index-embeddings-huggingface

2. Import the required packages

When working through the 2 Import the required stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.

from llama_index.core import VectorStoreIndex, Settings
from llama_index.readers.web
import SimpleWebPageReader
from llama_index.llms.groq import Groq
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from google.colab import userdata

3. Configure the Groq API key

When working through the 3 Configure the Groq stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the 3 Configure the Groq stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# Groq API Key
os.environ["GROQ_API_KEY"] = userdata.get("GROQ_APIKEY")

4. Configure the LLM and embedding model

The 4 Configure the LLM stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

# Set up the open-source LLM and embedding model
Settings.llm = Groq( model="openai/gpt-oss-120b", temperature=0.1 )
Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5" )

5. Load a web page

The 5 Load a web stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.

# Passing a URL which we want to load to our vector store
url = "https://mlds.analyticsindiamag.com/"
# Using SimpleWebPageReader to load the URL content
# html_to_text=True converts HTML into plain text
d1 = SimpleWebPageReader( html_to_text=True ).load_data([url])

6. Create the vector index

The 6 Create the vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The 6 Create the vector stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.

# Create a searchable index from the loaded document
index = VectorStoreIndex.from_documents(d1)

7. Create the query engine

For the 7 Create the query stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

# Creating query engine
query_engine = index.as_query_engine()

8. Ask a question

For the 8 Ask a question stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.

# Running a query against the loaded URL data
r1 = query_engine.query("What is MLDS?")
print(r1)

The Complete RAG Flow

For the The Complete RAG Flow stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state.

Web Page
   ↓
SimpleWebPageReader
   ↓
Extract Text
   ↓
Hugging Face Embeddings
   ↓
VectorStoreIndex
   ↓
User Question
   ↓
Relevant Context
   ↓
Groq LLM
   ↓
Answer

What’s Next?

Operational checklist