This article is published in English.
Practical notes: A Step-by-Step Guide for Developing Your Personal Agentic
Operable walkthrough of Practical notes: A Step-by-Step Guide for Developing Your Personal Agentic: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: A Step-by-Step Guide for Developing Your Personal Agentic System.. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent.
Hands-on Tutorials
For the Hands-on Tutorials stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
A complete guide to learn how to set up and create your own agentic LLM system with local databases and specialized for your task.
For the A complete guide to stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
An Introduction Towards Agentic Systems.
For the An Introduction Towards Agentic stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Takeaway 1: Your LLM is Not a Writer; It’s a CPU.
For the Takeaway 1 Your LLM stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Takeaway 2: You Don’t Need More Billions of Parameters; You Need a Specialized “Agentic Team”.
For the Takeaway 2 You Don stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Takeaway 2 You Don stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The Real-World Problem: Human vs. Machine.
When working through the The Real-World Problem Human stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Why Single Prompts Fail and Agentic Systems Work.
When working through the Why Single Prompts Fail stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
What Is an Agent Architecture?
When working through the What Is an Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the What Is an Agent stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Retrieval-Augmented Generation (RAG) Pipeline.
The Retrieval-Augmented Generation RAG Pipeline stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
The challenges of RAG systems
The The challenges of RAG stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Four steps are required in the RAG pipeline
The Four steps are required stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move. The Four steps are required stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Ranking and Aggregation of the Chunks
For the Ranking and Aggregation of stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
The Final Step is Generation.
For the The Final Step is stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Tasks That Work Remarkably Well Using This Pipeline.
For the Tasks That Work Remarkably stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Tasks That Work Remarkably stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Global Reasoning Workflow
When working through the Global Reasoning Workflow stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
The Two-Step Global Reasoning Strategy
When working through the The Two-Step Global Reasoning stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Query Rewriting
When working through the Query Rewriting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Query Rewriting stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Local and Global Reasoning Together
The Local and Global Reasoning stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Small Language Models Are Winning.
The Small Language Models Are stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
Set Up Your Own Local Model Using LM Studio.
The Set Up Your Own stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices. The Set Up Your Own stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
The LLMlight Library.
For the The LLMlight Library stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Prefer structured outputs with schema validation over free-form prose when the next step is code or a tool call.
Chunking Strategy
For the Chunking Strategy stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Search Strategy — Local Databases.
For the Search Strategy Local Databases stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Search Strategy Local Databases stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Embedding & Scoring Strategies
When working through the Embedding Scoring Strategies stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
Context Strategies
When working through the Context Strategies stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Prompt Optimization
When working through the Prompt Optimization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn. When working through the Prompt Optimization stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Exercise 1: Load A Single Model and Have A Simple Chat.
The Exercise 1 Load A stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Budget tokens per turn and per session. Agentic tools expand context aggressively; hard caps keep demos from becoming surprise invoices.
# Install the library
pip install llmlight
from LLMlight import LLMlight
# Initialize the LLMlight client
client = LLMlight(model='google/gemma-4-26b-a4b-qat', endpoint="http://localhost:1234/v1/chat/completions")
# Ask a question
response = client.prompt('What is the capital of France?')
print(response)
# [LLMlight.LLM] [INFO ] Model : google/gemma-4-26b-a4b-qat
# [LLMlight.LLM] [INFO ] Context strategy : disabled
# [LLMlight.LLM] [INFO ] Retrieval method : naive_rag
# [LLMlight.LLM] [INFO ] Embedding : {'memory': 'bert', 'context': 'bert'}
# [LLMlight.LLM] [INFO ] Alpha (sig. test): None
# [LLMlight.LLM] [INFO ] Chunk config : {'method': 'chars', 'size': 1000, 'overlap': 200}
# [LLMlight.LLM] [INFO ] LLMlight initialised.
# [LLMlight.LLM] [INFO ] Creating response with google/gemma-4-26b-a4b-qat..
# [LLMlight.LLM] [INFO ] No context strategy applied.
# [LLMlight.LLM] [INFO ] No context is provided into the prompt.
# [LLMlight.LLM] [INFO ] Running model: google/gemma-4-26b-a4b-qat
# The capital of France is Paris.
Exercise 2: Create a Local Knowledge Base.
The Exercise 2 Create a stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# Import library
from LLMlight import LLMlight
# Initialize model and memory
client = LLMlight(model='google/gemma-4-26b-a4b-qat', endpoint="http://localhost:1234/v1/chat/completions")
# Create (or load) database
client.memory_init(store_path='knowledge_base.db')
# Add a PDF file to the database (extracts and chunks text automatically)
url = 'https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf'
pdf_text = client.read_pdf(url)
# Write to db
client.memory_add(text=pdf_text)
# Show the chunks
client.memory_chunks(1)
# Store to disk (SQLite DB is persisted automatically)
client.memory_save()
# Query on the new knowledge
response = client.prompt(query='What are attention networks?', response_format='Summarize in 3 sentences.')
print(response)
Exercise 3: The Differences in Output Using Context Strategy Methods.
The Exercise 3 The Differences stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Exercise 3 The Differences stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
from LLMlight import LLMlight
# Initialize with NO context strategy
client = LLMlight(model='google/gemma-4-26b-a4b-qat', retrieval_method='naive_rag', context_strategy=None, top_chunks=6, endpoint="http://localhost:1234/v1/chat/completions")
# Create (or load) database
client.memory_init(store_path='knowledge_base.db')
# Query on the new knowledge
response = client.prompt(query='What are attention networks?', response_format='Summarize in 2 sentence')
print(response)
# Based on the provided text, attention networks utilize mechanisms like self-attention
# to perform tasks such as reading comprehension, summarization, and machine translation
# by capturing the syntactic and semantic structures of sentences.
# They operate by computing attention weights on "values" using "queries" and "keys" through
# a softmax function, with common types being additive or dot-product multiplicative attention.
from LLMlight import LLMlight
# Initialize with CHUNK-WISE context strategy
client = LLMlight(model='google/gemma-4-26b-a4b-qat', retrieval_method='naive_rag', context_strategy='chunk-wise', top_chunks=6, endpoint="http://localhost:1234/v1/chat/completions")
# Create (or load) database
client.memory_init(store_path='knowledge_base.db')
# Query on the new knowledge
response = client.prompt(query='What are attention networks?', response_format='Summarize in 2 sentence')
# [LLMlight.LLM] [INFO ] Chunk wise analysis on 6 chunks of text.
# Processing chunk: 0%| | 0/6 [00:00<?, ?chunk/s][08-06-2026 22:11:53] [LLMlight.LLM] [INFO ] Working on text chunk 1/6
# Processing chunk: 17%|█▋ | 1/6 [00:54<04:31, 54.24s/chunk][08-06-2026 22:12:47] [LLMlight.LLM] [INFO ] Working on text chunk 2/6
# Processing chunk: 33%|███▎ | 2/6 [01:18<02:26, 36.66s/chunk][08-06-2026 22:13:12] [LLMlight.LLM] [INFO ] Working on text chunk 3/6
# Processing chunk: 50%|█████ | 3/6 [01:30<01:15, 25.33s/chunk][08-06-2026 22:13:23] [LLMlight.LLM] [INFO ] Working on text chunk 4/6
# Processing chunk: 67%|██████▋ | 4/6 [01:51<00:47, 23.81s/chunk][08-06-2026 22:13:45] [LLMlight.LLM] [INFO ] Working on text chunk 5/6
# Processing chunk: 83%|████████▎ | 5/6 [02:47<00:35, 35.13s/chunk][08-06-2026 22:14:40] [LLMlight.LLM] [INFO ] Working on text chunk 6/6
# Processing chunk: 100%|██████████| 6/6 [03:14<00:00, 32.48s/chunk]
# [LLMlight.LLM] [INFO ] Running model: google/gemma-4-26b-a4b-qat
print(response)
# The provided context does not contain a formal definition of "attention networks."
# It only describes the mathematical mechanism of attention, which uses queries, keys, and values to
# compute output weights through methods like scaled dot-product or additive attention.
from LLMlight import LLMlight
# Initialize with GLOBAL-REASONING context strategy
client = LLMlight(model='google/gemma-4-26b-a4b-qat', retrieval_method='naive_rag', context_strategy='global-reasoning', top_chunks=6)
# Create (or load) database
client.memory_init(store_path='knowledge_base.db')
# Query on the new knowledge
response = client.prompt(query='What are attention networks?', response_format='Summarize in 2 sentence')
# [LLMlight.LLM] [INFO ] Global-reasoning on 6 chunks of text.
# Processing chunk: 100%|██████████| 6/6 [03:14<00:00, 32.48s/chunk]
# [LLMlight.LLM] [INFO ] Running model: google/gemma-4-26b-a4b-qat
print(response)
# Based on the provided text, attention networks are computational mechanisms that use queries, keys,
# and values to determine weights via a scaled dot-product and a softmax function.
# These networks can enhance model interpretability and, when combined with feed-forward layers, achieve a computational
# complexity similar to separable convolutions.
Exercise 4: Create A Discussion Between Two Agents On A Theme of Interest.
For the Exercise 4 Create A stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
# Initialize
from LLMlight import LLMlight
# temperature=0.1 Lower means a More factual debate
# temperature=1.0 Higher means more creative discussion
# top_chunks=10 Use more retrieved context
# embedding='bert' Strong semantic retrieval
# retrieval_method='naive_rag' Standard RAG retrieval
# ====================================================
# Agent A: Data Scientist
# ====================================================
agent_a = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
embedding="bert",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
# Set the database for Agent 1
agent_a.memory_init(store_path="agent_a.db")
# Add some background information database for agent 1
agent_a.memory_add("""
Large Language Models are one of the most important step we did in the field of AI
It helps the workload and the work easier and faster.
""")
agent_a.memory_add("""
Large Language Models use transformer architectures and are trained on
massive text corpora using self-supervised learning.
""")
# ====================================================
# Agent B: Farmer
# ====================================================
agent_b = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
embedding="bert",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
# Set the database for Agent 2
agent_b.memory_init(store_path="agent_b.db")
# Add some background information database for agent 2
agent_b.memory_add("""
The use of AI and machine learning consumes to much power and there is no need for this
new technology. The human work was good enough and there is no need to change that.
""")
agent_b.memory_add("""
Recent research shows that LLMs hallucinate and do not solve real world applications.
""")
# ====================================================
# Discussion Loop
# ====================================================
topic = "Discuss the importance of the use of Large Language Models and AI."
message = topic
for turn in range(5):
print(f"\n{'='*80}")
print(f"ROUND {turn+1}")
print(f"{'='*80}")
response_a = agent_a.prompt(
system='You are a Data Scientist.',
query=
f"""
Topic:
{message}
""",
response_format='Give your opinion in 1-2 paragraphs and ask a question to the other agent.',
)
print("\nAgent A:")
print(response_a)
response_b = agent_b.prompt(
system='You are a farmer.',
f"""
The Data Scientist said:
{response_a}
""",
response_format='Respond to the discussion in 1-2 paragraphs and ask a follow-up question.',
)
print("\nAgent B:")
print(response_b)
message = response_b
Introducing The Third Agent: The Moderator
For the Introducing The Third Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
from LLMlight import LLMlight
# ====================================================
# Agent A: Data Scientist
# ====================================================
agent_a = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
agent_a.memory_init(store_path="agent_a.db")
agent_a.memory_add("""
Large Language Models are one of the most important step we did in the field of AI
It helps the workload and the work easier and faster.
Large Language Models use transformer architectures and are trained on
massive text corpora using self-supervised learning.
""")
# ====================================================
# Agent B: Farmer
# ====================================================
agent_b = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
agent_b.memory_init(store_path="agent_b.db")
agent_b.memory_add("""
The use of AI and machine learning consumes to much power and there is no need for this
new technology. The human work was good enough and there is no need to change that.
Recent research shows that LLMs hallucinate and do not solve real world applications.
""")
# ====================================================
# Agent C: Moderator
# ====================================================
moderator = LLMlight(
model="openai/gpt-oss-20b",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.3, # lower temperature for objective summaries
)
moderator.memory_init(store_path="moderator.db")
# ====================================================
# Shared discussion memory
# ====================================================
shared_memory = LLMlight(model="openai/gpt-oss-20b")
shared_memory.memory_init(store_path="discussion.db")
# ====================================================
# Discussion Loop
# ====================================================
topic = "Discuss the importance of the use of Large Language Models and AI."
message = topic
for turn in range(5):
print(f"\n{'='*80}")
print(f"ROUND {turn+1}")
print(f"{'='*80}")
# --------------------------------------------
# Agent A responds
# --------------------------------------------
response_a = agent_a.prompt(
system='You are a Data Scientist.',
query=
f"""
Current discussion:
{message}
""",
instructions='Provide your opinion and ask a question to the Farmer.',
response_format='Response can be maximum 1-2 paragraphs.'
)
print("\nData Scientist:")
print(response_a)
# --------------------------------------------
# Agent B responds
# --------------------------------------------
response_b = agent_b.prompt(
system='You are a Farmer.',
query=f"""
The Data Scientist said:
{response_a}
"""
instructions='Respond and ask a follow-up question.',
response_format='Response can be maximum 1-2 paragraphs.'
)
print("\nFarmer:")
print(response_b)
# --------------------------------------------
# Moderator summarizes
# --------------------------------------------
moderator_summary = moderator.prompt(
system='You are a neutral moderator.',
query=
f"""
Data Scientist:
{response_a}
Farmer:
{response_b}
Perform the following tasks:
1. Summarize the key arguments.
2. Identify agreements.
3. Identify disagreements.
4. Propose one question that helps both agents move toward consensus.
""",
response_format='Keep the output concise.'
)
print("\nModerator:")
print(moderator_summary)
# Store discussion history
shared_memory.memory_add(response_a)
shared_memory.memory_add(response_b)
shared_memory.memory_add(moderator_summary)
# Next round starts from moderator guidance
message = moderator_summary
# ====================================================
# Final consensus
# ====================================================
consensus = moderator.prompt(
system='You are a neutral moderator.',
query=
"""
Review the discussion and provide:
- Main conclusions
- Remaining disagreements
- Final consensus statement
""",
response_format='Keep it under 200 words.'
)
print("\nFINAL CONSENSUS")
print("=" * 80)
print(consensus)
Controlling the Conversation With A Scoring Agent.
For the Controlling the Conversation With stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Controlling the Conversation With stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
from LLMlight import LLMlight
# ====================================================
# Agent A: Data Scientist
# ====================================================
agent_a = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
agent_a.memory_init(store_path="agent_a.db", overwrite=True)
# ====================================================
# Agent B: Farmer
# ====================================================
agent_b = LLMlight(
model="google/gemma-4-26b-a4b-qat",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.7,
)
agent_b.memory_init(store_path="agent_b.db", overwrite=True)
# ====================================================
# Moderator Agent (keeps discussion structured)
# ====================================================
moderator = LLMlight(
model="openai/gpt-oss-20b",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.3,
)
moderator.memory_init(store_path="moderator.db", overwrite=True)
# ====================================================
# Scoring Agent (decides convergence / stopping)
# ====================================================
scoring_agent = LLMlight(
model="liquid/lfm2-24b-a2b",
retrieval_method="naive_rag",
context_strategy=None,
top_chunks=5,
temperature=0.0, # deterministic scoring
)
scoring_agent.memory_init(store_path="scoring.db", overwrite=True)
# ====================================================
# Shared memory (optional logging)
# ====================================================
shared_memory = LLMlight(model="liquid/lfm2-24b-a2b")
shared_memory.memory_init(store_path="discussion.db", overwrite=True)
# ====================================================
# Discussion Loop with early stopping
# ====================================================
topic = "Discuss the importance of attention mechanisms in modern AI."
message = topic
MAX_ROUNDS = 5
AGREEMENT_THRESHOLD = 0.85 # stop if convergence is high enough
for turn in range(MAX_ROUNDS):
print(f"\n{'='*80}")
print(f"ROUND {turn+1}")
print(f"{'='*80}")
# --------------------------
# Agent A
# --------------------------
response_a = agent_a.prompt(
system='You are a Data Scientist.',
query=f"""
Topic:
{message}
""",
instructions='Ask a question.',
response_format='Respond in 1-2 paragraphs',
)
print("\nAgent A:")
print(response_a)
# --------------------------
# Agent B
# --------------------------
response_b = agent_b.prompt(
system='You are a Farmer.',
query=
f"""
Data Scientist said:
{response_a}
""",
instructions='continue the discussion with your own opinion.',
response_format='Respond in 1-2 paragraphs.',
)
print("\nAgent B:")
print(response_b)
# --------------------------
# Moderator summary
# --------------------------
moderator_summary = moderator.prompt(
system='You are a neutral moderator.',
query=f"""
Data Scientist:
{response_a}
Farmer:
{response_b}
""",
instructions=
"""
Summarize:
- agreements
- disagreements
- next question toward consensus
"""
)
print("\nModerator:")
print(moderator_summary)
# --------------------------
# Scoring Agent (convergence check)
# --------------------------
score_output = scoring_agent.prompt(
system='You are a scoring system.',
query=f"""
Data Scientist:
{response_a}
Farmer:
{response_b}
Moderator summary:
{moderator_summary}
""",
instructions='Evaluate agreement between the two agents.',
response_format=
"""
Return ONLY a number between 0 and 1:
- 1.0 = full agreement / consensus reached
- 0.0 = complete disagreement
""",
)
try:
score = float(score_output.strip())
except:
score = 0.0
print("\nAgreement Score:", score)
# --------------------------
# Store memory
# --------------------------
shared_memory.memory_add(response_a)
shared_memory.memory_add(response_b)
shared_memory.memory_add(moderator_summary)
# --------------------------
# Early stopping condition
# --------------------------
if score >= AGREEMENT_THRESHOLD:
print("\nConsensus reached early. Stopping discussion.")
break
# Next round context
message = moderator_summary
# ====================================================
# Final summary
# ====================================================
final_summary = shared_memory.prompt("""
Summarize the full discussion:
- final consensus
- key arguments
- remaining open points (if any)
""")
print("\nFINAL SUMMARY")
print("=" * 80)
print(final_summary)
# ========================
# ROUND 1
# ========================
# Agreement Score: 0.3
# ========================
# ROUND 2
# ========================
# Agreement Score: 0.7
# ========================
# ROUND 3
# ========================
# Agreement Score: 0.6
# ========================
# ROUND 4
# ========================
# Agreement Score: 0.8
# ========================
# ...
Good Instructions Are Key.
When working through the Good Instructions Are Key stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
response = client.prompt(
query="Explain attention mechanisms.",
instructions="""
Explain the concept for beginners.
Use exactly three paragraphs.
Include one real-world example.
""",
system="You are an experienced AI professor.",
context="some context", # This is autofilled too based on the database and RAG model.
response_format="markdown"
)
Takeaways Before Creating Language Models
When working through the Takeaways Before Creating Language stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Cache stable system instructions and tool schemas. Re-sending identical preamble is a common source of burn.
Wrapping Up: Don’t Go Fast, Go Structured
When working through the Wrapping Up Don t stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Wrapping Up Don t stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Software
The Software stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
References
The References stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for 24c6cd6fa849: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.
The hardening note 0 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 0/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 1 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 1/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 2 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 2/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 3 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 3/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 4 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 4/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 5 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 5/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 6 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 6/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 7 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 7/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 8 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 8/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 9 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Hardening detail 9/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 10 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments.
Hardening detail 10/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
When working through the hardening note 11 stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Hardening detail 11/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
The hardening note 12 stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion.
Hardening detail 12/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.
For the hardening note 13 stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph.
Hardening detail 13/872: measure wall time, error class, and token spend for this note, then decide whether to keep the change based on a fixed question set rather than anecdote.