This article is published in English.
Practical notes: How I turned my photo gallery into an autonomous AI Agent
Operable walkthrough of Practical notes: How I turned my photo gallery into an autonomous AI Agent: contracts, checks, and drop-in code slots for teams shipping this pattern.
This walkthrough rebuilds the path from raw materials to a working system for: How I turned my photo gallery into an autonomous AI Agent — The Complete Guide. The focus is operable steps, explicit checks, and code that you can drop into a repo without guessing intent. For the Overview stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Introduction
When working through the Introduction stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Why Traditional Photo Search Is Broken
When working through the Why Traditional Photo Search stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
The album approach collapses in real life
When working through the The album approach collapses stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the The album approach collapses stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keyword search only gets you so far
The Keyword search only gets stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Cloud gallery search: good, but at a cost
The Cloud gallery search good stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The real gap: no semantic understanding
The The real gap no stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The The real gap no stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The Concept: Semantic Search, Explained Simply
For the The Concept Semantic Search stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Where does language come in?
For the Where does language come stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Choosing the Right Tools (And Why It Took Weeks)
For the Choosing the Right Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Authenticate at the gateway and re-authorize at the data plane. A bearer token alone is not a tenancy boundary. For the Choosing the Right Tools stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
CLIP ViT-B/32 via FastEmbed — the eyes of the system
When working through the CLIP ViT-B 32 via stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface.
from fastembed import ImageEmbeddingModel, TextEmbeddingModel
# Loaded once, reused forever
image_model = ImageEmbeddingModel.from_pretrained("Qdrant/clip-ViT-B-32-vision")
text_model = TextEmbeddingModel.from_pretrained("Qdrant/clip-ViT-B-32-text")
# After embedding:
image_vector = embed_image("sunset_photo.jpg") # shape: (512,)
query_vector = embed_text("beautiful sunset") # shape: (512,)
# Cosine similarity — just a dot product on normalised vectors
similarity = np.dot(image_vector, query_vector)
# If image is a sunset → similarity ≈ 0.85 (strong match)
# If image is food → similarity ≈ 0.15 (weak match)
Qdrant Edge — the memory
When working through the Qdrant Edge the memory stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from qdrant_edge import(
Distance,
EdgeConfig,
EdgeShard,
EdgeVectorParams,
)
config = EdgeConfig(
vectors={
VECTOR_NAME: EdgeVectorParams(
size=EMBED_DIM,
distance=Distance.Cosine
)
}
)
_shard = EdgeShard.create(path=str(SHARD_DIR), config=config)
Putting it all together
When working through the Putting it all together stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Putting it all together stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
OpenClaw — the brain
The OpenClaw the brain stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Setting Up the Environment
The Setting Up the Environment stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Step 1: Clone the repo and set up a virtual environment
The Step 1 Clone the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Step 1 Clone the stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
# Clone the repo
git clone https://github.com/vatsala-singh/AI-Powered-Photo-Search-and-Tagging-Agent.git
cd AI-Powered-Photo-Search-and-Tagging-Agent
# Create and activate virtual environment
python -m venv .venv
source .venv/bin/activate # Mac/Linux
# .venv\Scripts\activate # Windows
Step 2: Install the dependencies
For the Step 2 Install the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
pip install -r requirements.txt
Step 3: Understand the project structure
For the Step 3 Understand the stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
AI-Powered-Photo-Search-and-Tagging-Agent/
│
├── main.py # FastAPI app + OpenClaw agent entry point
├── config.py # All configurable parameters in one place
├── requirements.txt
│
├── pipeline/
│ ├── embedder.py # CLIP embedding logic (image + text)
│ └── indexer.py # Batch photo processing and indexing
│
├── store/
│ └── qdrant_client.py # Qdrant Edge setup and collection management
│
├── tools/
│ ├── search.py # Semantic search tool
│ ├── tag.py # Zero-shot auto-tagging tool
│ ├── duplicates.py # Near-duplicate detection tool
│ └── albums.py # Smart album grouping tool
│
├── test/
│ ├── embedder_test.py
│ ├── indexer_test.py
│ ├── search_test.py
│ └── qdrant_edge_client_test.py
│
└── qdrant-edge-data/ # Auto-created at runtime
├── storage/ # Qdrant's internal shard data
├── models/ # Cached CLIP model weights
└── photos/ # Collection data
Step 4: Glance at config.py
For the Step 4 Glance at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Step 4 Glance at stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
# config.py
CLIP_IMAGE_MODEL = "Qdrant/clip-ViT-B-32-vision"
CLIP_TEXT_MODEL = "Qdrant/clip-ViT-B-32-text"
EMBEDDING_DIM = 512 # CLIP ViT-B/32 output dimension
COLLECTION_NAME = "photos"
QDRANT_PATH = "./qdrant-edge-data"
BATCH_SIZE = 32 # Images per indexing batch
TOP_K = 10 # Default search results returned
TAG_THRESHOLD = 0.20 # Min similarity score for a tag to apply
DUPLICATE_THRESHOLD = 0.97 # Min similarity to flag as duplicate
Step 5: Start the server
When working through the Step 5 Start the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
uvicorn main:app --reload --port 8000
Step 6: Verify the setup
When working through the Step 6 Verify the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
# Quick sanity check - should return {"status": "ok"}
curl http://localhost:8000/health
curl http://localhost:8000/api/status
Building the Image Embedding Pipeline
When working through the Building the Image Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Measure recall on a fixed question set before tuning prompts. Prompt churn rarely fixes a weak retrieval surface. When working through the Building the Image Embedding stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The embedder
The The embedder stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
# pipeline/embedder.py
from fastembed import ImageEmbeddingModel, TextEmbeddingModel
from PIL import Image
import numpy as np
from config import CLIP_IMAGE_MODEL, CLIP_TEXT_MODEL
# Load once, reuse for the lifetime of the process
# Models are large (~150MB each) — we never want to reload them per request
_image_model = ImageEmbeddingModel.from_pretrained(CLIP_IMAGE_MODEL)
_text_model = TextEmbeddingModel.from_pretrained(CLIP_TEXT_MODEL)
def embed_image(image_path: str) -> np.ndarray:
"""Convert an image file to a 512-d CLIP embedding."""
image = Image.open(image_path).convert("RGB")
embeddings = list(_image_model.embed([image]))
return np.array(embeddings[0]) # shape: (512,)
def embed_text(query: str) -> np.ndarray:
"""Convert a text string to a 512-d CLIP embedding."""
embeddings = list(_text_model.embed([query]))
return np.array(embeddings[0]) # shape: (512,)
The indexer
The The indexer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
# pipeline/indexer.py
import os
import uuid
from pathlib import Path
from datetime import datetime
from PIL import Image
from pipeline.embedder import embed_image
from store.qdrant_client import get_shard
from tools.tag import generate_tags
from config import BATCH_SIZE
SUPPORTED_FORMATS = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"}
def index_folder(folder_path: str) -> dict:
"""
Recursively index all images in a folder into Qdrant Edge.
Returns a summary: total found, indexed, skipped.
"""
folder = Path(folder_path)
shard = get_shard()
image_paths = [
p for p in folder.rglob("*")
if p.suffix.lower() in SUPPORTED_FORMATS
]
total = len(image_paths)
indexed = 0
skipped = 0
batch = []
for i, path in enumerate(image_paths):
try:
vector = embed_image(str(path))
tags = generate_tags(str(path))
img = Image.open(path)
point = {
"id": str(uuid.uuid4()),
"vector": vector.tolist(),
"payload": {
"filename": path.name,
"filepath": str(path.absolute()),
"tags": tags,
"timestamp": int(path.stat().st_mtime),
"width": img.width,
"height": img.height,
}
}
batch.append(point)
indexed += 1
except Exception as e:
print(f"Skipping {path.name}: {e}")
skipped += 1
# Flush every BATCH_SIZE images
if len(batch) >= BATCH_SIZE:
shard.upsert(points=batch)
batch = []
print(f" Progress: {i+1}/{total} images indexed...")
# Flush remaining
if batch:
shard.upsert(points=batch)
return {"total": total, "indexed": indexed, "skipped": skipped}
Indexing Images into Qdrant Edge
The Indexing Images into Qdrant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Indexing Images into Qdrant stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
How Qdrant Edge works here
For the How Qdrant Edge works stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Setting up the shard
For the Setting up the shard stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
def get_shard() -> EdgeShard:
"""
Return the singleton EdgeShard, creating it on first call.
- If SHARD_DIR does not exist → create a brand-new shard.
- If SHARD_DIR already contains data → reopen it (no config needed).
EdgeShard runs entirely in-process. No binary, no port, no network.
"""
global _shard
if _shard is not None:
return _shard
SHARD_DIR.mkdir(parents=True, exist_ok=True)
# Detect whether this is a fresh shard or an existing one.
# EdgeShard.create() fails if data already exists on disk.
shard_has_data = any(SHARD_DIR.iterdir())
if shard_has_data:
print(f"[qdrant_client] Reopening existing shard at '{SHARD_DIR}'")
_shard = EdgeShard.load(path=SHARD_DIR)
else:
print(f"[qdrant_client] Creating new shard at '{SHARD_DIR}'")
config = EdgeConfig(
vectors={
VECTOR_NAME: EdgeVectorParams(
size=EMBED_DIM,
distance=Distance.Cosine
)
}
)
_shard = EdgeShard.create(path=str(SHARD_DIR), config=config)
print(f"[store] Shard ready — vector: '{VECTOR_NAME}', dim: {EMBED_DIM}")
return _shard
The create vs load split
For the The create vs load stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the The create vs load stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
The singleton pattern
When working through the The singleton pattern stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
@app.on_event("shutdown")
def on_shutdown():
close_shard()
The payload schema
When working through the The payload schema stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from dataclasses import dataclass, field
from typing import List, Optional
@dataclass
class PhotoPayload:
filename:str
path:str
tags: list[str] = field(default_factory=list)
timestamp: Optional[str] = None
width: Optional[int] = None
height: Optional[int] = None
def to_dict(self) -> dict:
return {
"filename": self.filename,
"path": self.path,
"tags": self.tags,
"timestamp": self.timestamp,
"width": self.width,
"height": self.height
}
Writing points to the shard
When working through the Writing points to the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
from qdrant_edge import PointStruct
from store.qdrant_client import get_shard
from schema import PhotoPayload
shard = get_shard()
payload = PhotoPayload(
filename = "beach_sunset.jpg",
path = "/Users/me/Pictures/2024/Goa/beach_sunset.jpg",
tags = ["sunset", "beach", "outdoor"],
timestamp = "2024-05-01T18:42:00",
width = 4032,
height = 3024
)
point = PointStruct(
id = "3f7a2b1c-8e4d-4f9a-b2c1-7d8e9f0a1b2c",
vector = {"image": vector.tolist()}, # named vector matching VECTOR_NAME
payload = payload.to_dict()
)
shard.upsert(points=[point])
When working through the Writing points to the stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Automatic Photo Tagging
The Automatic Photo Tagging stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
The idea: zero-shot classification
The The idea zero-shot classification stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
TAG_VOCABULARY = [
"sunset", "sunrise", "beach", "ocean", "mountain", "forest", "city",
"night", "snow", "rain", "fog", "sunny", "cloudy",
"dog", "cat", "bird", "people", "crowd", "portrait", "selfie",
"food", "coffee", "restaurant", "travel", "architecture",
"car", "road", "nature", "flowers", "trees",
"indoor", "outdoor", "party", "celebration", "sport",
"screenshot", "document", "text", "map",
]
# tools/tag.py
from pipeline.embedder import embed_image, embed_text
from config import TAG_THRESHOLD
import numpy as np
TAG_LABELS = [...] # full list as above
# Pre-compute label embeddings once at module load —
# no point re-embedding the same 50 words on every photo
_label_vectors = {
label: embed_text(label)
for label in TAG_LABELS
}
def generate_tags(image_path: str) -> list[str]:
"""
Run zero-shot classification on an image.
Returns a list of tags whose similarity to the image
exceeds TAG_THRESHOLD (default: 0.20).
"""
image_vector = embed_image(image_path)
tags = []
for label, label_vector in _label_vectors.items():
similarity = np.dot(image_vector, label_vector) # cosine sim on normalised vectors
if similarity >= TAG_THRESHOLD:
tags.append(label)
return tags
Tags at index time vs query time
The Tags at index time stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Tags at index time stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
def generate_tags_from_vector(img_vec: np.ndarray, threshold: float = 0.20, max_tags: int = 6) -> list[str]:
"""
Generate tags for an image vector using zero-shot CLIP classification.
Tags with cosine similarity above threshold are included (up to max_tags).
This is a utility function used for generating tags during indexing
or when you already have an image vector.
"""
tag_vecs = _get_tag_vectors()
scores = {
tag: float(np.dot(img_vec, vec)) # both normalized → cosine similarity
for tag, vec in tag_vecs.items()
}
tags = sorted(
[t for t, s in scores.items() if s >= threshold],
key=lambda t: scores[t],
reverse=True,
)[:max_tags]
return tags
Using tags as filters
For the Using tags as filters stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
def search_photos(query: str, top_k: int = TOP_K, tags: list[str] = None) -> list[dict]:
#search photo library with a Natural language query
#takes in query, no of results to be displayed, and a list of tags
# returns list of dicts with photo metadata and relevance score
print(f"[search] Received query='{query}' with tags={tags} and top_k={top_k}")
shard = get_shard()
query_vector = embed_text(query)
# Over-fetch when tag filtering is requested to have enough candidates
# after post-filtering by tags
over_fetch_multiplier = 5 if tags else 1
fetch_limit = top_k * over_fetch_multiplier
results = shard.query(
QueryRequest(
query=Query.Nearest(query_vector.tolist(), using=VECTOR_NAME),
limit=fetch_limit,
with_vector=False,
with_payload=True,
)
)
print(f"[search] Found {len(results)} initial hits for query='{query}' with tags={tags}")
hits = []
untagged_hits = [] # Fallback results for images without tags
for point in results:
payload = point.payload or {}
point_tags = payload.get("tags", [])
result_dict = {
"path": payload.get("path"),
"filename": payload.get("filename"),
"tags": point_tags,
"timestamp": payload.get("timestamp"),
"score": round(point.score, 4)
}
# Post-filter by tags if specified
# (EdgeShard doesn't support complex filters, so we filter in Python
# after over-fetching more results than needed)
if tags:
if point_tags and any(t in point_tags for t in tags):
# Has tags and matches the filter
hits.append(result_dict)
elif not point_tags:
# No tags yet (images not auto-tagged), save as fallback
untagged_hits.append(result_dict)
else:
# No tag filter specified, include all results
hits.append(result_dict)
# Stop if we have enough tagged results
if len(hits) >= top_k:
break
# If we don't have enough tagged results, include untagged ones that match the query
if tags and len(hits) < top_k:
hits.extend(untagged_hits[:top_k - len(hits)])
return hits[:top_k]
Building the Search Agent
For the Building the Search Agent stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
curl --location 'http://localhost:8000/search' \
--header 'Content-Type: application/json' \
--data '{"query": "eiffel tower from rooftop","tags":[], "top_k": 1}'
{
"query": "eiffel tower from rooftop",
"results": [
{
"path": "/Users/vatsalasingh/Documents/Datasets/tag_phot/photo-1638051017225-0d9fcca18cf4.jpg",
"filename": "photo-1638051017225-0d9fcca18cf4.jpg",
"tags": [
"cloudy",
"city",
"rain",
"screenshot",
"architecture",
"travel"
],
"timestamp": "2021-12-09T17:27:56",
"score": 0.2864
}
]
}
Handling edge cases gracefully
For the Handling edge cases gracefully stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Handling edge cases gracefully stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Orchestrating it all with OpenClaw
When working through the Orchestrating it all with stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
How OpenClaw works
When working through the How OpenClaw works stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
@app.post("/chat")
def chat(req: ChatRequest):
"""
Conversational endpoint. Accepts user message and conversation history,
returns agent's reply after processing with tools.
"""
# Define tool functions that the agent can call
def search_tool(query: str, top_k: int = 10, tag_filter: list = None):
"""Search photos by natural language query"""
return search_photos(query=query, top_k=top_k, tags=tag_filter)
def duplicates_tool(threshold: float = 0.97):
"""Find duplicate or near-duplicate photos"""
return find_duplicates(threshold=threshold)
def tag_tool(image_path: str):
"""Generate and update tags for a specific photo"""
return generate_tags_from_vector(image_path=image_path)
# Create the agent with tools
agent = Agent(
tools=[search_tool, duplicates_tool, tag_tool],
system_prompt="""
You are a personal photo assistant. You help users search, organize,
and understand their local photo library. You have access to tools
for semantic search, duplicate detection, and tagging.
When helping users:
- Use the search tool to find photos by describing their content
- Use duplicates tool to find and clean up duplicate shots
- Use tag tool to inspect or update tags for specific photos
Always be concise and helpful. When returning photo results,
format them clearly with filenames, similarity scores, and tags.
Use emojis sparingly but helpfully.
"""
)
# Run the agent conversation
response = agent.chat(
message=req.message,
history=req.history
)
return {"response": response}
Real interaction flows
When working through the Real interaction flows stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Real interaction flows stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Why the agent layer matters
The Why the agent layer stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Running Everything Locally — And Why That Matters
The Running Everything Locally And stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Nothing leaves your device. Full stop.
The Nothing leaves your device stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The Nothing leaves your device stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
What it actually takes to run this
For the What it actually takes stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
What’s Next — Extending the System
For the What s Next Extending stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness.
Face clustering
For the Face clustering stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Put human approval on edges that spend money or change production data. Compile-time wiring does not equal business completeness. For the Face clustering stage, define the inputs, the owner of the step, and the exit criteria before changing code. Operators should be able to re-run the step from a known checkpoint without guessing hidden state. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
OCR for screenshots and documents
When working through the OCR for screenshots and stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Video frame indexing
When working through the Video frame indexing stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node.
Hybrid search: vector + keyword + metadata
When working through the Hybrid search vector keyword stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Checkpoint after expensive steps. Resume should not re-bill the same LLM call when an operator retries a later node. When working through the Hybrid search vector keyword stage, write down the contract first: required inputs, success signal, and what happens on partial failure. That checklist keeps later code changes honest. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Time and location aware retrieval
The Time and location aware stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Treat this stage as a contract between inputs and validated outputs. Name the artifacts, define success checks, and refuse silent partial completion. Separate chunking policy from retrieval policy. Changing one should not force a rewrite of the other when quality metrics move.
Final Thoughts
The Final Thoughts stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Record timings and token or query cost next to functional results. Cost visibility early prevents surprise bills when the path moves from demo to shared environments. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
References & Further Reading
The References Further Reading stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Keep configuration outside application code. Environment files, secret stores, and feature flags belong in one place operators can audit without reading the whole graph. Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts. The References Further Reading stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope. Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Operational checklist
The Operational checklist stage works best when treated as a measurable surface. Capture one golden transcript, one failure case, and the rollback note before expanding scope.
Document the happy path and the recovery path together. Retries, human gates, and dead-letter handling are part of the product, not later polish.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Add a smoke test that exercises the critical path in CI with fixtures, not live paid APIs, whenever budgets allow.
Prefer small, testable units over sprawling scripts. When a step fails, the failure should point at a single responsibility rather than a tangled pipeline.
Keep graph state flat and typed. Nested blobs hide which node wrote which field and break resume after interrupts.
Before promoting the stack, freeze versions, capture a golden transcript for the critical path, and confirm rollback steps. Shared environments need rate limits, tenancy checks, and a clear owner for secret rotation. Prefer boring reliability over clever one-off demos.
Batch note for b7b9768f8acd: keep provider keys out of the repo, set a per-session token ceiling, and store transcripts next to the eval fixtures so later model swaps stay comparable.