This article is published in English.
RDF to GraphRAG: Practical Ontology Layers with VOO Examples
From IRIs and triples through RDFS, OWL, and SHACL—when ontologies beat tables and how they steady GraphRAG.
An ontology sounds mystical; the idea is plain. It is an explicit agreement about a domain: which kinds of things exist, which relationships are legal, which rules those links must obey, and what a machine may infer from facts already on hand. Gruber’s classic line—“an explicit specification of a conceptualization”—just means a machine-readable description of how a team chose to understand one slice of the world.
A running example helps: the Vanguard S&P 500 ETF (VOO). Public materials say VOO seeks to track the S&P 500. A knowledge system should understand at least:
Vanguard S&P essentially 500 ETF is an ETF. It is managed by Vanguard. It tracks a S&P 500 Index.
Those three sentences already touch every major layer below.
Ontology and knowledge graph are not the same thing
The ontology defines the rules of the world. The knowledge graph holds facts about that world.
The ontology might declare ETF and AssetManager as types, managedBy as a link from ETF to manager, and tracksIndex as a link from ETF to index. The graph then stores instances: VOO is an ETF; VOO managedBy Vanguard; VOO tracksIndex S&P500Index.
Think board-game design versus the current position on the board. Saying “ontology ≈ schema, knowledge graph ≈ data” is a useful shortcut—even though ontologies can encode richer logic than typical database schemas.
Do you always need an ontology?
No. Exact-key lookups over a few fields can live in tables or JSON. Ontology work costs design, governance, validation, and upkeep. It pays when problems like these appear:
Problem 1. Different systems use different words
fund provider, asset manager, and management company may mean one concept. An ontology picks a preferred term and maps aliases.
Problem 2. Users ask questions that require relationships
“Which ETFs managed by Vanguard track a U.S. equity index?” needs a multi-hop path, not a keyword hit.
Problem 3. The system needs to detect invalid data
If managedBy must point at an organisation, VOO managedBy John Smith should fail—even if John is a portfolio manager. In regulated industries a single bad edge can undermine compliance and AI answers; ontology constraints act as intake guardrails.
Problem 4. The system should derive facts that were never directly stored
Subclass and inverse-property reasoning can materialise implied types and reverse links.
Problem 5. An LLM needs a reliable map of the domain
Models invent fluent structure; an ontology supplies a checked map for decomposition, grounding, and validation—especially beside GraphRAG.
One example, five technology layers
The VOO story climbs five layers: identifiers, RDF triples, RDFS schema, OWL semantics, and SHACL validation.
Layer 1: identifiers and namespaces
Stable IRIs avoid name collisions. A manager might be named:
https://example.org/finance/Vanguard
Prefixes keep files readable:
@prefix fin: <https://example.org/finance/> .
Which expands local names such as:
fin:Vanguard
Layer 2: RDF represents facts as triples
Every statement is subject–predicate–object:
Subject Predicate Object
VOO managedBy Vanguard
VOO tracksIndex S&P 500 Index
Turtle form:
@prefix fin: <https://example.org/finance/> .
fin:VOO fin:managedBy fin:Vanguard ;
fin:tracksIndex fin:SP500Index .
JSON-LD can carry the same graph for web stacks:
{
"@context": {
"fin": "https://example.org/finance/",
"managedBy": {
"@id": "fin:managedBy",
"@type": "@id"
},
"tracksIndex": {
"@id": "fin:tracksIndex",
"@type": "@id"
}
},
"@id": "fin:VOO",
"managedBy": "fin:Vanguard",
"tracksIndex": "fin:SP500Index"
}
Layer 3: RDFS introduces basic schema
RDFS adds classes, properties, domain/range, and subclass links. A simple product taxonomy:
@prefix fin: <https://example.org/finance/> .
@prefix rdf: <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .
fin:FinancialProduct a rdfs:Class .
fin:Fund a rdfs:Class ;
rdfs:subClassOf fin:FinancialProduct .
fin:ETF a rdfs:Class ;
rdfs:subClassOf fin:Fund .
fin:AssetManager a rdfs:Class .
fin:MarketIndex a rdfs:Class .
fin:managedBy a rdf:Property ;
rdfs:domain fin:Fund ;
rdfs:range fin:AssetManager .
fin:tracksIndex a rdf:Property ;
rdfs:domain fin:ETF ;
rdfs:range fin:MarketIndex .
Declaring an ETF under that tree:
FinancialProduct
└── Fund
└── ETF
Layer 4: OWL adds richer semantics
OWL can state inverses, cardinality, and disjointness. If managedBy is the inverse of manages, storing one direction can imply the other:
fin:VOO fin:managedBy fin:Vanguard .
Inverse sketch:
@prefix fin: <https://example.org/finance/> .
@prefix owl: <http://www.w3.org/2002/07/owl#> .
fin:managedBy a owl:ObjectProperty ;
owl:inverseOf fin:manages .
fin:ETF owl:disjointWith fin:AssetManager .
Human reading of the implication:
If VOO is managed by Vanguard,
then Vanguard manages VOO.
Asserted fact:
fin:VOO fin:managedBy fin:Vanguard .
Inferred reverse:
fin:Vanguard fin:manages fin:VOO .
Layer 5: SHACL validates graph data
Shapes catch instance errors ontology authors care about in production. A valid ETF description:
@prefix fin: <https://example.org/finance/> .
@prefix sh: <http://www.w3.org/ns/shacl#> .
fin:ETFShape
a sh:NodeShape ;
sh:targetClass fin:ETF ;
sh:property [
sh:path fin:managedBy ;
sh:class fin:AssetManager ;
sh:minCount 1 ;
sh:maxCount 1
] ;
sh:property [
sh:path fin:tracksIndex ;
sh:class fin:MarketIndex ;
sh:minCount 1
] .
Compact valid instance:
fin:VOO a fin:ETF ;
fin:managedBy fin:Vanguard ;
fin:tracksIndex fin:SP500Index .
fin:Vanguard a fin:AssetManager .
fin:SP500Index a fin:MarketIndex .
An invalid manager type should fail validation:
fin:BrokenFund a fin:ETF ;
fin:managedBy fin:Alice .
fin:Alice a fin:PortfolioManager .
How inference creates new knowledge
Subclass chains promote instances up the tree:
fin:ETF rdfs:subClassOf fin:Fund .
fin:Fund rdfs:subClassOf fin:FinancialProduct .
fin:managedBy rdfs:range fin:AssetManager .
fin:managedBy owl:inverseOf fin:manages .
From a narrow type assertion:
fin:VOO a fin:ETF ;
fin:managedBy fin:Vanguard .
Reasoners can conclude broader types:
fin:VOO a fin:Fund .
fin:VOO a fin:FinancialProduct .
fin:Vanguard a fin:AssetManager .
fin:Vanguard fin:manages fin:VOO .
Inference fills gaps; SHACL still guards what humans or extractors write.
Competency questions: design from the questions backward
Start from questions the system must answer—“Which Vanguard ETFs track U.S. equity indexes?”—then decide types, properties, and constraints. That keeps ontologies tied to use, not philosophical completeness.
A practical ontology-development workflow
Domain goals + user questions + source data
↓
Competency-question generation
↓
Concept and relationship extraction
↓
Initial ontology proposal
↓
Reasoning, SHACL, and query evaluation
↓
Human review and iterative revision
Iterate: goals and questions → conceptual model → formal ontology → populate graph → validate → query/reason → revise.
Conceptual sketch:
ETF ──managedBy──> AssetManager
ETF ──tracksIndex──> MarketIndex
Example SPARQL-shaped ask:
SELECT ?etf
WHERE {
?etf a fin:ETF ;
fin:managedBy fin:Vanguard ;
fin:tracksIndex fin:SP500Index .
}
Layered stack reminder:
RDF stores:
VOO managedBy Vanguard
RDFS understands:
VOO is a Fund, and Vanguard is an AssetManager
OWL can infer:
Vanguard manages VOO
SHACL checks:
Does VOO have exactly one valid AssetManager?
Does it track at least one MarketIndex?
What LLMs can automate — and what they should not decide alone
Models help draft labels, suggest properties, and propose competency questions. They should not silently own governance: final type systems, cardinality, and regulatory constraints need human owners and tests.
Where ontology helps GraphRAG
1. Query decomposition
Typed relations tell planners which hops are meaningful.
2. Entity resolution
Shared IRIs and synonym maps collapse “Vanguard” vs “The Vanguard Group”.
3. Retrieval control
Anchors and edge filters beat pure cosine when identifiers matter.
4. Answer validation
SHACL and shape-derived checks reject answers that invent illegal edges.
Five design principles worth keeping
- Separate schema (ontology) from instance data (graph).
- Design from competency questions, not from buzzwords.
- Prefer small, testable constraints over monolith axioms.
- Validate on write; reason where it earns its cost.
- Treat LLM help as drafting assistance under human governance.
Conclusion
Ontologies are agreements made machine-checkable. RDF stores facts; RDFS and OWL add structure and inference; SHACL enforces instance quality; GraphRAG consumes the result as a safer map than vectors alone. The VOO example is small on purpose: if three sentences can exercise five layers, a real product catalog can too—once competency questions and owners exist.
Keep a one-page decision log for each major class and property: why it exists, which competency question it serves, and which SHACL shape guards it. Version IRI policies the way APIs version routes; renaming without redirects breaks every downstream GraphRAG join. When extractors propose new edges, require a shape or an explicit “unconstrained” sandbox graph so noisy LLM output cannot poison the governed store. Measure ontology value with question coverage and validation catch-rate, not with axiom count. Finally, pair every production GraphRAG deployment with a fixtures pack of legal and illegal triples so CI fails when a “helpful” schema edit silently widens domain/range and lets bad managers attach to funds again.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.
Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.