Home / Articles / RDF to GraphRAG: Practical Ontology Layers with VOO Examples

This article is published in English.

RDF to GraphRAG: Practical Ontology Layers with VOO Examples

From IRIs and triples through RDFS, OWL, and SHACL—when ontologies beat tables and how they steady GraphRAG.

2162 words

An ontology sounds mystical; the idea is plain. It is an explicit agreement about a domain: which kinds of things exist, which relationships are legal, which rules those links must obey, and what a machine may infer from facts already on hand. Gruber’s classic line—“an explicit specification of a conceptualization”—just means a machine-readable description of how a team chose to understand one slice of the world.

A running example helps: the Vanguard S&P 500 ETF (VOO). Public materials say VOO seeks to track the S&P 500. A knowledge system should understand at least:

Vanguard S&P essentially 500 ETF is an ETF. It is managed by Vanguard. It tracks a S&P 500 Index.

Those three sentences already touch every major layer below.

Ontology and knowledge graph are not the same thing

The ontology defines the rules of the world. The knowledge graph holds facts about that world.

The ontology might declare ETF and AssetManager as types, managedBy as a link from ETF to manager, and tracksIndex as a link from ETF to index. The graph then stores instances: VOO is an ETF; VOO managedBy Vanguard; VOO tracksIndex S&P500Index.

Think board-game design versus the current position on the board. Saying “ontology ≈ schema, knowledge graph ≈ data” is a useful shortcut—even though ontologies can encode richer logic than typical database schemas.

Do you always need an ontology?

No. Exact-key lookups over a few fields can live in tables or JSON. Ontology work costs design, governance, validation, and upkeep. It pays when problems like these appear:

Problem 1. Different systems use different words

fund provider, asset manager, and management company may mean one concept. An ontology picks a preferred term and maps aliases.

Problem 2. Users ask questions that require relationships

“Which ETFs managed by Vanguard track a U.S. equity index?” needs a multi-hop path, not a keyword hit.

Problem 3. The system needs to detect invalid data

If managedBy must point at an organisation, VOO managedBy John Smith should fail—even if John is a portfolio manager. In regulated industries a single bad edge can undermine compliance and AI answers; ontology constraints act as intake guardrails.

Problem 4. The system should derive facts that were never directly stored

Subclass and inverse-property reasoning can materialise implied types and reverse links.

Problem 5. An LLM needs a reliable map of the domain

Models invent fluent structure; an ontology supplies a checked map for decomposition, grounding, and validation—especially beside GraphRAG.

One example, five technology layers

The VOO story climbs five layers: identifiers, RDF triples, RDFS schema, OWL semantics, and SHACL validation.

Layer 1: identifiers and namespaces

Stable IRIs avoid name collisions. A manager might be named:

https://example.org/finance/Vanguard

Prefixes keep files readable:

@prefix fin: <https://example.org/finance/> .

Which expands local names such as:

fin:Vanguard

Layer 2: RDF represents facts as triples

Every statement is subject–predicate–object:

Subject       Predicate       Object
VOO           managedBy       Vanguard
VOO           tracksIndex     S&P 500 Index

Turtle form:

@prefix fin: <https://example.org/finance/> .

fin:VOO fin:managedBy fin:Vanguard ;
        fin:tracksIndex fin:SP500Index .

JSON-LD can carry the same graph for web stacks:

{
  "@context": {
    "fin": "https://example.org/finance/",
    "managedBy": {
      "@id": "fin:managedBy",
      "@type": "@id"
    },
    "tracksIndex": {
      "@id": "fin:tracksIndex",
      "@type": "@id"
    }
  },
  "@id": "fin:VOO",
  "managedBy": "fin:Vanguard",
  "tracksIndex": "fin:SP500Index"
}

Layer 3: RDFS introduces basic schema

RDFS adds classes, properties, domain/range, and subclass links. A simple product taxonomy:

@prefix fin:  <https://example.org/finance/> .
@prefix rdf:  <http://www.w3.org/1999/02/22-rdf-syntax-ns#> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .

fin:FinancialProduct a rdfs:Class .
fin:Fund             a rdfs:Class ;
                     rdfs:subClassOf fin:FinancialProduct .
fin:ETF              a rdfs:Class ;
                     rdfs:subClassOf fin:Fund .
fin:AssetManager     a rdfs:Class .
fin:MarketIndex      a rdfs:Class .
fin:managedBy a rdf:Property ;
    rdfs:domain fin:Fund ;
    rdfs:range fin:AssetManager .
fin:tracksIndex a rdf:Property ;
    rdfs:domain fin:ETF ;
    rdfs:range fin:MarketIndex .

Declaring an ETF under that tree:

FinancialProduct
└── Fund
    └── ETF

Layer 4: OWL adds richer semantics

OWL can state inverses, cardinality, and disjointness. If managedBy is the inverse of manages, storing one direction can imply the other:

fin:VOO fin:managedBy fin:Vanguard .

Inverse sketch:

@prefix fin: <https://example.org/finance/> .
@prefix owl: <http://www.w3.org/2002/07/owl#> .

fin:managedBy a owl:ObjectProperty ;
    owl:inverseOf fin:manages .
fin:ETF owl:disjointWith fin:AssetManager .

Human reading of the implication:

If VOO is managed by Vanguard,
then Vanguard manages VOO.

Asserted fact:

fin:VOO fin:managedBy fin:Vanguard .

Inferred reverse:

fin:Vanguard fin:manages fin:VOO .

Layer 5: SHACL validates graph data

Shapes catch instance errors ontology authors care about in production. A valid ETF description:

@prefix fin: <https://example.org/finance/> .
@prefix sh:  <http://www.w3.org/ns/shacl#> .

fin:ETFShape
    a sh:NodeShape ;
    sh:targetClass fin:ETF ;
    sh:property [
        sh:path fin:managedBy ;
        sh:class fin:AssetManager ;
        sh:minCount 1 ;
        sh:maxCount 1
    ] ;

    sh:property [
        sh:path fin:tracksIndex ;
        sh:class fin:MarketIndex ;
        sh:minCount 1
    ] .

Compact valid instance:

fin:VOO a fin:ETF ;
    fin:managedBy fin:Vanguard ;
    fin:tracksIndex fin:SP500Index .

fin:Vanguard a fin:AssetManager .
fin:SP500Index a fin:MarketIndex .

An invalid manager type should fail validation:

fin:BrokenFund a fin:ETF ;
    fin:managedBy fin:Alice .

fin:Alice a fin:PortfolioManager .

How inference creates new knowledge

Subclass chains promote instances up the tree:

fin:ETF rdfs:subClassOf fin:Fund .
fin:Fund rdfs:subClassOf fin:FinancialProduct .
fin:managedBy rdfs:range fin:AssetManager .
fin:managedBy owl:inverseOf fin:manages .

From a narrow type assertion:

fin:VOO a fin:ETF ;
    fin:managedBy fin:Vanguard .

Reasoners can conclude broader types:

fin:VOO a fin:Fund .
fin:VOO a fin:FinancialProduct .
fin:Vanguard a fin:AssetManager .
fin:Vanguard fin:manages fin:VOO .

Inference fills gaps; SHACL still guards what humans or extractors write.

Competency questions: design from the questions backward

Start from questions the system must answer—“Which Vanguard ETFs track U.S. equity indexes?”—then decide types, properties, and constraints. That keeps ontologies tied to use, not philosophical completeness.

A practical ontology-development workflow

Domain goals + user questions + source data
                  ↓
       Competency-question generation
                  ↓
       Concept and relationship extraction
                  ↓
         Initial ontology proposal
                  ↓
     Reasoning, SHACL, and query evaluation
                  ↓
       Human review and iterative revision

Iterate: goals and questions → conceptual model → formal ontology → populate graph → validate → query/reason → revise.

Conceptual sketch:

ETF ──managedBy──> AssetManager
ETF ──tracksIndex──> MarketIndex

Example SPARQL-shaped ask:

SELECT ?etf
WHERE {
  ?etf a fin:ETF ;
       fin:managedBy fin:Vanguard ;
       fin:tracksIndex fin:SP500Index .
}

Layered stack reminder:

RDF stores:
VOO managedBy Vanguard

RDFS understands:
VOO is a Fund, and Vanguard is an AssetManager

OWL can infer:
Vanguard manages VOO

SHACL checks:
Does VOO have exactly one valid AssetManager?
Does it track at least one MarketIndex?

What LLMs can automate — and what they should not decide alone

Models help draft labels, suggest properties, and propose competency questions. They should not silently own governance: final type systems, cardinality, and regulatory constraints need human owners and tests.

Where ontology helps GraphRAG

1. Query decomposition

Typed relations tell planners which hops are meaningful.

2. Entity resolution

Shared IRIs and synonym maps collapse “Vanguard” vs “The Vanguard Group”.

3. Retrieval control

Anchors and edge filters beat pure cosine when identifiers matter.

4. Answer validation

SHACL and shape-derived checks reject answers that invent illegal edges.

Five design principles worth keeping

  1. Separate schema (ontology) from instance data (graph).
  2. Design from competency questions, not from buzzwords.
  3. Prefer small, testable constraints over monolith axioms.
  4. Validate on write; reason where it earns its cost.
  5. Treat LLM help as drafting assistance under human governance.

Conclusion

Ontologies are agreements made machine-checkable. RDF stores facts; RDFS and OWL add structure and inference; SHACL enforces instance quality; GraphRAG consumes the result as a safer map than vectors alone. The VOO example is small on purpose: if three sentences can exercise five layers, a real product catalog can too—once competency questions and owners exist.

Keep a one-page decision log for each major class and property: why it exists, which competency question it serves, and which SHACL shape guards it. Version IRI policies the way APIs version routes; renaming without redirects breaks every downstream GraphRAG join. When extractors propose new edges, require a shape or an explicit “unconstrained” sandbox graph so noisy LLM output cannot poison the governed store. Measure ontology value with question coverage and validation catch-rate, not with axiom count. Finally, pair every production GraphRAG deployment with a fixtures pack of legal and illegal triples so CI fails when a “helpful” schema edit silently widens domain/range and lets bad managers attach to funds again.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.

Document competency questions beside SPARQL fixtures so schema edits cannot drift from the asks production GraphRAG must still answer.