This article is published in English.
SHACL TDD for GraphRAG Agents: One Executive Cap Rule That Halts Bad Actions
Encode a 30% liability-cap executive-approval policy as SHACL, prove it with pytest, and watch an ontology firewall halt the agent on a $2.3M contract demo.
Reading architecture notes is not the same as shipping a governance rule. This piece picks up after a three-way comparison—plain RAG, GraphRAG, and GraphRAG wrapped with OWL/SHACL/policy—on a $2.3M sample agreement, and focuses on a transferable skill: encode one business policy as a shape, prove it with an automated check, and watch the agent refuse to proceed.
The Ontology RAG Firewall repository holds the cont: vocabulary, shape files, and offline demo used here. Clone it, confirm the suite is green on main, then optionally replay an older commit to experience the red-green loop firsthand.
Confirm the baseline on main
git clone https://github.com/cloudbadal007/ontology-rag-firewall
cd ontology-rag-firewall
pip install -e ".[dev]"
pytest -q # 18 passed (full suite)
python examples/demo_offline.py
Pinning to 6318929 is optional if the article’s last verified tip matters; tip of main may already be newer.
A healthy run shows 18 passed. The offline demo should already surface an executive-oriented warning on the indemnity section, along these lines:
Safe to act: 🚫 NO
- Flagged: 5
...
⚠️ EXECUTIVE APPROVAL: Liability cap is below 30% of contract value. Cap ratio: 25.00%. Agent action requires executive sign-off.
...
AGENT ACTION: HALTED. Routed to human review queue.
Total value protected: $2,300,000
That warning is exactly what the new shape introduced. The remainder reconstructs how it landed via tests-first development.
The policy being encoded
Eleven node shapes already live in contract_domain_shacl.ttl covering payment terms, notice periods, uptime SLAs, low-confidence extraction, remedy gaps, auto-renewal, missing indemnity language, direct-damages review, a 10% cap-to-value ratio, a high-value low absolute-cap case, and the executive ratio shape this walkthrough highlights.
On the demo deal, a $575K ceiling is 25% of $2.3M—above the 10% floor—so the older ratio shape stays silent. The executive shape closes that blind spot for expensive agreements.
Procurement’s extra requirement, in plain language:
Whenever agreement value is at least $500K and the indemnity ceiling sits under 30% of that value, an agent must obtain executive sign-off before acting.
That sentence becomes ExecutiveCapRatioShape plus a pair of pytest cases.
Step 1 — Fail first
Always author the assertion before the TTL.
On today’s main those checks already pass. To feel the failure, move to bbeb15e (pre-shape), insert the tests, observe red, add the shape from Step 2, then return to main.
Add or compare this case in tests/test_shacl_constraints.py:
def test_liability_cap_below_30_percent_on_high_value_contract() -> None:
"""
25% cap on a $2.3M contract must trigger ExecutiveCapRatioShape.
Existing shapes (10% ratio, $100K absolute) do not catch 575K / 2.3M.
"""
clause = ExtractedClause(
"test-cap-ratio",
"LiabilityClause",
"text",
{"liabilityCap": 575_000, "liabilityScope": "DirectDamagesOnly"},
0.9,
1,
)
graph = ClauseRDFBuilder().build(clause, 2_300_000)
conforms, violations, _ = SHACLContractValidator().validate(graph)
assert not conforms
assert any(
"30%" in v or "executive" in v.lower() for v in violations
), violations
def test_liability_cap_at_32_percent_no_executive_flag() -> None:
"""32.6% cap on $2.3M should not trigger the 30% executive rule."""
clause = ExtractedClause(
"test-cap-ratio-ok",
"LiabilityClause",
"text",
{"liabilityCap": 750_000, "liabilityScope": "FullDamages"},
0.9,
1,
)
graph = ClauseRDFBuilder().build(clause, 2_300_000)
_, violations, _ = SHACLContractValidator().validate(graph)
cap_ratio_hits = [
v for v in violations if "30%" in v or "executive" in v.lower()
]
assert len(cap_ratio_hits) == 0, cap_ratio_hits
Execute:
pytest tests/test_shacl_constraints.py::test_liability_cap_below_30_percent_on_high_value_contract -v
What red looks like
Before the shape exists (for example on bbeb15e):
FAILED tests/test_shacl_constraints.py::test_liability_cap_below_30_percent_on_high_value_contract
AssertionError: ... executive ...
On current main the identical invocation is green. Next comes the TTL itself—already merged upstream, reproduced so the pattern is reusable.
Step 2 — Author the shape
Append to ontologies/contract_domain_shacl.ttl. Keep cont: pointed at the OWL namespace via the raw GitHub IRI (avoid inventing a /contract# path that does not resolve):
https://raw.githubusercontent.com/cloudbadal007/ontology-rag-firewall/main/ontologies/contract_domain_owl.ttl#
Instance builders mint URIs under the same base (…#instance/).
Replaying on bbeb15e? Reuse the prefix URI already present in that revision’s SHACL file. On tip main, prefer the raw IRI so ontology, shapes, and tests agree.
cont:ExecutiveCapRatioShape a sh:NodeShape ;
sh:targetClass cont:LiabilityClause ;
sh:severity sh:Warning ;
sh:message "⚠️ EXECUTIVE APPROVAL: Liability cap is below 30% of contract value. Cap ratio: {?capRatio}%. Agent action requires executive sign-off." ;
sh:sparql [
a sh:SPARQLConstraint ;
sh:select """
PREFIX cont: <https://raw.githubusercontent.com/cloudbadal007/ontology-rag-firewall/main/ontologies/contract_domain_owl.ttl#>
PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>
SELECT $this ?capRatio WHERE {
?contract cont:hasLiabilityClause $this ;
cont:contractValue ?v .
$this cont:liabilityCap ?cap .
BIND((xsd:decimal(?cap) / xsd:decimal(?v) * 100) AS ?capRatio)
FILTER (xsd:decimal(?v) >= 500000)
FILTER (?capRatio < 30)
}
""" ;
] .
Three design choices are intentional:
- The
FILTERrequiring value>= 500000scopes the rule to high-value deals; the same percentage means something different on a $50K SOW. - Embedding
{?capRatio}in the human message gives reviewers the measured percentage instead of a vague warning. - The 30% threshold is organisational policy—edit the literal if Legal wants 40%. The file is the policy artifact.
Step 3 — Map violations onto readable clauses
Shapes emit machine violations; the firewall turns keyword hits into clause-level report lines. On main, EXECUTIVE is already listed among liability tokens in firewall.py:
"LiabilityClause": ["LIABILITY", "LOW CONFIDENCE", "LEGAL REVIEW", "HIGH-VALUE", "EXECUTIVE"],
When replaying bbeb15e, add that token alongside the shape—or the demo may compute the violation without attaching it to the indemnity clause.
Step 4 — Re-run suite and demo
pytest -q # 18 passed (entire repo)
pytest tests/test_shacl_constraints.py -v # 8 passed (this file)
python examples/demo_offline.py
Expect the full suite in a handful of seconds. The demo report gains a dedicated executive line on the indemnity clause (flag counts may stay flat when multiple violations share one clause):
⚠️ EXECUTIVE APPROVAL: Liability cap is below 30% of contract value. Cap ratio: 25.00%. Agent action requires executive sign-off.
AGENT ACTION: HALTED. Routed to human review queue.
Total value protected: $2,300,000
One policy encoded, proven, and visible. That loop is what scales to the next domain rule.
Why not “just prompt it”?
Stuffing the same 30% guidance into a system prompt fails when counsel rephrases the clause, when the instruction sinks in a long context, when someone edits prompts without compliance context, or when auditors ask which build applied which rule on which day.
A shape is deterministic over typed RDF, lives in version control, carries a regression test, emits structured evidence including the measured ratio, and cannot vanish because someone chased fluency elsewhere.
Governance for agentic systems needs formal constraints beside retrieval and generation. Policy lives in the shape; proof lives in the test; the violation text is the audit artifact.
Repeatable extension recipe
docs/extending.md spells this out; the short form is:
- State the rule in language a compliance owner recognises.
- Extend OWL only when new types or properties are required.
- Add one shape per rule—small and separable beats a monolith.
- Land a pytest that is red before the shape and green after.
Target state: every shape has a test; every test maps to a business consequence. Legal edits TTL; CI runs proofs; posture stays reviewable.
Roadmap beyond a single agreement
The firewall today exercises one-agreement paths. Two production gaps remain: batch summaries across a portfolio (examples/demo_batch_processing.py is a seed), and vendor multi-hop context (holds, incidents) once a linked property graph exists—not a stub URL.
Until then, practice the loop: write the shape, prove it, watch the halt, then carry the same pattern into the next domain.
Keep a sidecar ledger for each shape—owner, effective date, source memo id, and pytest node id—so a git blame line expands into an audit trail. When thresholds move, version the message string and add boundary tests so old cutoffs cannot silently return. Treat keyword→clause maps as API surface: snapshot demo output in CI so refactors cannot drop EXECUTIVE from the indemnity token list without a failing job. Prefer additive shapes over editing shared SPARQL blobs; independent shapes revert cleanly when a policy experiment fails. Finally, publish the measured ratio in every human-facing warning—reviewers trust numbers they can recompute from the RDF more than generic “needs approval” banners.