Frank Coyle, meshed with Kingsley Idehen's companion theses and the agent-rdf-memory system

Why Agentic Systems Need Ontologies

Probabilistic reasoning inside. Logical guardrails outside — and a working proof of the pattern.

Executive SummaryBy Frank Coyle · Medium · 2026-08-03

Synopsis

Frank Coyle's thesis that probabilistic AI agents need symbolic ontology guardrails, meshed with Kingsley Idehen's companion theses that LLMs are the missing generic RDF client and that a Semantic Web Project was waiting for AI, illustrated by the agent-rdf-memory system as a working ontology-guardrail deployment.

“Hallucination is a property, not a defect.”

View this analysis as a KG entity
Why Agentic Systems Need OntologiesFrank Coyle · Medium · the anchor thesis
LLMs as Powerful Generic RDF ClientsKingsley Uyi Idehen · LinkedIn
The Semantic Web Project Didn't FailKingsley Uyi Idehen · LinkedIn
agent-rdf-memoryOpenLink Software · GitHub · working proof-point
The Setup

Framing: A Talk That Filled the Room

Coyle's AI Engineer World's Fair talk on ontologies for agentic systems drew far more interest than expected, and most inbound questions were about hallucination — which the essay reframes before addressing directly.

The Argument

The Five-Move Argument

The sequential structure of Coyle's case for ontology guardrails in agentic AI.

95.1%0.995Five 99%-reliable agents chained end to end. Roughly five points lost to composition alone — no single component degraded.
36.6%0.99100A hundred-step workflow at 99% per step fails nearly twice as often as it succeeds.
60.6%0.995100Halving the per-step error rate nearly doubles end-to-end success. The same exponential, working in your favour.
1 in 1700.95100At 95% per step a hundred-step workflow does not degrade. It collapses.
Sequential reliabilities multiply — they do not average
End-to-end success rate as a function of chain length, for three per-step reliability levels.
0%25%50%75%100%0255075100Steps in the chain60.6%99.5% per step36.6%99% per step0.6%95% per step
Nobody’s intuition works this way: we look at a pipeline of excellent components and estimate that the whole thing is excellent, when the correct operation is exponentiation and the exponent is the length of the workflow. The gap between the amber and green curves is the argument for spending real engineering money on step-level reliability rather than hoping a retry loop covers it.
1

Hallucination Is a Property, Not a Defect

Hallucination and correct novel reasoning are the same generative machinery — sampling from a probability distribution — differing only in whether the output lands inside the truth; a model restricted to only what it has seen is a lookup table, not an agent worth building.

2

The Arithmetic Nobody Runs

Sequential per-step reliabilities multiply, not average: a chain of five agents each 99% reliable succeeds end to end only about 95% of the time; at 95% per step, a hundred-step workflow succeeds only about once in 170 attempts; validation placed only at the end of a pipeline turns debugging into forensics rather than prevention.

3

Ontologies to the Rescue

If the generative core must stay probabilistic, something outside it has to hold the line; an ontology supplies that line in two separable dimensions — a shared domain vocabulary, and logical constraints (RDFS/OWL domain, range, cardinality, disjointness) that ground agent output and catch violations at the step that produced them.

4

Wait — Expert Systems? Didn't We Try This Already?

Expert systems and CYC failed because they tried to hand-encode an unbounded model of general common sense; today's task is narrower — a domain expert specifying the bounded, dozens-to-hundreds-of-rules constraint set for one field — and authoring cost, the objection that actually sank the first attempt, has been largely dissolved by LLMs that can read a corpus and draft the tedious parts for a human to correct.

5

The Pendulum

Neural networks were dismissed after Minsky and Papert, then returned at scale as LLMs; expert systems were dismissed for reasons correct at the time; symbolic methods are now returning as the guardrails LLMs need before anyone can trust them with something consequential — ontologies are not a complete answer, since what counts as true in a domain stays contested, but the economics of building one, not the difficulty of being right, is what changed.

“Probabilistic reasoning inside. Logical guardrails outside.”Frank Coyle — the two separable dimensions of an ontology, once agents are running
Consensus layer

Shared Understanding of the Domain

Entities, relationships, and the vocabulary of what exists and how it connects — the dimension that forces a team to agree what a customer, a claim, or a policy is before anyone writes a prompt.

Non-probabilistic

Logical Constructs That Act as Guardrails

RDFS and OWL constraints — domain and range restrictions, cardinality, disjointness, class hierarchies with inferential teeth — grounding agent output so a violating entity is caught at the step that produced it, not a hundred steps downstream.

Cited Work

References

The Mesh

Synthesis: Guardrails, Clients, and a Working Proof

An agent-authored reading connecting Coyle's ontology-guardrail thesis to Idehen's LLM-as-RDF-client and Semantic-Web-Yin-Yang theses, using the agent-rdf-memory system as a concrete, running instance of the pattern Coyle argues for in the abstract.

Coyle reframes hallucination as a property of the generative substrate, not a defect to be prompted away. Idehen's RDF-clients thesis names the concrete fix at the systems level: connecting an LLM to an RDF-based Knowledge Graph anchors output to reliable, dereferenceable data and mitigates exactly the ungrounded interpolation Coyle describes — the ontology dimension is what makes that anchor meaningful rather than merely present.

Coyle's 'logical constructs that act as guardrails' — domain/range restrictions, cardinality, disjointness, class hierarchies with inferential teeth — is a description of RDFS and OWL by function rather than name; Idehen's RDF-clients thesis supplies the missing other half, arguing LLMs are the accessible interface RDF/OWL constraint graphs always needed to become usable rather than merely correct.

The agent-rdf-memory system that generated this very document is a running instance of the pattern Coyle argues for in the abstract: standing behavioral rules are encoded as schema:HowToStep instances in an RDF graph, not as prose an LLM can drift from mid-chain. Each generation step is checked against that graph before the next step consumes its output — directly addressing the composition-arithmetic problem Coyle names, where validation placed only at the end of a pipeline turns debugging into forensics. Catching a violation at step two, rather than after a hundred-step chain has silently compounded it, is precisely the economics Coyle says has changed since expert systems' first failure: authoring cost is low because the constraint graph is read and applied by the same LLM the guardrail is meant to check.

Meshes withagent-rdf-memory

Idehen's Yin/Yang thesis — that the Semantic Web Project's core ideas were sound but stayed hidden for two decades for lack of an accessible interface, until LLMs supplied one — is a specific historical instance of the general pendulum Coyle describes for symbolic AI overall. Both arguments locate the missing ingredient in usability rather than validity: what changed was not whether RDF/OWL or ontology-guarded agents were right, but whether a non-specialist could use them.

Responds toThe Pendulum
How-To

How-To Guide

1

Name the domain vocabulary before writing a prompt

Agree as a team on the shared understanding of the domain — what an entity is, what a relationship means, what a value can and cannot attach to — the first guardrail dimension. Most of the value shows up during the writing, in the arguments it starts.

2

Encode the bounded rule set as RDFS/OWL constraints

Specify domain, range, cardinality, and disjointness constraints for the dozens-to-hundreds of rules that matter in this one domain — not an unbounded model of general knowledge — using RDFS/OWL so a reasoner can act on class hierarchies with real inferential teeth.

3

Ground every agent step's output against the constraint graph

Check each step's output against the constraint graph before the next step consumes it, so an entity that violates the model is caught at the step that produced it — not a hundred steps downstream when someone notices the invoice went to a country that doesn't exist.

4

Let the LLM draft the tedious parts of the ontology

Use the LLM itself to read the domain corpus, propose structure, and draft the tedious parts of the ontology for a human expert to correct — the authoring-cost fix that dissolved the objection that sank ontologies the first time.

5

Treat the ontology as living, not settled

Expect deprecated concepts to accumulate and revisit the constraint graph over time — an ontology is one of the better tools for pulling probabilistic drift back toward the plausible, not a permanent, contest-free answer to what is true in the domain.

FAQ

Frequently Asked Questions

No. Coyle argues hallucination and correct novel reasoning share the same generative machinery — sampling from a probability distribution over what comes next — differing only in whether the output happens to land inside the truth. A model that could only emit what it had literally seen would be a lookup table, which defeats the purpose of building it.

Sequential per-step reliabilities multiply rather than average. Five agents each 99% reliable succeed end to end only about 95% of the time (0.99^5 ≈ 0.951) — a loss of roughly 5 percentage points from composition alone, with no individual component degrading.

It collapses: a system right 95 times out of 100 at every step succeeds end to end only about once in every 170 attempts across a hundred-step workflow, despite each individual step looking like it works fine.

Halving the per-step error rate moves a hundred-step workflow's end-to-end success from 36.6% to 60.6% — nearly doubling it — because the exponential that punishes long chains works just as sharply in your favor when per-step reliability improves.

A shared understanding of the domain — entities, relationships, and vocabulary the team must agree on before writing a prompt — and logical constructs that act as guardrails: RDFS/OWL domain, range, cardinality, and disjointness constraints that ground agent output and are not probabilistic, so they either pass or they don't.

Yes — CYC hit a wall trying to hand-encode a complete symbolic model of everything a person knows, an unbounded target that never scaled. What's being asked of ontologies today is narrower: a bounded set of dozens-to-hundreds of domain-specific rules, not a model of the world.

Authoring cost. What sank ontologies previously was the labor of building and maintaining them by hand; LLMs are now good at reading a corpus, proposing structure, and drafting the tedious parts for a human expert to correct — the economics of building an ontology changed, not the difficulty of being correct within one.

The pairing of probabilistic generation with symbolic verification: LLMs supply fluency and the ability to handle open-ended, messy input, while symbolic systems supply certainty within a bounded domain — neither does the other's job well, which is why the combination works.

Idehen argues LLMs are the missing accessible client RDF and the Semantic Web always needed — able to translate triples into natural language, mesh disparate data sources, and anchor output to structured data. That anchoring function is the concrete systems-level mechanism behind Coyle's abstract claim that ontologies keep probabilistic agents grounded.

Idehen's Yin/Yang thesis argues the Semantic Web's core ideas — HTTP URIs for naming, RDF for representation, ontology-guided navigation — were sound from inception but stayed hidden for two decades for lack of a usable interface, until LLMs and tools like OPAL supplied one. That is Coyle's neural-symbolic pendulum, already played out once at the scale of an entire W3C initiative.

agent-rdf-memory encodes standing behavioral rules as schema:HowToStep instances in an RDF graph an agent must consult before acting, rather than as prose it can drift from mid-task. That is a deployed, non-probabilistic constraint layer sitting outside the LLM's generation — exactly the second guardrail dimension Coyle describes, and it catches violations at the step that produced them rather than at the end of a long chain.

Coyle notes that validation placed only at the end of a pipeline turns debugging into forensics: a bad entity that entered at step two has every subsequent step faithfully reasoning over garbage. Catching it at the step it appeared is the cheapest possible intervention, and it's the specific failure mode a step-level RDF constraint graph like agent-rdf-memory is built to prevent.

No. Somebody still has to decide what's true in a domain, that decision is contested, and it doesn't stay decided — every production ontology accumulates a graveyard of deprecated concepts. What changed is the economics of building one, not the underlying difficulty of being right.

Domain and range restrictions (what a property may connect), cardinality constraints (how many values a property may take), disjointness axioms (which classes cannot overlap), and class hierarchies with real inferential consequences — all specified in RDFS and OWL 2, both formally standardized by the W3C.

Glossary

Glossary of Terms

Hallucination (Coyle's usage)

A generative-model output that is novel but wrong, produced by the same probability-sampling machinery that produces correct novel reasoning — a property of the substrate, not a removable defect.

Neurosymbolic AI

The pairing of probabilistic neural generation with non-probabilistic symbolic verification, so that fluency over messy input and certainty within a bounded domain each come from the component suited to it.

OWL (Web Ontology Language)

A W3C standard for expressing rich class hierarchies and logical constraints — cardinality, disjointness, domain/range restrictions — with inferential consequences a reasoner can act on.

RDFS (RDF Schema)

The base W3C vocabulary for describing classes, properties, and their domain/range relationships in an RDF graph, forming the foundation OWL builds richer constraint expressiveness on top of.

Expert System

A 1980s-90s AI approach that hand-encoded a domain expert's rules into a symbolic knowledge base; it hit a scaling wall when the target domain (general common sense, in CYC's case) was unbounded.

CYC

Douglas Lenat and Ramanathan Guha's decades-long project to hand-encode a complete symbolic model of common-sense knowledge, cited by Coyle as the monument to how high the expert-systems scaling wall was.

Composition Reliability (exponentiation over chain length)

The rule that sequential per-step reliabilities multiply rather than average, so a long agent chain's end-to-end success rate collapses even when every individual step looks reliable in isolation.

Guardrail Layer

A non-probabilistic component sitting outside an LLM's generation, checking agent output against logical constraints and either passing or failing it — the role Coyle assigns to ontologies and this document's synthesis section assigns concretely to RDFS/OWL and to agent-rdf-memory.

agent-rdf-memory

An RDF-Turtle-encoded behavioral memory system for AI coding agents, cited in this document's synthesis as a working, deployed instance of Coyle's ontology-guardrail pattern.

Semantic Web

The W3C-led effort to make the Web's data machine-interpretable using HTTP URIs, RDF, and ontologies; per Idehen's companion thesis, its core ideas were sound but stayed largely hidden until LLMs supplied a usable interface.

Linked Data

The practice of publishing structured data using dereferenceable HTTP URIs and RDF links so that facts about one entity can be traversed to related facts across independently published datasets.

Agentic System

A software system in which one or more LLM-driven agents take multi-step, tool-using actions toward a goal, often chaining several such steps or agents together — the composition-reliability problem Coyle centers his argument on.

Knowledge Graph Explorer 134 nodes · 326 links

Interactive graph visualization derived from the companion RDF. Click nodes to resolve, drag to explore. Graph data embedded from companion RDF at generation time.

Why Agentic Systems Need Ontologies

Nodes: 0 Links: 0
Click SVG to activate zoom, click outside to release | Drag nodes to pin, double-click to unpin
Classes Properties Instances

SPARQL Workbench 3 sample queries

Query this knowledge graph on URIBurner. The editor opens on the canonical SAMPLE entity-type summary (DAV named graph). Pick a recipe, edit freely, then run live or copy.

Query editor

▶ Run live on URIBurner SELECT: text/x-html+tr | DESCRIBE/CONSTRUCT: text/x-html-nice-turtle