Mehdi Akiki
Published on

An Ontology Is an Executable Data Contract, Not a Diagram

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Reference · Derived state

The word ontology can make a practical engineering problem sound academic. I used to imagine a large diagram, many abstract classes, and long meetings about naming.

The useful starting point is smaller:

An ontology states which things exist in a product, how they relate, and which distinctions the software must preserve.

Once ingestion, validation, search, rules, or automation depend on those statements, the ontology is not only documentation. It is a product contract.

A schema tells shape; an ontology adds meaning

A JSON schema can tell me that a record contains two strings:

{
  "subject": "account:42",
  "owner": "person:7"
}

It does not necessarily tell me:

  • whether owner means legal owner, operator, or creator;
  • whether an account can have several owners;
  • whether a team can be an owner too;
  • whether ownership is current or historical;
  • what should happen when the owner is deleted.

Those are semantic decisions. They affect behaviour across the product.

An ontology can make concepts and relationships explicit:

Person ──memberOf──> Organization
Account ──controlledBy──> Organization
Subscription ──appliesTo──> Account

The notation is not the important part. The shared meaning is.

The contract lives between systems

An ontology becomes especially valuable when several parts of a system need the same answer.

external data
      ↓
mapping ──> shared concepts ──> validation
                              ├──> search
                              ├──> rules
                              └──> product behaviour

Without a shared model, every consumer interprets source fields again. One service treats account.owner as a person. Another accepts an organization. A third uses the record creator because the names looked similar.

The bugs do not look like type errors. Each component works according to its local assumption.

The ontology gives these assumptions one place where they can disagree visibly.

Model distinctions that change behaviour

It is easy to model too much. A team can spend weeks deciding whether a company is an organization, a legal entity, or an economic actor before one product feature needs the distinction.

I prefer one question:

If I merge these two concepts today, which behaviour becomes incorrect or impossible later?

If the answer is "none that I know," one concept may be enough. If permissions, validation, billing, or automation behaves differently, the distinction is probably real.

This gives me a practical rule:

  • model a distinction when the product needs to preserve it;
  • write an example showing the difference;
  • delay distinctions that have no observable consequence yet.

The goal is not to describe the whole world. It is to give the product a language precise enough for its decisions.

Identity is part of meaning

Two records with the same name are not necessarily the same entity. Two records with different names may describe the same one.

An ontology needs an identity strategy, not only classes and properties. For integrated data, it can help to keep source identity separate from product identity:

Source observation:
  provider = crm-a
  type = company
  external_id = 42

Product entity:
  organization_id = org_9f3...

Mapping one to the other is evidence, not a free assumption. The product may know that several source records refer to the same organization, but this needs a rule, confidence, or human decision.

If identity is vague, relationships will be vague too.

Constraints are not the same as inference

This difference is important in semantic systems.

An ontology can express knowledge from which more knowledge is inferred. Validation asks a different question: does this data satisfy the shape required here?

For example, not seeing a birth date does not prove that a person has no birth date. It may only mean that the graph does not contain this fact. This is related to the open-world style of reasoning used by OWL.

A product input can still require exactly one birth date for a particular workflow. That is a validation constraint.

The W3C overview of OWL 2 describes ontologies in terms of classes, properties, individuals, and data values. SHACL provides a separate way to validate an RDF data graph against shapes.

Keeping these jobs separate avoids a common confusion:

ontology: what can this concept mean?
constraint: what must this operation receive?

The product usually needs both.

Put the model on an execution path

A model that lives only in a diagram will become stale. A model used by software receives feedback.

Useful execution points include:

  • validating mapped source records;
  • generating typed interfaces or documentation;
  • checking whether a relationship is allowed;
  • driving search facets;
  • explaining why a rule matched;
  • testing example entities in continuous integration.

This does not require generating the whole application from an ontology. One narrow executable use is enough to reveal whether the model is precise.

If engineers must bypass the model for every real case, the model is giving a useful warning. It may be too abstract, too strict, or missing a product concept.

Version meaning, not only syntax

Changing owner to controlledBy may be a rename. Changing its allowed target from Person to Person | Organization changes possible product states.

Ontology evolution needs questions similar to API evolution:

  • Can old and new data coexist?
  • Does a new class change existing inference?
  • Will a stricter constraint reject stored entities?
  • Does a relationship rename preserve identity?
  • Can consumers understand unknown concepts?
  • Is a backfill required?

For important changes, keep representative example graphs and expected outcomes. Run the old and new model against them. A version number alone does not explain the semantic difference.

Use examples as executable conversations

Abstract definitions become clearer with small examples:

Given:
  Sam is a member of Northwind.
  Account 42 is controlled by Northwind.

Expected answer:
  Which organization controls Account 42?

Unsupported inference without another rule:
  Sam personally controls Account 42.

This example gives product, domain, and engineering people something concrete to challenge. It can later become a test fixture.

Examples are also where missing temporal context appears. "Sam is a member" may need validFrom and validUntil. I learn this from a product question, not from making the diagram larger in advance.

What I learned from working with semantic models

The main lesson for me was that modeling and implementation cannot be separate phases. Data arrives with source-specific assumptions. A shared model translates those assumptions. Runtime behaviour then reveals whether the translation was useful.

The examples in this article are generic and invented. The principle is reusable: the model earns its place when it makes product decisions safer and easier to explain.

A practical ontology checklist

For every new concept, ask:

  • Which product behaviour needs this distinction?
  • What gives an entity its identity?
  • Is this a source fact or my interpretation?
  • Which relationships are allowed?
  • Which constraints belong to a specific workflow?
  • What can safely remain unknown?
  • Which example proves the model is useful?
  • How will a semantic change be detected and migrated?

A good ontology is not the largest description. It is the smallest shared language that keeps important meaning intact.