Mehdi Akiki
Published on

External APIs Will Change. Your Product Model Should Survive

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Derived state

External API changes are often discussed as versioning problems: a field was renamed, an endpoint moved, or a new version appeared.

The harder changes keep the same JSON and change its meaning.

A status called active starts including trial accounts. A timestamp changes from creation time to import time. An omitted field used to mean false, then begins to mean "not available to this caller."

The payload still parses. The product can still be wrong.

The principle is:

Protect product meaning from both schema changes and semantic changes at the source boundary.

There are at least three contracts

When a provider sends data into a product, I often hear "the API contract" as if there is only one. I find it more useful to separate three contracts:

  1. Transport contract: Can the boundary decode the payload?
  2. Mapping contract: Can the mapper translate source fields into product concepts?
  3. Product contract: Is the resulting meaning safe for product behaviour?

Consider this payload:

{
  "id": "42",
  "state": "active"
}

It can pass JSON validation while failing the mapping contract because active is a new enum value. It can pass mapping while failing the product contract because the source changed what active includes.

Parsing is necessary. It is not proof of understanding.

Do not let the source own your vocabulary

If a provider's state field is stored directly as the product's state, every provider change becomes a product change. A translation layer creates a deliberate decision point:

function mapLifecycle(source: ProviderRecord): Lifecycle {
  switch (source.state) {
    case "ready":
    case "active":
      return "usable";
    case "suspended":
      return "blocked";
    default:
      return "unknown";
  }
}

This code is simple, but it records a product decision: ready and active mean usable here. That decision can be tested, reviewed, and changed independently from the source decoder.

The companion article Why Every API Integration Needs a Stable Internal Model develops this boundary in more detail.

Unknown must remain different from false

One common schema-evolution bug collapses three states into two:

  • yes;
  • no;
  • not known.

If a field disappears because of permissions, false may be the wrong default. If a new enum value appears, mapping it to the nearest known value hides the change.

An explicit unknown state is sometimes inconvenient. It forces callers to decide what to do. This inconvenience is valuable when uncertainty is real.

For critical behaviour, fail closed or quarantine the record. For non-critical display, show a neutral state. The choice depends on the consequence, but it should be visible.

Preserve enough source evidence

When a mapping changes, old normalized data may need to be rebuilt. That is possible only if the source can be fetched again or I retained enough evidence from the original observation.

Useful provenance can include:

  • provider and resource type;
  • external identifier;
  • time observed;
  • source API version;
  • selected response headers;
  • payload hash;
  • mapper version;
  • raw or redacted source payload, when policy allows it.

This does not mean storing sensitive responses forever. Retention, encryption, access, and deletion requirements still apply. It means choosing consciously whether normalized data can be reproduced.

Without provenance, a strange internal value becomes an archaeology project.

Treat additive changes as signals

Many decoders are intentionally tolerant of unknown fields. This helps old clients survive when a provider adds data. Protocol Buffers is designed for this kind of evolution; its documentation distinguishes wire-safe, wire-compatible, and unsafe changes.

Tolerant decoding should not mean zero visibility. I record sampled schema fingerprints, unexpected enum values, or validation warnings. This lets the pipeline continue while showing that the outside contract moved.

There is a useful balance:

strict enough to protect meaning
tolerant enough to permit safe evolution
observable enough to notice the difference

Maximum strictness makes harmless additions an outage. Maximum tolerance turns semantic breakage into quiet data corruption.

Roll out mapping changes like code changes

A new mapping can rewrite product meaning for millions of stored records. It deserves a controlled release.

A safer sequence is:

  1. Make the reader understand both old and new representations.
  2. Observe the new source values without changing behaviour.
  3. Add the new internal representation.
  4. Backfill a bounded sample.
  5. Compare counts and important invariants.
  6. Expand the backfill and monitor it.
  7. Remove old handling only after old data and old writers are gone.

This is expand-and-contract applied to data meaning. The important part is that mixed versions exist during rollout. A migration plan that assumes one atomic deployment will eventually meet reality.

Test semantics with fixtures

Contract tests should include more than one ideal payload. I like fixtures for:

  • the smallest valid record;
  • all known enum values;
  • missing optional fields;
  • unknown additional fields;
  • an unknown enum value;
  • a record from the previous provider version;
  • values at important boundaries;
  • permissions that hide selected fields.

Mapping tests can then state product meaning directly:

expect(mapAccount(suspendedFixture)).toMatchObject({
  lifecycle: "blocked",
});

These tests do not prove that the provider will never change. They make the current interpretation executable.

Watch behaviour, not only validation errors

Semantic changes may pass every fixture. Operational comparisons help catch them:

  • sudden changes in records per lifecycle state;
  • a field becoming empty for one tenant or credential scope;
  • unexpected growth in unknown mappings;
  • freshness changing after a pagination update;
  • large differences between old and new mapper output.

For risky changes, run both mappings on the same observations and compare the results before switching the product to the new one.

What I took from working across system boundaries

Working across ingestion, data modeling, and runtime boundaries made schema evolution a product concern for me, not just connector maintenance. External data has to arrive, retain explicit meaning, and remain usable as both sides evolve.

The examples here are invented. The lesson comes from owning a boundary: a source schema, an internal meaning model, and runtime behaviour are related contracts, but they are not the same contract.

A schema-change checklist

When a provider changes, ask:

  • Did only the shape change, or did the meaning change too?
  • Can old and new records exist at the same time?
  • What does missing mean for this field?
  • How are unknown enum values represented?
  • Can normalized records be traced and rebuilt?
  • Which product decisions depend on this mapping?
  • Can old and new mapping results be compared before rollout?
  • What signal will tell me that the interpretation is wrong?

An integration survives change when the product owns its meaning and the boundary makes outside change visible.