The contract comes before the implementation
API-first is routinely mistaken for having an API. The discipline is about order of work: the contract is designed, reviewed and agreed while the implementation is still an implementation detail, so disagreement happens while it is cheap.
Write the schema first, in OpenAPI or JSON Schema, and generate or validate against it. The artefact under review should describe behaviour, errors and authorisation, not a bare list of endpoints. A schema of routes alone tells you nothing you could not read from the code later.
Treat an internal API as a public one. The moment two teams depend on it, changing it becomes a coordination problem rather than a refactor, and you have already built the habit of versioning.
- Schema written and reviewed before any handler code
- Contract committed to version control beside the implementation
- Client code generated from the contract rather than written by hand
- Errors, authorisation and pagination designed as part of the contract
- Reviewed by an actual consumer, not only by the author
Model states and transitions, not just CRUD
An API that exposes only create, read, update and delete pushes the state machine into every consumer, and each consumer gets it slightly wrong. If a claim moves from submitted to assessed to approved, that transition belongs in the contract.
Make invalid transitions explicit. Return a distinct error naming the attempted transition and the current state, so the client can recover instead of guessing. It documents the domain for anyone integrating later.
Prefer specific errors over a generic failure with a message attached. A machine-readable code, a human-readable explanation and a hint about the next action save every consumer from writing its own interpretation.
error:
code: invalid_transition
message: "A claim in state 'closed' cannot be resubmitted."
current_state: closed
attempted: resubmit
allowed: [reopen, escalate]
Idempotency is part of the design, not a later patch
Networks retry, clients retry and queues redeliver. Any endpoint that creates or moves something will eventually be called twice with the same intent, and the design has to decide what the second call does.
Accept an idempotency key on unsafe operations, fingerprint the request, store the outcome with the key, and replay the stored response on a repeat instead of re-executing. Define how long keys are retained, because a client retrying tomorrow still needs a correct answer.
The same rule applies inbound. A webhook handler that processes a delivery twice should end in the same state as one that processed it once, which usually means a natural key on the business object rather than a synthetic event identifier.
- Idempotency key on every create, update and side-effecting operation
- Request fingerprint stored with the key so misuse is detectable
- Stored response replayed rather than re-executed
- Inbound handlers keyed on business identity, not event id
idempotency:
header: Idempotency-Key
scope: organisation + endpoint
fingerprint: sha256(canonical_body)
retention: 24h
conflict_on_mismatch: 409
Version without regret
Most API breakage is avoidable. Adding an optional field, adding an endpoint, adding a value to an enum that clients switch over, changing the meaning of an existing field: each has broken a consumer that had no way to notice.
Additive change is the default, and the exceptions deserve real scrutiny. Widening an enum, making an optional field required, tightening validation and reinterpreting a field are all breaking even when the shape is identical.
Deprecate with evidence rather than announcement. Send a sunset header, measure actual usage per field and per consumer, and remove something only once telemetry shows it unused. Announcing a deprecation to consumers you cannot observe is a guess.
- Additive change by default, with enum widening treated as breaking
- Field semantics documented and never silently changed
- Deprecation measured with per-consumer usage telemetry
- A stated deprecation window with the date exposed in headers
- Version boundaries at the API rather than inside every client
Authentication, authorisation and quota are one design
Authentication answers who is calling. Authorisation answers what they may do. Quota answers how much of it they may consume. Designed separately, these produce APIs that are secure in a demonstration and awkward in production.
Use OAuth 2.0 or OIDC for user-facing access and scoped service credentials for machine callers, with scopes expressed in the vocabulary of the business rather than in table names. A new consumer can then be granted a capability without understanding the database.
Per-consumer quotas are as much a product decision as a technical one. They protect the platform from a single integration and give you a lever when a partner genuinely needs more. Issue keys with an expiry and a rotation path, and keep every call attributable to a known credential.
- OAuth 2.0 and OIDC for user-facing access
- Scoped, rotatable credentials for service-to-service calls
- Per-consumer rate limits and quotas with a documented review path
- Authorisation checked at the resource, not only at the route
- Secrets held in a secret store, never in source or logs
Test the contract, not only the handler
Unit tests on handlers prove the code runs. They do not prove that two independently deployed components still agree. Consumer-driven contract tests do, by asserting the provider still returns what the consumer expects.
Validate every response against the published schema in the pipeline. A handler that starts returning an undocumented field, or an undocumented null, has changed the contract and should fail the build rather than surprise a consumer.
Keep recorded request and response pairs for the critical operations and replay them against changes to validation or authorisation. That is where regressions hide, because the code path still looks correct while the rule has narrowed.
- Consumer-driven contract tests between producer and consumer
- Response validation against the schema on every build
- Recorded fixtures replayed against validation and authorisation changes
- A sandbox with realistic, non-production data
- Rate-limit and throttling behaviour tested rather than assumed
If only one team consumes the API, contract discipline is the only thing protecting the next team.
Takeaways
- Design and review the contract before writing the handler
- Expose state transitions, not only CRUD
- Treat idempotency and retention as part of the contract
- Assume change is additive, and measure before deprecating
- Test the contract, not just the endpoint