Cloud & DevOps

Cloud-native application architecture

Cloud-native means stateless where it pays, boundaries that own their data, and failure designed for. Moving a lift-and-shift application into Kubernetes achieves none of it.

Guide 9 min read All insights

Define cloud-native honestly

Cloud-native is often used to mean running in a managed container platform, which delivers none of the properties the term implies. Lift a tightly coupled application into Kubernetes and you have added a scheduler without removing any coupling.

The properties that matter are unglamorous: stateless components where the work allows it, configuration declared rather than discovered, delivery automated and reversible, dependencies treated as unreliable, and every component observable before it fails.

Decide which of those you are buying. Managed databases and queues reduce operational work substantially and are usually worth the vendor dependence. Portability is rarely worth paying for on its own.

Ask what the architecture lets you change on a Tuesday, and what it lets you diagnose at 2am.

Boundaries own their data

A shared database between services is the most common way a microservices estate becomes a distributed monolith. Two services writing the same tables cannot be deployed, scaled or reasoned about independently, whatever the diagrams say.

Split by business capability rather than by technical layer. A service that owns enrolment owns the enrolment tables and publishes facts about changes; other services consume those facts instead of reading the source.

Distributed transactions then need an answer. Long-running business processes work better as a saga, where each step commits locally and a failure triggers compensating actions rather than a distributed rollback.

  • One owner per table, enforced by access policy rather than convention
  • Events published on change, with a stable schema
  • Sagas with explicit compensating actions instead of distributed transactions
  • Read models assembled for consumers rather than shared writes
order:
  owns: [orders, order_lines]
  publishes: order.confirmed, order.cancelled
  consumes: inventory.reserved, payment.captured
  compensation: release_inventory, void_authorisation

If two services can corrupt each other’s data, they are one component. Say so and stop pretending otherwise.

Stateless where it pays, stateful where it is honest

Statelessness is worth pursuing where it enables horizontal scaling and simple failover. It is not worth pursuing for a component whose entire job is to hold state, and forcing a session store to be stateless in the name of purity produces more moving parts, not fewer.

Decide deliberately where each kind of state lives: request-scoped data on the request, durable business state in a database, uploaded files in object storage behind expiring signed URLs, and short-lived derived data in a cache that is allowed to be wrong.

A cache is not a system of record. Anything you cannot regenerate or refetch should not be in one, and treating a cached value as authoritative produces a class of bug that only appears on eviction.

  • Durable state in a store chosen for its guarantees, not its speed
  • Object storage with expiring signed URLs for uploads and downloads
  • Caches treated as disposable, with an explicit expiry policy
  • Session affinity avoided unless a specific constraint requires it
  • File and message transports treated as untrusted input

Assume every dependency is down

In a distributed system, partial failure is the normal operating condition. A request will cross four services and one of them will be slow. The question is whether that becomes an outage or a degraded response.

Set a timeout on every call, including the ones to your own database. Retry only operations that are safe to repeat, with exponential backoff, jitter and a retry budget, so a struggling dependency is not amplified into a self-inflicted outage.

Use bulkheads so each dependency gets its own pool, and a circuit breaker to stop calling something clearly failing. Idempotency is what makes retrying safe, so the two decisions belong together.

  • Timeout on every outbound call, internal ones included
  • Retries with backoff, jitter and a budget that stops runaway amplification
  • Circuit breakers and bulkheads per dependency
  • Idempotent handlers plus a dead-letter path once retries are exhausted
  • Graceful degradation with a defined reduced response
call_policy:
  connect_timeout_ms: 500
  request_timeout_ms: 2000
  retry: { attempts: 3, backoff: exponential_jitter, budget: 0.25 }
  breaker: { error_threshold: 0.5, cooldown_s: 30 }
  idempotent: true

Delivery that can be repeated and reversed

Build one artefact and promote it. Rebuilding per environment reintroduces exactly the drift that made staging useless, and it means the thing you tested is not the thing you shipped.

Separate configuration from the artefact, and separate schema changes from application deploys. Expand and contract: add the new column, deploy code that writes both, backfill, switch reads, then remove the old column in a later release. That sequence is what makes rollback possible.

Rolling back code is easy. Rolling back data is not, which is why migration order matters more than deployment tooling.

  • One immutable artefact promoted across environments
  • Configuration supplied per environment and never baked in
  • Expand and contract migrations so deploys stay reversible
  • Feature flags for unfinished behaviour, each with an expiry
  • Health checks that gate traffic rather than merely restarting workloads

Observe before you scale, and count the cost

Instrument the request path before the first incident, not during it. Structured logs with a correlation identifier, metrics for the operations that matter, and traces that cross asynchronous boundaries are what turn an outage into a query.

Async is where tracing usually stops. A request that enqueues work and returns has not finished, and without context propagated into the queue you will reconstruct the relationship from timestamps.

Attribute cost to the component that incurs it. A per-service cost view is uncomfortable and useful, because the most common saving in a mature platform comes from removing a service or a duplicated capability rather than from tuning anything.

  • Structured logs, metrics and traces on every service boundary
  • Trace context propagated into queues and scheduled jobs
  • Objectives agreed with the business rather than chosen in a dashboard
  • Alerts phrased as user-visible symptoms
  • Per-service cost and resource limits reviewed on a schedule

Takeaways

  • Define what cloud-native buys you before choosing the platform
  • Give every service exclusive ownership of its data
  • Put timeouts, retry budgets and circuit breakers on every call
  • Use expand and contract so releases remain reversible
  • Propagate trace context into queues, and watch cost per service

Facing something similar?

These articles describe how we think about a problem. A conversation is where we would apply it to yours, against your actual constraints.

Start a conversation All insights