---
name: system-design-checks
description: 15 rules from the Noesa course "System design, understood". For developers comfortable with basic APIs and databases who use AI to propose architecture but need to judge load, state, failure, consistency, security, and evolution. No distributed-systems background required.
---

# System design, understood — the rules

Use with: Claude Code or Claude (save as a skill), Cursor (save under .cursor/rules as .mdc), ChatGPT or any other assistant (paste the text below into custom instructions or a project's instructions).

15 rules, taken from the course at https://noesa.leafsoft.online/c/system-design

Each heading is one thing the course teaches. Most are checks to run on your own output before presenting it as done; a few are background you are expected to have. 15 also name a mistake models make by default, under "Watch for".

Apply these to the thing you are producing — the type, the schema, the query, the copy — not only to how you explain it. Where a rule names a field, a format or an identifier, that name belongs in the output.

## Draw the request path

Trace one Slot booking request from a customer tap to a confirmed record and back.

**Watch for:** If AI draws a clean request path from a feature description, what important part may still be fictional?

It may invent which component owns truth and which dependencies the user must wait for. Those choices depend on your service's actual correctness promise and failure tolerance, context absent from a generic feature description. Verify every arrow against the intended user outcome.

## Separate state from computation

Label each Slot component as stateful or stateless and explain the recovery consequence.

**Watch for:** An AI calls your workers “stateless, scales horizontally”. What do you check before you believe it?

AI can call a worker stateless because the pattern is common, without knowing whether an accepted booking is committed durably before the response returns. Hidden in-memory state passes review and fails the restart test silently. Trace each fact to its durable owner and confirm no user-visible outcome depends on a single worker's memory.

## Estimate load before choosing architecture

Make a rough Slot traffic estimate, state its assumptions, and identify the busiest interval.

**Watch for:** If AI recommends an architecture “for ten thousand users,” why is the number insufficient?

User count does not reveal actions per user, read-to-write ratio, payloads, concurrency, or bursts. AI defaults to common scale patterns when workload context is absent. Supply an explicit load envelope and ask which assumption changes the recommendation.

## Find the bottleneck

Identify Slot's current bottleneck from capacity and timing evidence instead of guessing from the diagram.

**Watch for:** If AI names your system's bottleneck from the architecture diagram alone, why might it point at the wrong component?

The limiting step depends on measured latency, capacity limits, and the current workload, none of which a diagram shows. AI tends to blame familiar suspects such as the database while the real wait sits on a slow external call. Supply per-step timings and one measurement that would disprove the diagnosis before you act on it.

## Cache without lying

Choose what Slot may cache, set a freshness rule, and protect booking correctness.

**Watch for:** If AI adds caching to make every endpoint faster, what judgment can it miss?

It may apply a common performance pattern without knowing the consequence of stale data for each field. your service's tolerated staleness and authoritative invariants are product context. Ask what each cache may return incorrectly and where the final rule is enforced.

## Queue work that can wait

Decide which Slot work belongs in a queue and define what happens when workers fall behind.

**Watch for:** If AI moves a slow step to a queue to speed up the response, what must you decide that the pattern itself does not?

Queuing changes the promise from "done" to "accepted," and the pattern says nothing about what happens when consumers fall behind. Whether the work may wait at all, its maximum acceptable age, and what your service tells users once that limit is crossed are product decisions. Confirm nothing the customer needs before proceeding was quietly moved off the response path.

## Partition data with care

Choose a partition key for Slot and identify a hotspot or cross-partition operation it creates.

**Watch for:** Why can an evenly drawn AI partition diagram still hide a severe imbalance?

The drawing lacks your service's real key distribution and per-tenant traffic. A common partition pattern assumes records are balanced, but a few large businesses can dominate. Supply skew evidence and test the hottest key, not only the average partition.

## Replicate for survival

Explain how replication helps Slot survive a storage failure and what lag changes.

**Watch for:** If AI adds replicas "for high availability," which harder question does that phrase leave unanswered?

More copies do not say whether a confirmed booking can vanish inside the replication-lag window, nor how a single current primary is chosen during failover. Those answers depend on your service's booking promise and require synchronous durability or coordinated promotion, not additional replicas. Make the tolerated loss window and the promotion rule explicit, and judge the design against them rather than against copy count.

## Choose consistency deliberately

Choose a consistency rule for Slot booking, profile, and reporting data during a network split.

**Watch for:** Why can “use eventual consistency for scale” be dangerous advice for a booking service?

It applies a broad pattern without the business invariant that one slot has one owner. Ask which operation may return stale or conflicting data, for how long, and with what consequence. The answer can differ field by field.

## Make retries safe

Make Slot's booking operation idempotent and explain how it handles an unknown first outcome.

**Watch for:** If AI adds retries around every failed network call, what hidden condition makes that unsafe?

The original call may have succeeded, and the operation may not recognize repeated intent. Retry helpers default to transport behavior, while effect identity belongs to the domain. Verify idempotency, atomic storage, and retry limits before enabling repetition.

## Design for partial failure

Choose how Slot degrades when one dependency is slow or unavailable.

**Watch for:** Why might AI's fallback keep the service online yet violate the product's core promise?

It may optimize for a response without knowing which fact cannot be guessed. Common fallbacks serve cached data, but stale availability cannot authorize ownership. State the invariant first, then judge whether reduced behavior preserves it.

## Observe what users experience

Choose Slot signals that connect a customer symptom to the responsible path.

**Watch for:** If AI adds "full observability" to your service, why can the dashboards still stay green while bookings fail?

AI defaults to infrastructure signals such as CPU and request counts, which can look healthy while correct booking decisions fall. The indicator that measures your service's actual promise, and the restraint to keep sensitive fields out of logs, come from your judgment about what users feel. Define the user-relevant SLI first, then decide whether the proposed signals reveal that outcome or merely the machines beneath it.

## Protect boundaries and secrets

Mark Slot's trust boundaries, separate authentication from authorization, and place secrets outside untrusted clients.

**Watch for:** An AI adds a login to your service. Why can one customer still read another customer’s data?

Authentication is a common pattern, but your service's membership and object-ownership rules are local context. AI may prove who called without scoping what they may access. Inspect every object query and action for server-side authorization.

## Evolve without breaking clients

Plan a compatible Slot interface and data change across old and new clients.

**Watch for:** Why can an AI-generated migration plan look complete while breaking live clients?

It often sees the desired end state but lacks client-version distribution, queued old work, and rollback constraints. A shortest-path replacement assumes coordinated deployment. Supply the coexistence window and demand staged compatibility evidence.

## Defend an AI-generated architecture

Audit an AI-generated Slot design, reject unsupported complexity, and defend a smaller architecture against named requirements.

**Watch for:** What is the strongest reason not to ask AI for one final “best architecture”?

There is no context-free best design. AI can compare mechanisms, but your service's traffic, promises, team capacity, failure tolerance, and change constraints determine the tradeoffs. Use it to expose alternatives and counterarguments, then require every recommendation to trace back to evidence you own.
