Reviewed: 9 October 2026. This is an original fictional system-design discussion. The service, timings and failure records are supplied practice inputs, not a deployed system, measured service-level agreement or company interview question.
Begin with the failure behaviour the user needs
A system-design interview can become more concrete when you explain what the caller should receive during a failure. Adding a cache or repeating a request is not enough: define what data may be returned, when the caller must stop waiting and whether repeating the operation could create a duplicate effect.
Microsoft Learn's retry-pattern guidance discusses transient faults, bounded retries, delays, idempotency and the extra load that layered retries can create. It also notes that a request may complete even when its response is lost. This is narrow design context. The exercise below supplies its own service contract and timings; it does not reproduce a Microsoft blueprint or establish a production guarantee.
Ask which operations are reads, which change state, how old a fallback may be and what the deadline includes. If the brief does not answer one of those questions, keep it unresolved rather than making a hidden assumption to justify your diagram.
Original brief: a public opening-hours lookup
The fictional service returns a published opening-hours record for a community desk. A request contains a desk identifier and asks for the currently published record. The operation does not create a booking or send a message. Its read-only contract is an explicit supplied assumption, not something inferred from an HTTP method name.
| Requirement | Supplied boundary |
|---|---|
| Caller deadline | Return a response within 500 milliseconds from receipt, including local processing and waits. |
| Local work | Reserve 50 milliseconds for local processing and response completion. |
| Upstream attempt | Each attempted lookup has a 150-millisecond bound in the exercise. |
| Retry wait | A retry, if permitted, waits 50 milliseconds after the prior failure. |
| Fallback | A stored published record may be returned if its recorded age is at most 10 minutes, visibly labelled with that age. |
| No usable record | Return an unavailable response rather than inventing opening hours. |
These numbers are stipulated for arithmetic, not benchmark results. The brief does not establish that a real network call or local task can always meet those bounds. A production design would need actual measurements and cancellation behaviour before making a deadline promise.
Compare two attempts with three
One initial attempt plus one retry means two total attempts. At the supplied bounds, the budget is 50 milliseconds local work + 150 first attempt + 50 wait + 150 second attempt = 400 milliseconds. That leaves 100 milliseconds inside the supplied 500-millisecond budget.
One initial attempt plus two retries means three total attempts. Its corresponding bound is 50 + 150 + 50 + 150 + 50 + 150 = 600 milliseconds. It exceeds the deadline by 100 milliseconds. Calling that policy “only two retries” does not make its three attempts fit.
An original answer could be: “Under this brief's stated bounds, I can allocate at most two attempts with one 50-millisecond wait while retaining the reserved local work. Three attempts do not fit. I would track the remaining overall deadline, stop launching attempts that cannot fit and verify cancellation and response completion in implementation. The arithmetic is a proposed budget, not a measured latency result.”
That answer connects the policy to the total deadline. It avoids presenting an upstream timeout as the only time cost. If the local reserve changes or an operation includes additional work, recompute the budget rather than silently spending the existing slack twice.
State what each failure observation actually means
The fictional failure packet contains an upstream timeout after the first attempt. It does not say whether the upstream completed the lookup or whether a late response exists. The caller knows only that it did not receive a usable response within that attempt's bound.
Because the supplied lookup is read-only, repeating it does not create a new booking or send a duplicate notification under this exercise's contract. That reasoning would not transfer automatically to a state-changing operation. A request named “lookup” could still have side effects in a different system; check the actual contract.
A hypothetical second attempt might return a valid published record, fail again or report a permanent invalid identifier. Those outcomes require different handling. An invalid identifier is not made valid by waiting and retrying. The exercise supplies no actual second-attempt response, so the answer should describe proposed branches without inventing a successful recovery.
Evaluate the fallback against the contract
Consider a separately supplied stored record whose age is seven minutes. It meets the brief's maximum age of ten minutes. If it is the requested desk's published record and is otherwise usable, the proposed fallback may return it with a visible age label. Passing the age condition alone does not establish those other facts.
An eleven-minute record fails the supplied age condition. Do not return it as current simply to avoid an unavailable response. The brief explicitly permits unavailability when there is no usable record. That response may be less convenient than old hours, but it honours the defined information boundary.
The exercise does not supply a real stored record or prove that a cache remains available during an upstream outage. A diagram with a fallback box is therefore a proposed design element. Its existence, permissions, population and failure dependencies would need verification before claiming that it works in a deployed system.
Avoid multiplying retries across layers
Suppose a hypothetical caller layer permits three attempts, and each of those invokes a client layer that independently permits three downstream attempts. Under the supplied nesting assumption that each inner sequence runs fully and the outer layer repeats after failure, there could be nine downstream attempts: three times three.
That is not six attempts, and it is not a universal bound for every layered system. Overall deadline cancellation or different policies could stop the sequences earlier. The example identifies a reason to inspect where retries occur, rather than claiming that nine attempts were observed here.
For this original brief, a proposed design assigns one layer responsibility for the overall deadline and finite attempt policy. Any underlying client's behaviour still needs to be inspected; naming one owner does not prove hidden retries are disabled. Include that inspection in the implementation checks.
Show a complete proposed request path
The service first validates the desk identifier against its actual accepted identifier rules, which the brief has not specified in detail. It then starts the bounded lookup while tracking the total deadline. If a usable response arrives, it returns that record. If the first attempt has a retryable failure and the remaining budget permits another attempt, it waits once and makes the second attempt.
If no usable upstream response is available, it checks a usable stored published record against the requested identity and age condition. A permitted fallback is labelled with its age. Without one, it returns the defined unavailable response. A permanent invalid-identifier response should follow the agreed invalid-input contract, rather than being disguised as temporary unavailability.
This path is a proposal. There is no executed service, storage implementation, monitoring result or uptime measurement in the packet. An interviewer can ask about each unresolved implementation boundary without the answer pretending that those boundaries are already solved.
Test the design with observable assertions
An implementation test would need to cover a successful first response, a retryable first failure, a permanent failure, a late response, an eligible fallback and a too-old fallback. It should inspect total elapsed time, actual downstream attempt count and the response's stated data age.
Also check that abandoned calls do not unexpectedly continue consuming resources or overwrite a chosen response. The exact cancellation mechanism is not supplied here. Record what was actually observed instead of inferring cancellation merely because the caller stopped waiting.
For a state-changing concurrency exercise, see the reservation system-design guide. For a finite algorithm and its boundary tests, use algorithm practice. These are distinct scopes: a correct array function does not demonstrate service availability.
Frequently asked questions
Does a timeout prove that a request did not complete?
No. It establishes that the caller did not receive a usable response within the stated bound. Remote completion may still be unknown.
Is a fallback always better than unavailability?
Follow the actual contract. This exercise permits only a usable published record within the stated age limit, with its age labelled.
Does the 400-millisecond calculation prove the service meets its deadline?
It proves the arithmetic under the supplied bounds. Actual implementation, network behaviour and cancellation still need testing and measurement.
