Editorially revised on 9 October 2026.
Software engineering practice should connect implementation with requirements, tests and operational boundaries. This guide contains six original questions about a small event-processing service. It is a worked practice set, not a claim that a named employer uses these questions or that every software interview follows one sequence.
The code is a self-contained Python exercise. The service examples are hypothetical and do not run against production infrastructure. Explain assumptions before extending the design; a short function is not a complete distributed system.
Question 1: how would you define the contract?
Suppose an exercise asks you to accept a list of event identifiers, remove duplicates and preserve the first occurrence of each identifier. Identifiers must be non-empty strings without surrounding whitespace. Invalid input raises ValueError. The function returns a new list and does not modify the input.
Clarify those requirements before coding. Should identifiers be case-sensitive? In this exercise they are. Does “remove duplicates” mean sort them? Here it does not. Should whitespace be trimmed? Here it is rejected instead, because silently changing an identifier could change its meaning.
These choices are exercise assumptions, not universal API requirements. A different contract might permit normalisation, but it must define that behaviour. Write down what should happen on empty input, invalid elements and repeated identifiers.
Question 2: can you implement the agreed behaviour?
Here is an original implementation for the stated list-of-strings contract:
def unique_event_ids(ids):
if not isinstance(ids, list):
raise ValueError("Expected a list")
seen = set()
result = []
for item in ids:
if not isinstance(item, str):
raise ValueError("Expected strings")
if not item or item != item.strip():
raise ValueError("Invalid identifier")
if item not in seen:
seen.add(item)
result.append(item)
return result
For ["b", "a", "b"], the result is ["b", "a"]. The input order determines first appearance, and the set tracks membership. The Python data-structures tutorial describes sets as collections of distinct elements; the separate result list preserves the required output order.
Under usual average-case hash-table assumptions, processing n identifiers takes expected O(n) membership operations, with O(k) extra storage for k distinct identifiers. String hashing and comparison also depend on identifier lengths, so do not claim constant total cost independent of input size or content. An in-memory set is not a durable record of previously processed requests.
Question 3: which tests would expose a wrong implementation?
Use examples that distinguish the contract from common mistakes. An empty list should return an empty list. All-unique input should retain order. Repeated values should retain the first occurrence. Mixed case should stay distinct. Invalid input should raise the agreed exception rather than partly return a result.
The following finite checks exercise those cases. They are original educational checks, not an exhaustive proof:
assert unique_event_ids([]) == []
assert unique_event_ids(["b", "a", "b"]) == ["b", "a"]
assert unique_event_ids(["A", "a"]) == ["A", "a"]
original = ["x", "x", "y"]
assert unique_event_ids(original) == ["x", "y"]
assert original == ["x", "x", "y"]
for bad in (None, "x", [1], [""], [" x"]):
try:
unique_event_ids(bad)
except ValueError:
pass
else:
raise AssertionError("Invalid input accepted")
A sorted-output implementation would fail the b, a order case. A case-folding implementation would fail the mixed-case case. A function that changes the supplied list would fail the mutation check. Testing only a single duplicate example would miss those distinctions.
You could add property checks: every output appears in the input, no output appears twice, and applying the function again returns the same sequence. Those properties complement concrete cases; they do not replace validation tests for the input contract.
Question 4: does this make a network request idempotent?
No. Deduplicating identifiers inside one call does not establish how the service handles two requests, a restart or two concurrent workers. Durable processing needs a separately defined operation identity and storage behaviour.
RFC 9110 defines idempotency through the intended server effect of repeating an identical request. It does not mean every response must be identical or that a client can retry every POST safely. Discuss the actual API contract and side effects before proposing retries.
For a fictional event endpoint, ask whether a repeated identifier with different content is accepted, rejected or treated as a new operation. Ask how long the deduplication record is retained and what happens during concurrent submissions. Those questions affect the design; a set local to one process does not answer them.
A timeout creates another uncertainty: did the server commit the operation before the client stopped waiting? Avoid saying that a timeout proves no work occurred. A retry strategy should account for that uncertainty and the endpoint's documented semantics.
Question 5: what belongs in the database transaction?
Imagine storing an accepted event and updating an internal processing record. If the two database changes must succeed or fail together, explain the transaction boundary. PostgreSQL's transaction tutorial describes grouping statements with BEGIN and COMMIT and using ROLLBACK to cancel a transaction.
A transaction is not a promise that an external email, remote service or file write rolls back with the database. If an operation includes an external side effect, discuss its failure and retry handling separately. Do not claim “exactly once” merely because two SQL statements share a transaction.
For the hypothetical service, a database uniqueness rule may help coordinate concurrent inserts for the operation identity. It still needs a defined response when the record already exists and a strategy for any downstream work. Explain those decisions without pretending the brief has supplied every requirement.
Now consider reading accepted events a page at a time. A keyset design might use a unique, immutable, ordered identifier and request entries greater than the last identifier. State those assumptions explicitly. An identifier that can change or an order with unresolved ties changes the correctness argument.
Even with a defined order, new writes during traversal require a consistency decision. Are you presenting a moving feed or a fixed snapshot? A cursor alone does not establish snapshot consistency. Ask what the product needs before promising that every row will appear exactly once in a changing dataset.
Question 6: how would you investigate a repeated event?
Suppose a user reports that an event appears twice. Start by distinguishing duplicate identifiers, duplicated display rows and two different operations that happen to look similar. The symptom does not tell you which layer is responsible.
A hypothetical investigation could correlate permitted request identifiers, operation records and downstream processing states. Check whether both submissions reached the service, whether a uniqueness constraint was used as intended and whether the display query duplicated a join result. Use the minimum necessary information and follow the system's access and privacy rules.
Do not begin by disabling safeguards or changing production records to make the symptom disappear. Establish the facts, reproduce a safe finite case where possible and explain what evidence would support each hypothesis. If a fix is proposed, identify the behaviour it changes and a test that would detect recurrence.
Organise your explanation around trade-offs
| Decision | Boundary to explain |
|---|---|
| In-memory deduplication | Limited to the inputs and lifetime of that function or process. |
| Durable identity record | Needs retention, conflict and concurrency rules. |
| Database transaction | Groups database work; external side effects need separate handling. |
| Pagination cursor | Needs ordering and a consistency policy. |
| Retry | Depends on operation semantics and an uncertain prior result. |
These are prompts for reasoning, not a complete architecture prescription. A small offline tool and a high-volume service may require different choices. Explain why a design meets the stated requirements before adding components whose purpose is unclear.
For more basic code-reading exercises, use technical interview practice. For hypothetical teamwork and judgement responses, use situational interview responses. Keep technical evidence and imagined outcomes distinct.
Frequently asked questions
Do these questions represent every SDE interview?
No. They are an original practice set. Check the employer's actual assessment instructions and the role's responsibilities before choosing your preparation priorities.
Are the finite code tests exhaustive?
No. They check the stated sample boundaries. Broader input requirements, adversarial performance and integration behaviour need additional analysis and tests appropriate to the actual system.
Can I claim this exercise as production experience?
Describe it as practice or a personal exercise. Do not relabel a small function or hypothetical design as a deployed service you operated.
