Editorially revised on 9 October 2026.
Prepare explanations you can check
A useful data-science interview answer connects a question to the data, assumptions, calculation and limits of the conclusion. Naming a library or memorising a definition is only part of that work. You should also be able to explain what the result measures and what additional information is needed before using it.
The exercises below are original and fictional. They are practice questions, not a prediction of every employer's interview process or a claim about actual business performance. No dataset was drawn from a real candidate, patient or customer. The numerical examples are intentionally small so you can independently check them.
Prepare against the advertised role. A role involving reporting might emphasise data checks and communication; another might require model development or software implementation. Those are possible differences, not mandatory round structures. Ask the recruiter about the actual assessment format when instructions are unclear.
Question one: what would you check before modelling?
Suppose a fictional table contains an item ID, creation timestamp, category and an outcome label. Before selecting a model, explain what each row represents. Is it one item, one event or a repeated observation? Ask when each field became available, how the outcome was defined and whether the intended prediction time is before that outcome.
Then examine missing fields, duplicate records, inconsistent categories and impossible values in context. Two rows sharing an item ID might be a duplicate or two legitimate events. Removing one without understanding the row definition can discard useful information. A missing category might mean unknown, not applicable or a collection failure; those meanings can require different treatment.
Practice answer: I would first confirm the unit of observation and prediction time. I would inspect the label definition, field availability and repeated IDs, then check missingness and value validity. I would record unresolved data questions before choosing a cleaning rule. I would not apply a blanket deletion or filling rule solely because a function makes it easy.
Follow-up: Does finding no missing values prove the dataset is ready? No. The labels may still be inconsistent, the sampling may not match the intended use and a feature may reveal information unavailable at prediction time. A data check establishes what was checked, not that every possible defect has been excluded.
Question two: read a confusion matrix with explicit labels
Consider a fictional classifier marking items for manual review. Label 1 means requires review; label 0 means does not require review. The held-out exercise contains 100 items, of which 12 actually require review. The classifier marks 10 items positive: 8 are genuinely positive and 2 are false positives. Four positive items are missed, and 86 negative items are correctly left unflagged.
Declare the axes before calculating. With actual labels as rows and predicted labels as columns, ordered as 0 then 1, the matrix is:
| Actual label | Predicted 0 | Predicted 1 | Total |
|---|---|---|---|
| Actual 0 | 86 true negatives | 2 false positives | 88 |
| Actual 1 | 4 false negatives | 8 true positives | 12 |
| Total | 90 | 10 | 100 |
Accuracy is the correctly classified count divided by the total: (86 + 8) / 100 = 94%. Positive-class precision is 8 / (8 + 2) = 80%. Positive-class recall is 8 / (8 + 4) = 2/3, approximately 66.7%. These values answer different questions, so an interviewer should be able to see each denominator.
The scikit-learn documentation defines the confusion-matrix axes and the positive-class precision and recall formulas. Check label ordering and averaging options when using a library; this exercise uses a single explicitly identified positive class. Confusion matrix, precision, recall.
Follow-up: Is 94% accuracy enough to choose this model? That number alone does not establish suitability. Ask about the consequences of missed reviews and unnecessary reviews, the evaluation population and the available alternatives. The exercise supplies counts, not an approved operating threshold or a cost-benefit decision.
Question three: compare the all-negative baseline
On the same fictional 100-item set, a classifier predicting 0 for every item has 88 correct negatives and misses all 12 positives. Its accuracy is 88%. Its positive-class recall is 0 / 12 = 0. It never predicts a positive, so the precision denominator is zero: 0 / 0 is undefined as a ratio.
That last distinction matters. A library can apply a documented zero-division convention and return a chosen value or warning. That convention does not make the underlying ratio mathematically defined. In an interview, state the counts, identify the undefined denominator and explain which library policy you used if reporting a computed score.
Practice answer: The model improves accuracy from the all-negative baseline's 88% to 94% on this exercise and finds eight of twelve positives. I would still review the error trade-off and evaluation design. I would report the baseline's positive precision as undefined, or clearly label the library's zero-division convention, because it makes no positive predictions.
Follow-up: What if the evaluation set contains no actual positives? Positive recall then also has a zero denominator. The precision and recall documentation explains handling options. Do not silently treat every empty-denominator case as evidence of good or poor real-world performance.
Question four: show how preprocessing can leak information
Use another fictional example: the training feature values are 0 and 10, while a held-out test value is 100. A mean estimated from the training values is 5. A mean estimated from all three values is 110 / 3, approximately 36.67. The second estimate uses the held-out value and therefore changes the transformation using information outside the training set.
The practical issue is the source of the learned preprocessing parameter, not whether the arithmetic is difficult. Fit learned preprocessing on training data and apply the fitted transformation to held-out data. During cross-validation, preprocessing needs to be fitted within the relevant training fold. Scikit-learn's common-pitfalls guidance explains this principle and how pipelines help organise it. Preprocessing and leakage guidance.
Practice answer: I would split according to the evaluation design, fit the learned transformation on the training portion and apply it without refitting on the held-out portion. Within cross-validation I would keep that fitting inside the fold's training workflow. I would also check whether features exist at prediction time.
Follow-up: Does using a pipeline eliminate all leakage? No. A pipeline can organise transformations, but it cannot by itself correct a feature containing future outcome information or decide whether related rows should stay together. The feature definition and split design still require review.
Question five: choose a split for the actual question
A random row split is an option, not a universal default. If the intended use concerns future observations, investigate whether the evaluation should preserve time order. If multiple records belong to the same item or person, investigate whether putting related records on both sides would produce an unrealistic estimate. Describe the structure before choosing the split.
This guide does not prescribe one splitter for every dataset. A defensible answer states the intended use, the relationship between rows and the information available at the time of prediction. If those facts are missing, identify the missing facts rather than claiming a particular evaluation is valid.
Practice answer: I would decide the split after confirming whether the model will be used on future events, new entities or additional records from known entities. I would check repeated entities and timestamps. Then I would explain what the evaluation includes and excludes, instead of relying on a score without its sampling context.
Present the answer as a short review record
For each exercise, write five lines: the question, the data definition, the method, the result and the limitation. Ask a practice partner to reconstruct the calculation from your record. If the partner cannot identify the positive class or denominator, clarify those points before adding more terminology.
For practical task rehearsal, see the technical interview preparation guide. Its validation task is a separate software exercise; it does not substitute for evaluating a statistical model. Keep these different kinds of evidence distinct when describing your preparation.
Common questions
Should I memorise all library defaults?
Know the tools relevant to the advertised role, but verify exact options against the version in use. An explicit statement of labels, averaging and undefined-score handling is often more useful than assuming a default. These examples do not establish the installed version in an employer's environment.
Can a high score prove the model is useful?
A score describes a calculation under an evaluation design. Suitability also depends on the data, error consequences and intended use. Explain the limits of the supplied evidence instead of turning a practice result into a deployment recommendation.
What if I cannot finish the calculation?
State the quantities you know and the step you need to check. Correct labels and an honest limitation are better foundations for discussion than an invented number. After practice, independently recompute the example and repair the specific gap.
