Reviewed: 9 October 2026.
Machine learning interview preparation should include reasoning about data, evaluation and the decision a prediction supports. Naming algorithms is useful background, but a worked answer should explain the quantities being calculated and what the evidence cannot establish.
This guide provides original practice questions using a small regression fixture and a grouped-data split. The predictions are supplied for arithmetic; no model was trained to produce them. The examples are not reported questions from an employer and do not establish real-world model performance.
Question 1: What problem are you trying to predict?
Define the target, the observation unit and the information available at prediction time. For a fictional exercise predicting task duration, the target might be minutes per task, the unit one completed task and the input information only what was available before that task began.
If an input contains a field generated after completion, ask whether it would be available during actual prediction. A model can appear useful in an evaluation that accidentally gives it information unavailable in operation. Do not assume removing one field repairs every possible leakage source; inspect the data collection and splitting process too.
The remaining arithmetic uses duration in minutes. It is a finite calculation, not evidence that a task-duration model should be deployed.
Question 2: Calculate MAE and MSE
The four supplied true durations are [2,4,6,8]. Candidate prediction set A is [3,3,5,9]. Prediction set B is [2,4,6,12]. Values correspond by position, and all four observations have equal weight.
Mean absolute error, MAE, is the average absolute prediction error. Mean squared error, MSE, is the average squared prediction error. The official MAE reference and MSE reference also describe weighting and multiple-output options; this exercise uses neither.
All table values below are minutes.
| No. | True | A | B |
|---|---|---|---|
| 1 | 2 | 3 | 2 |
| 2 | 4 | 3 | 4 |
| 3 | 6 | 5 | 6 |
| 4 | 8 | 9 | 12 |
A's absolute errors are [1,1,1,1], and its squared errors are also [1,1,1,1]. B's absolute errors are [0,0,0,4], while its squared errors are [0,0,0,16].
For A, absolute errors sum to 4 and squared errors sum to 4. Dividing each by four observations gives MAE 1 minute and MSE 1 square minute.
For B, absolute errors also sum to 4, but squared errors sum to 16. B therefore has MAE 1 minute and MSE 4 square minutes. Both sets have the same MAE, while B's one larger error contributes more to the squared-error average.
An answer should state the units. MSE is not measured in minutes. Its square root, RMSE, would be 1 minute for A and 2 minutes for B. Do not call an MSE value an RMSE value.
Question 3: Which prediction set is better?
On this supplied four-observation fixture, A has lower MSE and equal MAE compared with B. That is the complete numerical conclusion. It does not establish which trained model would generalise, whether either set meets an operational requirement or whether the errors occur on important cases.
Ask what the decision requires. If large errors are particularly costly, the error distribution and an appropriate evaluation metric matter. If timing mistakes have asymmetric consequences, a symmetric average may not capture them. Choose the evaluation plan from the intended use rather than changing metrics after seeing which one favours a preferred result.
There is no supplied training history, validation split or representative sample. A and B are prediction sets for an arithmetic exercise; do not describe their difference as an experimentally verified modelling improvement.
Question 4: What changes if B's last prediction becomes 10?
The revised B values are [2,4,6,10]. Absolute errors are [0,0,0,2], summing to 2. MAE becomes 2 divided by 4, or 0.5 minute. Squared errors are [0,0,0,4], so MSE becomes 1 square minute and RMSE 1 minute.
The revised fixture has lower MAE than A and equal MSE. It still does not prove a method caused an improvement; the final prediction was deliberately changed in the supplied example. Describe the altered input and recomputed quantities without inventing a training procedure.
Question 5: How can a row split hide group overlap?
Consider a separate fictional dataset with eight task records from four centres. The centres are group labels, not numerical features:
| Rows | Centre |
|---|---|
| R1, R2 | Cedar |
| R3, R4 | Maple |
| R5, R6 | Pine |
| R7, R8 | Willow |
Suppose the proposed training rows are R1, R3, R5 and R7, while evaluation rows are R2, R4, R6 and R8. Each side contains all four centres. If the intended question is performance on entirely unseen centres, this split does not withhold any centre from training.
A repaired split for that specific aim could train on R1–R4, covering Cedar and Maple, and evaluate on R5–R8, covering Pine and Willow. The two centre sets have no overlap. This checks the intended separation; it does not supply a score or guarantee better performance.
Scikit-learn's cross-validation guidance discusses separate evaluation and group-aware splitting. Apply the split to the generalisation question you actually need to answer. Keeping centres separate does not automatically handle time order, duplicated entities across centres or other dependencies.
Question 6: Where do training, validation and test data fit?
Use training data to fit the model and learn preprocessing. Use a validation plan, potentially cross-validation within the development data, for model and hyperparameter choices. Keep a final test set separate from that selection process.
If you compare many settings on the final test set and choose the best, you have used that set for selection. Do not then report its score as an untouched final evaluation. The appropriate next step depends on the available data and evaluation design; renaming the same rows does not restore their independence from your decisions.
For grouped data, preserve the intended group separation in development folds and the final evaluation where relevant. The eight-row fixture above demonstrates overlap checking only. It does not implement a complete three-way split or a cross-validation run.
The official data leakage guidance also explains why preprocessing must be learned from training data. Learn imputation or scaling parameters within the relevant training portion, then apply them to held-out observations. In cross-validation, fit those transformations separately within each training fold. Learning them from the full dataset uses held-out information.
Question 7: What baseline would you use?
A simple baseline helps interpret a more complicated model, but specify how it is produced. For regression, a development plan might compare against a constant derived from the training targets. The test targets must not determine that constant.
Do not introduce a numerical baseline score for this exercise: no training targets or fitted constant have been supplied separately. State the proposed comparison and what data would be needed to carry it out.
An interviewer can learn more from this boundary than from a claimed score based on an unspecified dataset. If a result is hypothetical, keep it hypothetical when explaining the conclusion.
Question 8: What would you inspect before deployment?
Explain the input contract, missing-value behaviour, evaluation population and downstream response to errors. Ask how predictions are used and whether someone can identify a bad output before acting on it.
Inspect errors across meaningful groups and periods where sufficient data and a suitable process exist. Define who owns monitoring and what action follows a material change. A proposed dashboard is not evidence that drift detection or operational response already works.
The four supplied duration observations cannot answer deployment questions. A finite arithmetic success and a group-disjoint row assignment are useful checks, but neither is an operational acceptance test.
Practise a complete spoken answer
For the first comparison, an original concise answer is: “Both supplied prediction sets have MAE one minute. A has MSE one square minute, while B has four because its last error is four minutes and squares to sixteen. A therefore has lower MSE on these four equal-weight observations. I would need the intended error costs and an appropriate independent evaluation before making a model-selection or deployment decision.”
Then practise the follow-up: “If the aim is unseen-centre performance, I would inspect centre overlap in the split. The first eight-row assignment shares all four centres; the repaired assignment separates two training centres from two evaluation centres. It changes the evaluation design, but no measured score improvement is supplied.”
For broader technical discussion, continue with the system design interview guide. For presenting an actual project honestly, use the STAR method guide.
Frequently asked questions
Do all machine learning roles require the same tools?
Prepare from the role description and assessment instructions. This guide does not establish universal language, framework or deep-learning requirements.
Is a lower metric always a better business outcome?
No. A metric summarises a specified evaluation. Its relevance depends on the use, error costs, data and evaluation design.
Were models trained for these examples?
No. The predictions and group assignments were supplied to make reasoning and arithmetic explicit. Training, fitting and empirical performance are not claimed.
