KarmSakha
Help

Interview & Soft Skills

ML Interview Questions: Dataset, Metrics and Checked Answers

Reviewed: 9 October 2026

Machine learning interview questions are useful practice when you can explain an answer, work a small example and state its limits. A memorised definition will not help much if you confuse predicted labels with actual labels, choose a threshold using the final test set, or report a score without explaining its denominator.

This guide contains original questions and a complete numerical exercise. The records and scores are invented teaching data, not an employer's interview bank, a trained production model or evidence that a particular approach is ready to deploy. Confirm the actual role's assessment topics and permitted tools before an interview.

Question 1: What do training, validation and test data do?

Training data is used to learn model parameters and any learned preprocessing. Validation data, or suitable cross-validation within the development data, helps compare choices. A held-out test set is reserved for evaluating the final chosen procedure. Repeatedly adjusting the procedure to improve its test score makes that set part of development rather than an untouched final evaluation.

For example, suppose you compare thresholds using a development set and choose one before evaluating new held-out records. That is different from looking at the test labels, selecting whichever threshold scores best there and presenting the result as an independent estimate. Keep a record of which data informed each choice. The scikit-learn evaluation guide explains this separation and the risk of tuning to the test set.

Question 2: Where can preprocessing cause leakage?

If a transformation learns a mean from data, learn that mean on training data and apply the same learned transformation to later data. The official common-pitfalls guidance distinguishes fitting from applying transformations and describes pipelines as a way to keep those steps together.

Here is an original centring example. Training values are 2, 4 and 6, whose mean is 4. A held-out value is 20. Subtracting the training mean gives training values −2, 0 and 2, and a held-out centred value of 16. If you instead use all four values, the mean becomes 8, and the held-out centred value becomes 12. Its value has influenced the learned transformation.

This arithmetic illustrates contamination of the mean. It does not show how much any model's accuracy would change; no model is trained in this example. A pipeline also cannot repair a dataset split that puts related observations on both sides when the evaluation requires independent groups.

Question 3: How do you read a confusion matrix?

Specify the positive class and the axis order. In the convention used by scikit-learn's confusion matrix, rows represent actual classes and columns represent predicted classes. With class order 0 then 1, the first row contains true negatives and false positives; the second contains false negatives and true positives.

Reversing an axis changes what a cell means. Saying “three mistakes” without identifying their type loses information that may matter to the decision. The following original exercise makes each count visible.

The complete ten-record classification exercise

A fictional system supplies a score for each record. Actual class 1 is the positive class; 0 is negative. Predict 1 when the score is at least 0.50, otherwise predict 0. Scores are supplied numbers for practice; they are not established calibrated probabilities. No fitting data or model is claimed.

RecordActual classSupplied scorePrediction at 0.50
R110.901
R210.450
R300.701
R400.200
R510.801
R600.551
R710.300
R800.100
R900.601
R1000.050

There are four actual positives: R1, R2, R5 and R7. The system predicts five positives: R1, R3, R5, R6 and R9. Two predictions are true positives, R1 and R5. Three are false positives, R3, R6 and R9. Two actual positives are missed, R2 and R7. The remaining three records are true negatives, R4, R8 and R10.

The counts sum to ten: TP 2 + FP 3 + FN 2 + TN 3 = 10. In the stated matrix convention, the result is:

Actual / predictedPredicted 0Predicted 1
Actual 033
Actual 122

Question 4: Calculate accuracy, precision, recall and F1

For these records, accuracy is correctly classified records divided by all records: (2 + 3) / 10 = 0.50.

Positive-class precision asks how many predicted positives are correct: 2 / (2 + 3) = 0.40. Recall asks how many actual positives are found: 2 / (2 + 2) = 0.50. F1 is 2TP / (2TP + FP + FN) = 4 / 9, approximately 0.4444. The official precision, recall and F-score API documents these quantities and undefined-denominator handling.

Name the positive class and averaging convention when reporting metrics. The numbers here are for class 1, with no sample weights. They do not establish performance in a new population or on another dataset.

Question 5: Does a higher accuracy always mean a better choice?

A rule predicting 0 for every supplied record gets six of ten correct: 0.60 accuracy. It finds none of the four positives, so recall is 0. Precision has no mathematical value here because there are no predicted positives; its denominator is zero. F1 by the count formula is 0, because the four false negatives keep that denominator positive.

Whether this rule is useful depends on the actual objective and consequences of errors. Its greater accuracy than the 0.50-threshold rule does not show that it meets a need to identify positives. Conversely, you cannot declare the other rule preferable without knowing that need. A toy score table is insufficient for a deployment decision.

When using a library, state its zero-division setting. A returned numeric zero for undefined precision is a convention, not proof that the missing denominator exists. In the code below, None explicitly marks an undefined ratio.

Run an original calculation

This Python code uses the exact supplied records. It needs no ML library and does not train a model.

truth = [1,1,0,0,1,0,1,0,0,0]
scores = [.9,.45,.7,.2,.8,
          .55,.3,.1,.6,.05]

def ratio(a, b):
    return a / b if b else None

def evaluate(pred):
    tp = fp = fn = tn = 0
    for actual, guess in zip(truth, pred):
        if actual == 1:
            if guess == 1:
                tp += 1
            else:
                fn += 1
        elif guess == 1:
            fp += 1
        else:
            tn += 1
    return {
        'counts': [tn, fp, fn, tp],
        'accuracy': ratio(tp+tn, 10),
        'precision': ratio(tp, tp+fp),
        'recall': ratio(tp, tp+fn),
        'f1': ratio(2*tp, 2*tp+fp+fn)
    }

pred = [int(s >= .5) for s in scores]
print(evaluate(pred))

The count order in the output is [TN, FP, FN, TP], giving [3, 3, 2, 2]. Accuracy is 0.5, precision 0.4, recall 0.5 and F1 approximately 0.4444. This is a calculation check, not an evaluation of a fitted classifier.

Question 6: What changes at a different threshold?

For a separate arithmetic exercise, predict 1 at scores at least 0.70. That selects R1, R3 and R5, including R3 exactly on the boundary. Counts become TN 5, FP 1, FN 2, TP 2. Accuracy is 0.70, precision 2/3, recall 0.50, and F1 4/7, approximately 0.5714.

At a threshold of 0.80, R1 and R5 are positive. Counts become TN 6, FP 0, FN 2, TP 2. Accuracy is 0.80, precision 1, recall 0.50, and F1 2/3, approximately 0.6667. These are comparisons on fixed invented records. They are not a legitimate way to select a threshold on a final held-out test set and then claim unbiased performance.

Practise a complete explanation

Try this answer: “I would first confirm the positive class and threshold rule. At 0.50, the supplied ten records give two true positives, three false positives, two false negatives and three true negatives. Positive precision is 0.40 and recall is 0.50. The all-negative rule has higher accuracy but finds no positives. I cannot choose a deployment rule without the objective, error consequences and appropriate evaluation evidence.”

Then ask a partner to change one score and identify which count changes. If R2 moves from 0.45 to exactly 0.50, it becomes a true positive under the original boundary rule. TP becomes 3, FN becomes 1, while FP and TN remain 3 each. Accuracy is 0.60, precision 0.50, recall 0.75 and F1 0.60.

Questions people ask

Are these actual employer questions?

No. These are original teaching questions around documented concepts and an invented dataset. They do not predict a company's question bank or hiring outcome.

Can I call the scores probabilities?

The supplied exercise does not establish probability calibration. Explain them as scores and state the threshold. Do not add a probability interpretation merely because values lie between zero and one.

What should I do if a denominator is zero?

Identify why it is zero and state your chosen reporting convention. Do not quietly invent a precision value or drop records to avoid the issue.

For related preparation, use SQL interview practice and the worked case interview exercise. Their fixtures are separate from this ML exercise; prepare the topics required by your actual role.

Related guides

Ask KarmSakha AI