Data analysis connects a question to a defined dataset, checks the records and communicates a result with its limits. The useful output is not simply a chart or a list of software skills. Another person should be able to identify what was counted, reproduce the calculation and see what the data cannot establish.
This guide completes an original small analysis of invented request records. It includes a data contract, duplicate and missing-value decisions, descriptive statistics, a changed-input check and a concise report. No real company dataset, customer performance or hiring outcome is represented.
Start with the question and unit of observation
The fictional task asks: “What do the supplied records show about the recorded completion durations of closed requests at this snapshot?” One unique request identifier is the unit of observation. Each record gives its status at the same supplied snapshot; these are not sequential events in a request's history.
That distinction changes the analysis. Two events for one request could both be useful in an event log. Here, the task stipulates one snapshot record per unique request, so an identical repeated row should not count as another request.
The durations are supplied minutes associated with completed requests. The exercise does not supply the operational definition used to measure them, a service target or a representative sample from a business. We can summarise the given numbers without claiming they measure actual service performance.
Keep preparing
Continue with Sarkari Resume Templates₹299 — coaching के एक महीने से काफ़ी सस्ता / far cheaper than a month of coachingKeep exploratory questions open
NIST's introduction to exploratory data analysis describes an approach to inspecting data, uncovering structure and checking assumptions, with substantial use of graphical techniques. The worksheet below is an independently written small descriptive exercise. It does not reproduce a NIST dataset or establish a statistical model.
Before selecting a summary, inspect the statuses, missing values and duplicate identifiers. A mean calculated from an unexplained column can be numerically correct while answering the wrong question.
Keep the question narrow enough to match the records. This task concerns five supplied closed-request durations, not every request a business has ever received. There is no before-and-after intervention or randomised comparison from which to infer an improvement.
Read the eight supplied rows
All identifiers and values below are invented. “Absent” means no duration is supplied; it is not a zero-minute completion.
| Row and identifier | Supplied status and duration |
|---|---|
| 1 — R1 | Closed; 10 minutes |
| 2 — R2 | Closed; 20 minutes |
| 3 — R3 | Closed; 40 minutes |
| 4 — R4 | Pending; duration absent |
| 5 — R5 | Closed; 30 minutes |
| 6 — R6 | Cancelled; duration absent |
| 7 — R7 | Closed; 50 minutes |
| 8 — R1 repeated | Closed; 10 minutes |
The last row exactly repeats R1's identifier, status and duration. It is not a later update in this stipulated snapshot. Preserve the original eight-row input and record the decision to retain R1 once.
If a real dataset repeated an identifier with conflicting values, the same simple decision would not resolve it. You would need an applicable rule or clarification. Do not arbitrarily retain the first conflicting record and report the discrepancy as fixed.
Create an explicit analysis record
After removing the one identical repeated row for this task, there are seven unique requests: five closed, one pending and one cancelled. Two unique requests have no supplied duration. Both are outside the requested closed-duration summary.
Retain them in the request-count report. Excluding them from a particular statistic is not permission to delete them from the dataset or pretend every request closed. The denominator for a closed-duration summary is five; the denominator for unique-request status counts is seven.
An original analysis note can state:
The source contains eight rows, including one identical repeat of R1. The snapshot analysis retains seven unique requests. Completion-duration summaries use the five closed requests with supplied minute values. Pending R4 and cancelled R6 remain in the status counts and have no imputed duration.
This note records the rule and its consequences. It does not establish that missing fields in every dataset should be handled the same way.
Calculate the descriptive summaries
The five relevant durations are 10, 20, 40, 30 and 50 minutes. Their sum is 150 minutes. The mean is 150 ÷ 5 = 30 minutes.
Sort them as 10, 20, 30, 40, 50. The middle value, and therefore the median for these five observations, is 30 minutes. The minimum is ten, the maximum is fifty and the range is 50 − 10 = 40 minutes.
| Summary | Supplied-data result |
|---|---|
| Closed durations included | 5 |
| Sum | 150 minutes |
| Mean | 30 minutes |
| Median | 30 minutes |
| Minimum / maximum | 10 / 50 minutes |
| Range | 40 minutes |
These results were independently calculated from the supplied synthetic rows. They do not establish a typical duration for a wider population, an acceptable service level or a causal effect of staffing or software.
The equal mean and median do not prove a particular distribution. Five deliberately supplied values are insufficient evidence for a broad claim about how operational durations behave.
Show how two cleaning errors change the answer
If the repeated R1 were incorrectly counted as another closed request, the six included durations would sum to 160 minutes. The resulting mean would be 160 ÷ 6, approximately 26.67 minutes. That answers a six-row question containing a duplicate, rather than the intended five-unique-request question.
If the two absent values were instead replaced with zero after deduplication, the sum would remain 150 but the denominator would become seven. The mean would be 150 ÷ 7, approximately 21.43 minutes. That would silently treat pending and cancelled requests as zero-minute completions.
Both errors produce lower averages, but neither establishes a faster process. The exercise demonstrates why a changed statistic needs a traceable data rule rather than a favourable interpretation.
Do not choose a cleaning rule because it makes the result look better. Choose the applicable rule for the question, document it and retain excluded records or unresolved conflicts for review.
Check a changed input without rewriting the original
For a separate variation, the fixture changes R7's supplied duration from fifty to 80 minutes. It changes no identifier or status. The five durations then sum to 180 minutes, giving a mean of 36 minutes.
Sorted values become 10, 20, 30, 40, 80. The median remains 30 minutes, while the range becomes 80 − 10 = 70 minutes. The different response of the mean and median is visible in the arithmetic; no new cause for the larger value is supplied.
Keep the original and changed-input results labelled separately. Do not call the variation a later business month, an observed delay or an outlier removed from a real dataset. It is a supplied teaching change.
A later investigation could ask whether an unusual value reflects measurement, task differences or a data error. Those are questions, not answers established by the changed fixture.
Choose a display that matches the question
For this small exercise, the complete input table and summary table make the calculation inspectable. A plot could display the five durations, but a decorative chart would not resolve the unknown measurement definition or sampling scope.
If you create a display, label the unit as minutes and identify the included observations. Avoid a headline such as “Team achieves a 30-minute completion standard”: no standard, team or achieved target is supplied.
A relevant headline is “Five supplied closed requests have a mean and median of 30 minutes.” It names the scope and statistic. Keep the status count nearby so the reader does not infer that all seven unique requests completed.
Write the result and its limitations together
An original short report can read:
After removing one identical repeated row, the supplied snapshot contains seven unique requests: five closed, one pending and one cancelled. The five closed-request durations total 150 minutes, with a mean and median of 30 minutes and a range of 40 minutes. Pending and cancelled requests have no supplied duration and were not treated as zero-minute completions. The exercise does not provide a service target, sampling basis, duration-measurement definition or evidence about causes.
The report is useful because another person can reproduce it from the table. It does not require a claim of business improvement, a forecast or an advanced model.
If a decision depends on the missing measurement definition or a representative period, request that evidence before extending the conclusion. A descriptive result can remain valid within its scope while being insufficient for a wider decision.
Build an honest learning artifact
Retain the original rows, data contract, cleaning note, calculations, changed-input check and report. Name software only if you actually used it and can explain the relevant operations. A manually checked small example does not establish proficiency with every analytics platform.
For a portfolio, label the data as synthetic and the output as independent practice. You can describe identifying a duplicate and preserving missing values. You cannot describe improved customer service, a deployed dashboard or employer-approved analysis unless those events actually occurred.
The next useful learning task is one whose question and evidence are clear enough to check. A transparent small analysis gives you a stronger explanation than an unsupported list of tools or a fabricated percentage improvement.
