KarmSakha
Help

Career Advice

A/B Testing: Assignment, Rates and a Guardrail Decision

An A/B test compares defined alternatives under an experiment plan. A useful analysis needs more than two percentages: it needs a clear assignment unit, consistent measurement, an outcome window and a decision rule. A larger observed rate is a descriptive result; it does not by itself establish statistical significance or a safe release.

This guide builds an original fictional registration experiment, calculates its supplied results and applies an invented practical error rule. No website, user assignment or statistical test was executed for this article. The counts are teaching inputs, not a result achieved on KarmSakha or another business.

Define the question and the two alternatives

The exercise concerns a fictional community workshop page. Version A retains a short invitation sentence. Version B replaces that sentence with a clearer description of what the workshop covers. The form, schedule, availability and confirmation logic remain the same in the proposed design.

The question is whether B's wording changes the specified registration outcome for the eligible population. The exercise provides no real price, workshop, live page or publishing authority. The planned change has not been approved or launched.

NIST's completely randomised design introduction describes random assignment of a factor's levels to experimental units. The user-level protocol below is independently constructed; the source does not verify an online allocation system or supply this example.

Keep preparing

Continue with Sarkari Resume Templates₹299 — coaching के एक महीने से काफ़ी सस्ता / far cheaper than a month of coaching
₹299
Buy this eBook

Write the protocol before reading results

The proposed protocol uses newly arriving eligible users during a supplied fourteen-day enrollment window. Each receives A or B with a probability of one half, and the same assignment is retained across their repeat visits. Known automated test accounts are excluded under a rule declared before assignment.

The assignment unit is a unique eligible user, identified through the proposed stable identifier. It is not a page view, session or button click. The exercise supplies no implementation of the identifier or allocation routine, so both need verification before any real experiment.

Protocol fieldOriginal exercise definition
Assignment unitUnique eligible user
AlternativesA or B invitation sentence
EnrollmentFixed fourteen-day window
Outcome windowSeven days after first assignment
Primary denominatorAll eligible assigned users in the arm
Analysis timingAfter the last user's outcome window matures

With enrollment ending fourteen days after the starting instant and a seven-day window for each user, analysis waits until 21 days after that starting instant. These durations are fictional design inputs, not a recommendation about sufficient sample size or power.

NIST's experiment-planning discussion includes objectives, design, measurement, execution and interpretation. It supports planning the analysis deliberately; it does not prescribe this guide's durations or guarantee a particular outcome.

Specify success and retain the denominator

The primary metric is the number of unique assigned users with at least one valid completed registration within their seven-day window, divided by all eligible assigned users in that arm. Each user contributes at most one success.

A user who abandons the form or encounters an error remains in the denominator under this contract. Do not calculate the primary rate only among people who submitted something, since that would answer another question.

The fixture supplies matured aggregate records of 600 assigned users in A and 600 in B. These equal counts do not prove that randomisation, exposure or instrumentation worked. They are inputs to the arithmetic, not an audit of the proposed system.

A real investigation would need the corresponding assignment and measurement evidence. If identifiers or outcome records conflict, the aggregate percentages cannot explain those conflicts away.

Calculate the supplied registration results

A has 48 converting users out of 600. B has 63 out of 600. Their rates are:

  • A: 48 ÷ 600 = 8%.
  • B: 63 ÷ 600 = 10.5%.
  • B minus A: 2.5 percentage points.
  • Relative difference against A: (10.5% − 8%) ÷ 8% = 31.25%.

B's record contains fifteen more converting users. A has 552 nonconverting users, and B has 537. Those counts complete each arm's supplied total of 600.

The absolute and relative changes use different denominators. A 2.5-point difference is not a 2.5% relative increase. The relative formula would also be undefined if the reference rate were zero; do not divide by zero or silently substitute an absolute difference.

These calculations were independently executed on the fictional inputs. They do not establish profit, retention, adequate sample size or a conversion increase on a real service.

Check a denominator counterexample

For a separate arithmetic example, consider 70 conversions among 1,000 assigned users. Its rate is 7%, lower than A's 8% despite its larger raw conversion count.

This is not a third arm or another supplied experiment result. It demonstrates why a count needs its denominator before comparison. Keep the counterexample separate from the A/B result table.

Likewise, repeat visits should not automatically increase the unique-user denominator under the original protocol. A session-based question would require its own definition and analysis. It cannot be introduced after seeing the counts just because it produces a more attractive rate.

Apply the original error guardrail

The fixture defines an error measure as assigned users with at least one submission error, divided by all assigned users in that arm. It supplies 12 of 600 in A and 24 of 600 in B: 2% and 4%, an observed worsening of two percentage points.

A user can encounter an error and later complete registration. Error and conversion counts may therefore overlap; do not add them as if they were mutually exclusive categories.

Before seeing the records, the exercise declares this practical rule: do not approve expansion when the observed error-rate increase exceeds 0.5 percentage point; investigate the errors first. The supplied two-point increase exceeds that margin.

Original decision itemSupplied result
Registration rateB higher: 10.5% versus 8%
Error rateB higher: 4% versus 2%
Invented error marginAt most 0.5 point increase
Expansion decision under ruleWithhold approval; investigate

The margin is an invented teaching condition, not an industry standard or a statistical noninferiority test. Applying it does not prove B caused the errors. It demonstrates that the attractive primary metric does not satisfy every declared requirement.

Keep monitoring and interpretation within the plan

A Microsoft-hosted primary experimentation paper discusses outcome, guardrail and data-quality metrics, along with misleading denominators and early stopping. For standard fixed-horizon testing, repeatedly changing the stopping time while seeking significance can increase false positives. Safety monitoring has a different purpose. This guide implements no sequential inference method or company experiment from that paper.

In the original plan, retain the declared enrollment and outcome windows. If an authorised owner pauses a real test because of a verified blocking problem, record that event and its effect on interpretation. The fixture establishes no pause, owner permission or live release.

Do not call the sample sufficient simply because each arm contains 600 users or the plan spans fourteen days. No power calculation was performed. The practical error rule likewise should not be reported as a statistical test merely because it has a numerical threshold.

Do not invent the missing inference

NIST's comparison-of-proportions section describes statistical methods with relevant sample and independence assumptions. Calculating the two rates is not the same as executing such a method.

No confidence interval, p-value or significance test was calculated for this exercise. Statistical significance is therefore unestablished, rather than positive or negative. The source does not certify the fixture's assignment integrity or tell us that B is a winner.

Random assignment in a proposed protocol also differs from random sampling of a wider population. The fixture does not establish that the eligible users represent every future user or that repeated users, shared experiences and measurement gaps have been handled in a real implementation.

Produce a bounded result note

An original report can read:

In the supplied fictional matured records, A has 48 registrations among 600 assigned users, or 8%; B has 63 among 600, or 10.5%. The observed difference is 2.5 percentage points, or 31.25% relative to A. B's supplied error rate is 4% compared with A's 2%, exceeding the invented 0.5-point practical margin. Under that declared rule, expansion approval is withheld pending investigation. Assignment integrity, statistical significance, causal effects and real-world release outcomes have not been established.

The complete learning packet is the protocol, input counts, calculations, error decision and unresolved-evidence record. It provides something concrete to inspect without turning fictional arithmetic into a measured business improvement.

Related guides

Ask KarmSakha AI