LearnPath
DiscoverFindPricingBlog
Sign inGet started
Home›Blog›A/B Testing for Product Managers: The 2026 One-Week Plan

Learning guides13 min readUpdated October 9, 2026

A/B Testing for Product Managers: The 2026 One-Week Plan

A 2026 five-day plan with StatQuest, Emma Ding and Lenny's Podcast, so you can read an A/B test readout, question its p-value and size your next experiment.

By LearnPath Team

Example path

Your learning path

Pick a topic. We'll turn YouTube into a course.

Tell us your level. We pick the best lessons, put them in order and check you understood each one before the next.

Building your path is free. Ready in about a minute.

Quick Answer: what a product manager needs from A/B testing statistics before a launch call

An A/B test is a randomized experiment that splits users between the current version (A) and a changed version (B), then compares one metric chosen in advance, so a difference bigger than chance can be credited to the change. To read, question and plan one, a product manager needs five ideas: the p-value, the confidence interval, power and sample size, the ways tests lie, and how a full readout fits together. This plan covers them in about five hours over five days, one hour a day, from free YouTube videos, finishing in time for your next experiment review.

Why it matters at a launch decision: most ideas do not work. In "The Surprising Power of Online Experiments" (Harvard Business Review, September-October 2017), Ron Kohavi and Stefan Thomke report that at Google and Bing only about 10% to 20% of experiments generate positive results, and that at Microsoft as a whole one-third of ideas prove effective, one-third neutral and one-third negative. If you approve launches on gut feel, you are approving a lot of neutral and negative changes.

The same article shows the upside: a Bing change to how ad headlines were displayed increased revenue by 12%, which on an annual basis would come to more than $100 million in the United States alone, without hurting key user-experience metrics. Source: Kohavi and Thomke, HBR 2017.

Do this before you watch anything: get the last readout

Before Day 1, find the most recent experiment readout your team produced and write down the one metric it was judged on. Every day of this plan ends with a question you ask of that document. Without it you are learning statistics in the abstract, and the week will not stick to anything you actually decide.

A slide or a three-line message counts. Note the primary metric, the result for A and B, any p-value or confidence interval, how long it ran, how many users landed in each group, and what the team decided. If you start on Monday 12 October 2026, Day 5 lands on Friday 16 October, which is enough for most review cycles. If your review is sooner, do Days 1, 3 and 4 first.

Day 1: what a test result actually means

By the end of today you can say in one sentence what the null hypothesis is for your last test, and what its p-value does and does not prove. This is the foundation. Every argument in an experiment review is, underneath, an argument about what the p-value means, and it is easy to get slightly wrong.

The idea: an A/B test starts by assuming the change did nothing. That assumption is the null hypothesis. The test then asks how surprising the observed difference would be if that were true. The p-value is that surprise, as a number. A small p-value means the data would be rare under "no effect", so you reject "no effect".

What it is not: the chance that B is better, or a measure of how big the effect is. Those are the misreadings you will hear most often in a review.

Watch: "Hypothesis Testing and The Null Hypothesis, Clearly Explained!!!", StatQuest with Josh Starmer, 14:41, https://www.youtube.com/watch?v=0oc49DyA3hU. Why it is enough: it builds the null hypothesis from a plain example before any formula, which is the order a product manager needs.

Watch: "p-values: What they are and how to interpret them", StatQuest with Josh Starmer, 11:21, https://www.youtube.com/watch?v=vemZtEM63GY. Why it is enough: the title is the whole job, defining a p-value and how to interpret one, from the same teacher so the vocabulary carries over.

On your own readout today: write the null hypothesis for your last test in one sentence ("the new checkout changed nothing about conversion"). Then write what its p-value lets you say, and one thing it does not let you say.

Day 2: how big the effect is, not just whether there is one

By the end of today you read the confidence interval before the p-value. The p-value says whether there is likely an effect. The confidence interval says how big it plausibly is, and that is the number a launch decision actually turns on. A real but tiny lift can still be a bad trade for the work it costs.

A confidence interval is a range of effect sizes consistent with the data. If your readout reports a lift whose interval runs from barely above zero to something large, the honest summary is "probably positive, size unclear". If the interval crosses zero, the test has not ruled out "no effect" at all, whatever the headline says.

Watch: "Confidence Intervals, Clearly Explained!!!", StatQuest with Josh Starmer, 6:42, https://www.youtube.com/watch?v=TqOeMYtOc1w. Why it is enough: short, visual, and focused on what the interval means rather than how to compute it by hand.

Watch: "Hypothesis test for difference in proportions example | AP Statistics", Khan Academy, 10:28, https://www.youtube.com/watch?v=HOyNEU94xJQ. Why it is enough: a difference in proportions is exactly what a conversion test is, and one worked example shows every number your readout reports.

On your own readout today: find the confidence interval for the primary metric. If there is none, ask for it. Write the smallest and largest plausible effect in plain words, and say whether the smallest one would still justify shipping.

Day 3: power and sample size, or how long a test must run

By the end of today you can estimate how many users a test needs before it starts, and so how long it must run. This is where product managers add the most value, because the decision about test length is usually made under pressure to ship, and it decides whether the result is worth reading at all.

Power is the chance a test detects an effect that really exists. A test with too few users is underpowered: it will often miss real improvements and call them "no effect", which quietly kills good ideas. The usual target is 80% power with a 5% significance level.

There is a rule of thumb for sample size per variant: n ≈ 16σ²/δ², for a two-sided test at α = 0.05 and 80% power, where δ is the smallest difference you want to detect. A 2023 arXiv paper on sample-size calculations for A/B testing describes it as widely adopted in academia and industry. For a conversion rate p, the variance σ² is p(1 - p).

A worked example (the arithmetic is ours): your baseline conversion is 5% and you want to detect an absolute lift of 0.5 points, from 5.0% to 5.5%. Then σ² = 0.05 × 0.95 = 0.0475, and n ≈ 16 × 0.0475 / 0.005² = 30,400 users per variant. Divide that by your daily eligible traffic and you have a minimum run time. Because δ is squared, halving the lift you want to detect quadruples the sample size.

Watch: "Statistical Power, Clearly Explained!!!", StatQuest with Josh Starmer, 8:19, https://www.youtube.com/watch?v=Rsc5znwR5FA. Why it is enough: it explains what power is before asking you to calculate anything.

Watch: "Power Analysis, Clearly Explained!!!", StatQuest with Josh Starmer, 16:45, https://www.youtube.com/watch?v=VX_M3tIyiYk. Why it is enough: it turns power into the planning step that sets sample size, which is the decision you own.

Watch: "Sample Size Estimation in A/B Testing: Easy Explanation for Data Science Interviews", Emma Ding, 7:38, https://www.youtube.com/watch?v=FKPec6RoJOg. Why it is enough: it puts the same idea in A/B testing terms, in under eight minutes, the way a data scientist would explain it to you.

On your own readout today: take the baseline rate and the lift your last test reported, run the rule of thumb, and compare the answer with how many users the test actually had. If it had far fewer, the result was probably underpowered.

Day 4: the ways a test lies

By the end of today you can name the three failures that make a clean-looking readout wrong: peeking, p-hacking and a broken split. Each one produces a confident result that is not real. Catching them is the most useful thing a product manager can do in a review, because none of them shows up in the headline number.

Peeking. Checking the results every day and stopping as soon as they look significant gives chance many tries to cross the line. A change that does nothing will eventually look like a winner if you look often enough. The fix is to fix the end date in advance, or to use a testing method built for continuous monitoring.

P-hacking. Testing many metrics, segments or variants and reporting the one that came out significant. With enough comparisons, something always does. Ask how many metrics were checked and whether the winning one was the primary metric chosen before launch.

Sample ratio mismatch (SRM). If you planned a 50/50 split and the groups came out noticeably uneven, the assignment or the logging is broken, and the result cannot be trusted at all. It is a cheap check and it should come before anyone reads the metric. Kohavi and Thomke's 2017 HBR article cites Twyman's law, which fits this moment well: "Any figure that looks interesting or different is usually wrong."

Watch: "p-hacking: What it is and how to avoid it!", StatQuest with Josh Starmer, 13:45, https://www.youtube.com/watch?v=HDCOUXE3HMM. Why it is enough: it covers both what p-hacking is and how to avoid it, in the vocabulary from Day 1.

Watch: "Peeking at A/B Tests: Why it matters, and what to do about it", KDD2017 video, 20:32, https://www.youtube.com/watch?v=6G8F22NSCxk. Why it is enough: it is the conference talk for a KDD 2017 paper on exactly this problem, covering both why peeking matters and what to do instead.

Watch: "TLC: Sample Ratio Mismatch (SRM) with Lukas Vermeer", TLC - Test & Learn Community, 54:30, https://www.youtube.com/watch?v=CJ9KinpJplg. Why it is enough: a whole session on one check every readout should pass. It is close to an hour, so watch at speed and stop once you can explain how SRM is checked.

On your own readout today: check three things. Did the test stop on its planned date? Was the reported metric the one chosen before launch? Are the user counts in A and B close to the planned split? Write down anything you cannot confirm.

Day 5: read a real readout end to end

By the end of today you can walk through a readout in the order that catches problems: the split, the planned sample size, the primary metric's confidence interval, then the p-value, then the decision. Days 1 to 4 gave you the pieces. Today puts them in one sequence you can use in any review.

The order matters. The split comes first because an SRM makes everything after it meaningless. Sample size second, because an underpowered test cannot tell you "no effect" with confidence. The interval before the p-value, because the size of the effect is what pays for the launch. The podcast episode is long: skim it or watch at speed.

Watch: "A/B Testing Made Easy: Real-Life Example and Step-by-Step Walkthrough for Data Scientists!", Emma Ding, 13:39, https://www.youtube.com/watch?v=VuKIN9S8Ivs. Why it is enough: one real-life example from design to decision, which is the shape of the readout you are about to review.

Watch, selectively: "The ultimate guide to A/B testing | Ronny Kohavi (Airbnb, Microsoft, Amazon)", Lenny's Podcast, 1:23:08, https://www.youtube.com/watch?v=hEzpiDuYFoE. Why it is enough: a product podcast, so it is framed around the decisions product managers make, not the formulas. Pick the parts that match a question your own readout raised.

On your own readout today: go through it in the order above and write five lines, one per step. End with one sentence: "Based on this test, I would ship, kill, or rerun, because..." That sentence is what you bring to the review.

How the options compare

There is more than one realistic way to get this skill before a review. A video plan is fastest, a book goes deepest, your platform's documentation is closest to your tool, and a data scientist is closest to your company. LearnPath adds order and quizzes. Pick by how much time you have before the decision.

OptionCostTimeGood forWeak spot
This five-day YouTube planFree videosAbout five hoursReading and questioning a readout before a set dateYou pick and pace it yourself, and nothing checks you understood
"Trustworthy Online Controlled Experiments" (Kohavi, Tang and Xu, 2020)Paid bookWeeks, read in partsDepth on how experiment programs work and failToo long to finish before most reviews
Your experimentation platform's documentationFree with the toolAn hour or twoHow your own tool computes intervals, power and checksExplains the tool, not the judgment calls
Pairing with your team's data scientist on one readoutFree, costs their timeOne or two sessionsYour company's metrics, conventions and past mistakesDepends on their availability, and leaves no structured record
LearnPathFree covers step 1 of each path; later steps and the certificate need ProDepends on the pathAn ordered path of free YouTube videos with a quiz generated from each video's transcriptOnly step 1 on a free account; it does not know your company's metrics

The book, "Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing" by Ron Kohavi, Diane Tang and Ya Xu (Cambridge University Press, 2020), is an optional paid extra, not a starting point for this week. Details are at experimentguide.com. If you go further into data work after this, our guide on how to learn data science from YouTube is the longer road.

What to skip this week

Skip Bayesian A/B testing, sequential testing math, variance reduction methods, multi-armed bandits and deriving formulas by hand. All of them are real and several are useful later. None of them is needed to read a readout, question it and plan the next test before your review.

Bayesian versus frequentist debates. Learn to read the approach your tool uses; the argument is for another month.

Sequential testing and variance reduction. Know that the first answers peeking and the second lets tests finish with fewer users. Whoever runs the platform needs the inside detail; you do not.

Multi-armed bandits. They shift traffic toward the winner during the test, which answers a different question than "did this change work?"

Deriving formulas by hand. The Day 3 rule of thumb is enough to sanity-check a plan.

The pattern is the same one our one-week Google Analytics 4 plan for marketers uses: learn the parts that let you ship one decision, and put the rest on a later list. If your launch also needs a design review, the two-week Figma plan for product managers follows the same shape.

Frequently Asked Questions

What statistics does a product manager need for A/B testing?

Five ideas, in order. What the null hypothesis and the p-value say. What a confidence interval says about the size of the effect. How power and sample size set the length of a test. How peeking, p-hacking and a broken split make results lie. And how to read a full readout with all of that in mind.

Can you learn A/B testing statistics in a week?

Enough to read a readout, question it and plan the next test, yes. Five days at about an hour each covers p-values, confidence intervals, sample size, the common failure modes and one real walkthrough. It will not make you a statistician, and it will not cover sequential testing or variance reduction methods.

What does a p-value actually tell me about an A/B test?

It tells you how surprising your result would be if the change did nothing at all. A small p-value means a difference this big would be rare under that assumption. It does not tell you the chance that B is better, and it says nothing about whether the effect is large enough to matter.

How long should an A/B test run?

Long enough to reach the sample size you calculated before launch, and usually in whole weeks so weekday and weekend behavior both count. The sample size comes from your baseline rate, the smallest lift worth detecting, and the power you want. Stopping the moment a result looks significant is the mistake to avoid.

Why is peeking at an A/B test a problem?

Each time you look at the results and are willing to stop on a significant one, you give chance another opportunity to cross the line. Look often enough and a change that does nothing will eventually look like a winner. Either fix the stopping date in advance or use a method designed for continuous monitoring.

Does LearnPath give a certificate for finishing an A/B testing learning path?

Yes, on the Pro plan; free accounts do not get one. LearnPath builds an ordered path of free YouTube videos and generates a quiz from each video's transcript. Free accounts get step 1 of each path, and later steps plus the completion certificate need Pro. Treat it as proof you finished the plan.

Proving you did it

At the end of the week you should have three things written down. Your five-line review of your last readout. A sample size estimate for your next test, with the arithmetic shown. And one sentence you can say in the review: ship, kill or rerun, and the evidence for it. That sentence is the deliverable.

Be honest about what five days bought. Not a statistics degree, but the right questions about a result, asked in the right order, before your name goes on a launch.

LearnPath turns a topic into an ordered path of free YouTube videos and generates a quiz from each video's transcript, so you find out whether the watching stuck. Free accounts get step 1 of each path; later steps are on the Pro plan. Completion certificates are Pro-only, and free accounts do not get one. A certificate proves you finished a plan. What proves you can do the job is a launch decision that still holds up a month later.

Build a work-ready A/B testing path before your next experiment review.

On this page

Quick AnswerDo this before you watch anythingDay 1Day 2Day 3Day 4Day 5How the options compareWhat to skip this weekFrequently Asked QuestionsProving you did it

Your learning path

Any topic, as a course

The right lessons in order, a short quiz after each.

Build my pathFree to start. Any topic.
Any topic, in the right orderBuild your path free
Start free

Keep reading

Learning guides

Can You Learn SQL in a Month From YouTube? (2026 Plan)

Find out what SQL you can really learn from YouTube in one month, which courses to follow, and a four-week plan that ends with queries you can show an employer.

11 min read
Learning guides

Python for Excel Users: The 2026 Two-Week Automation Plan

A 2026 ten-day plan using freeCodeCamp, Corey Schafer and Tech With Tim, so you can turn this week's raw exports into a finished Excel report with one scheduled script.

13 min read
Learning guides

We Studied 2,151 YouTube Lessons: What Actually Teaches

Data from 1,340 learners and 54,000 YouTube search results: how long the lessons that teach really are, why half of them are years old, and which channels show up most.

10 min read

More free YouTube learning guides

  • 10 Best AI YouTube Channels in 2026 for Learning AI
  • 10 Best DevOps YouTube Channels in 2026 (Kubernetes & CI/CD)
  • 10 Best YouTube Channels for Cloud Engineering (AWS, Azure, GCP) in 2026 (Ranked)
  • 10 Best YouTube Channels for Go Programming in 2026 (Ranked)
  • 10 Best YouTube Channels to Learn Linux in 2026 (Ranked)
  • 11 Best YouTube Channels for Kubernetes in 2026 (Ranked)
  • 11 Best YouTube Channels for System Design in 2026 (Ranked for FAANG Interview Prep)
  • 11 Best YouTube Channels to Learn SQL in 2026 (Ranked)
  • 8 Best AI & Machine Learning YouTube Channels 2026 (Math to Research)
  • 8 Best Unity Game Dev YouTube Channels in 2026 (Beginner to Pro)
  • 8 Best YouTube Channels for Android & Kotlin Development in 2026 (Ranked)
  • 8 Best YouTube Channels for Python in 2026 (Ranked)
  • 8 Best YouTube Channels for React Native in 2026 (Ranked)
  • 8 Best YouTube Channels for Vue.js in 2026 (Ranked)
  • 9 Best Cybersecurity YouTube Channels in 2026 (Ethical Hacking to SOC)
  • 9 Best Game Development YouTube Channels in 2026 (Unity, Godot, Unreal)
  • 9 Best YouTube Channels for Data Engineering in 2026 (Ranked)
  • 9 Best YouTube Channels for Data Science in 2026 (Ranked)
  • 9 Best YouTube Channels for iOS Development in 2026 (Ranked)
  • 9 Best YouTube Channels for JavaScript in 2026 (Ranked)
  • 9 Best YouTube Channels for React in 2026 (Ranked)
  • 9 Best YouTube Channels for Rust Programming in 2026 (Ranked)
  • Full-Stack Web Development on YouTube: The 2026 Course
  • How to Become a Site Reliability Engineer: The 2026 Path
  • Learn Python by Building Projects: The 2026 YouTube Course
  • Power BI Certification Prep on YouTube: The 2026 Path
  • Python for AI in 2026: The YouTube Course and Roadmap
  • Terraform Certification Prep on YouTube: The 2026 Path
  • TypeScript Tutorial on YouTube: The 2026 Course, Ranked

Ready when you are

Pick a topic. We'll find the lessons and put them in order.

Type what you want to learn. We build the path in about a minute.

Free to start. Any topic.
LearnPath · Learn it. Don't just watch it.
AboutPricingSupportPrivacyTerms