Hypothesis Tests and P Values Done Right
Run tests, read p values correctly and report effects that actually matter
A taste of a lesson
Our test of a new checkout design gave p = 0.03. Can I tell my team there's a 97 percent chance it works?
That would overstate it. p = 0.03 means: if the new design truly made no difference, results at least this extreme would happen about 3 percent of the time. That is evidence against 'no difference', but it is not the probability the design works, and it says nothing about how big the improvement is. Report the estimated lift with its confidence interval, for example plus 1.2 points, interval 0.1 to 2.3. Exercise: rewrite your team message using the estimate and interval instead of the p value.
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- State null and alternative hypotheses for a real question
- Define a p value precisely and reject common misreadings
- Explain type I and type II errors and the role of power
- Separate statistical significance from practical importance
- Handle multiple comparisons and recognise p hacking
Lesson plan
- 1 The logic of a test Set up a null hypothesis, alternative and test statistic. Start
- 2 What a p value is and is not Define the p value precisely and avoid the classic misreadings. Start
- 3 Errors and power Understand false positives, false negatives and power. Start
- 4 Common tests in concept Match the question and data type to a suitable test. Start
- 5 Significance versus importance Report effect sizes and intervals alongside p values. Start
- 6 Many tests and p hacking Control false positives when testing many hypotheses. Start
Try asking
About this tutor
An intermediate tutor for anyone who runs or reads statistical tests and wants to stop misusing p values. You will set up null and alternative hypotheses, understand test statistics and the exact meaning of a p value, and learn the two kinds of error and the role of power. Lessons cover common tests in concept (t tests, chi squared tests, tests for proportions), statistical versus practical significance, multiple comparisons and p hacking. You finish reporting results with effect sizes and confidence intervals, and able to spot overclaiming in papers and dashboards.
Reviews
4.7
3 ratingsSample
- Priyanka D.Sample
Strong on multiple comparisons. Our dashboard was testing dozens of segments and nobody had corrected for it.
- Liam O.Sample
Balanced and honest. It did not tell me to abandon p values, just to use them with effect sizes.
- Victor N.Sample
The true or false drills on p value wording were humbling. I had been misreading them for years.
About the teacher
Statistics in plain language, from averages to Bayesian reasoning
9 tutors 350 lessons taught Sample
I teach statistics to people who were put off by it the first time. My approach is to start from a question someone actually has, simulate or count our way to an answer, and only then name the formula. I have worked as an analyst on survey and health research projects, so I have a soft spot for messy samples,...
See Lina's profile and tutorsMore like this
Other tutors on the same or nearby topics.