Critical Value vs P-Value
The critical value and p-value approaches always reach the same decision. How they relate, when each is easier, and how to report results either way.
Critical Value vs P-Value: Same Decision, Two Routes
Most students hit this wall: your textbook shows one method using a critical value that you compare to a test statistic, then the next chapter shows a p-value that you compare to alpha. They feel like different subjects. They are not. The critical value vs p value distinction is a difference in route, not destination. Both methods always produce the same reject-or-not decision for the same test. The confusion comes because textbooks teach them as separate procedures and rarely show them side by side. This walks both approaches on one problem, then explains why they never disagree.
The failure case that hits most often: you compute a p-value and a test statistic, get numbers that look unrelated, and have no idea which one to trust. Trust both. They are the same procedure expressed differently. The p-value is the probability of seeing a test statistic as extreme as yours if the null is true. The critical value is the threshold that test statistic must cross at your chosen alpha. If the test statistic crosses the critical value, the p-value is smaller than alpha. If it does not, the p-value is larger. They are mirror images.
The Two Approaches Side by Side
Route one: critical value approach. Pick your significance level (alpha, almost always 0.05 or 0.01). Look up the critical value from the appropriate distribution (Z, t, chi-square, or F). Compute your test statistic from the sample data. If the test statistic is more extreme than the critical value, reject the null hypothesis. Decision rule: |test statistic| > critical value → reject.
Route two: p-value approach. Compute the same test statistic. Find the probability of getting a value that extreme or more extreme under the null distribution. That probability is the p-value. Compare it to your alpha. Decision rule: p-value < alpha → reject.
Every textbook that teaches both routes, OpenIntro Statistics (4th ed.) by Diez, Cetinkaya-Rundel, and Barr, or Moore, McCabe & Craig's Introduction to the Practice of Statistics (9th ed.), presents them as separate chapters. They are the same logic. The critical value is just the p-value turned inside out: the critical value is the test statistic that would exactly produce a p-value equal to alpha.
Why They Always Agree
The agreement is mathematical, not coincidental. For a test statistic T and a critical value C at a significance level alpha, the relationship is always: p-value = P(T > observed T | H₀). The critical value C is defined as the value such that P(T > C | H₀) = alpha. If your observed test statistic is larger than C, then the area in the tail beyond it (the p-value) must be smaller than alpha. If your test statistic is smaller than C, the tail area is larger than alpha. The two rules, compare test statistic to critical value, compare p-value to alpha, are exactly equivalent.
The one detail that trips people: for a two-tailed test, you compare |test statistic| to the two-tailed critical value, and the p-value is the two-tailed probability (the sum of both tails). The equivalence holds for symmetric distributions. For asymmetric distributions like chi-square and F, the same logic works: the critical value is the single threshold in the right tail, and the p-value is the right-tail probability beyond your test statistic. The ASA Statement on Statistical Significance and P-Values (Wasserstein & Lazar 2016) emphasises that neither method measures effect size or practical importance, both only assess whether the observed data are surprising under the null.
When Textbooks and Software Use Each
Introductory textbooks lean on the critical value approach because it is visual: you draw the sampling distribution, shade the rejection region, and see whether the test statistic falls inside it. OpenIntro Statistics uses critical values for all its worked examples in the hypothesis testing chapters. Moore, McCabe & Craig do the same, then introduce p-values as a separate concept for reporting results.
Software and calculators default to p-values. Excel and Google Sheets return p-values from functions like T.TEST and CHISQ.TEST. The TI-84's STAT > TESTS menu gives p-values, not critical values. The reason: p-values give a continuous measure of surprise (a p-value of 0.04 is more borderline than 0.001), while a critical value only tells you pass or fail at one alpha. Advanced users need p-values because they can report them directly in papers ("p = 0.04") without committing to an alpha level.
But the critical value approach survives where the p-value is hard to compute by hand, for example, when using a Z-table for a one-tailed test, you can look up 1.645 and compare directly without converting the test statistic to a probability. The p value approach vs critical value approach question is really about what you have at hand: a table of critical values or a computer that can integrate the tail.
P-Value From Test Statistic
Converting a test statistic to a p-value is a three-step operation that applies to any distribution (Z, t, chi-square, F). Step one: compute the test statistic from your data. Step two: identify the correct distribution (Z if sigma is known, t if sigma is estimated from sample SD, chi-square for variance or goodness-of-fit tests, F for ANOVA). Step three: find the area in the tail(s) beyond that test statistic.
Z and t (Symmetric Distributions)
For a two-tailed test, the p-value is 2 × P(T > |test statistic|). For a right-tailed test, it is P(T > test statistic). For a left-tailed test, it is P(T < test statistic). On the TI-84, use normalcdf(lower bound, upper bound, mean, SD) for Z or tcdf(lower bound, upper bound, df) for t. In Excel, use NORM.S.DIST(z, cumulative) for Z or T.DIST(t, df, cumulative).
Chi-Square and F (Asymmetric, Right-Tailed Tests)
The p-value is always the right-tail probability: P(χ² > test statistic) or P(F > test statistic). In Excel, use CHISQ.DIST(test stat, df, FALSE) for the PDF or CHISQ.DIST.RT(test stat, df) for the right-tail area directly. For F, use F.DIST.RT(test stat, df1, df2). The most common mistake at this stage is using the left-tail instead of the right-tail, chi-square and F hypothesis tests are universally right-tailed, so the critical value is always the upper quantile.
Reporting Results: Which Method Do You Publish?
In a paper, you report the p-value and the test statistic, not the critical value. The critical value is internally used to determine significance but is not published. A typical result reads: "t(23) = 2.45, p = 0.03" or "χ²(4) = 11.8, p = 0.02". The critical value (e.g., t(23) at α=0.05 two-tailed = 2.069) is your private threshold, not part of the published output, unless you are reporting a confidence interval, which uses the same critical value as the corresponding hypothesis test.
The ASA Statement on Statistical Significance and P-Values (Wasserstein & Lazar 2016) warns that reporting only a p-value and a significance decision ("p < 0.05") without the effect size or confidence interval is insufficient. The critical value and p-value together tell you whether the result is surprising under the null; they do not tell you whether the effect is large enough to matter. A tiny p-value from a huge sample can come from a trivial effect. Report the test statistic and the p-value, then interpret the practical size of the effect separately.
The single most practical thing you can do: next time you run a hypothesis test, compute both the test statistic and the p-value. Verify that the decision matches, if |test statistic| > critical value, the p-value must be < alpha. If they disagree, you have used the wrong distribution, the wrong tail, or the wrong alpha. That self-check catches the errors that sink most students' homework: using Z when sigma is unknown, using the wrong degrees of freedom, or forgetting to halve alpha for a one-tailed versus two-tailed test.
Worked Example: Same Test Both Ways
A researcher wants to test whether the mean exam score of a class is different from 75. Population standard deviation is unknown. Sample size is 16. Sample mean is 80. Sample standard deviation is 10. Alpha is 0.05.
Route one: critical value. Test statistic: t = (80 − 75) / (10 / √16) = 5 / 2.5 = 2.00. Degrees of freedom: 15. Critical value from t-table at α=0.05 two-tailed, df=15: 2.131. Compare: |2.00| < 2.131. Fail to reject H₀. No significant difference.
Route two: p-value. Same test statistic: t = 2.00 with df=15. Use tcdf(2.00, ∞, 15) on TI-84, double for two-tailed. One-tailed p-value: approximately 0.036. Two-tailed p-value: 0.072. Compare p = 0.072 to α = 0.05. p > 0.05. Fail to reject H₀. Same decision.
The critical value of 2.131 is the t-value that would produce a two-tailed p-value of exactly 0.05. Your test statistic of 2.00 is smaller, so its tail area (0.072) is larger than 0.05. The two numbers look different but the logic is identical. This is why the p value approach vs critical value approach is not a choice between methods, it is a choice of arithmetic.
Common Questions
Which method is easier to learn first?
Most students find the critical value approach easier to visualise, you draw the distribution, shade the rejection region, and see whether your test statistic falls inside it. Once you understand that, the p-value becomes the area in the tail beyond the test statistic. The p-value approach is more flexible for reporting because it does not require a specific alpha.
Do I ever need to compute both?
In practice, you compute one or the other. But computing both once and verifying they agree is the best way to catch errors. If your test statistic crosses the critical value but your p-value is greater than alpha, you used the wrong distribution or the wrong tail.
Why does my textbook show a different critical value table for t than for Z?
Because the t-distribution has heavier tails than the normal distribution when degrees of freedom are small. As df increases, the t critical value approaches the Z critical value. At df=30, t(30) for α=0.05 two-tailed is 2.042, very close to Z=1.96. But at df=5, the t critical value is 2.571, much larger.
Can I use the p-value method for a one-tailed test?
Yes. For a one-tailed test, you do not double the tail area. If your test statistic is 2.00 and the one-tailed p-value is 0.036, you compare that 0.036 directly to alpha. If alpha is 0.05, reject. The critical value for a one-tailed test at α=0.05 is 1.645 for Z; your test statistic of 2.00 crosses it.
What if my test statistic is exactly equal to the critical value?
You are exactly at the threshold. The p-value equals alpha exactly. In practice, this is vanishingly rare with continuous distributions. Convention says you reject the null, the test statistic is in the rejection region (the boundary is included). Most textbooks and software treat the critical value as the cut-off.
Does the critical value change if I use a two-tailed test vs a one-tailed test?
Yes, and this is the most common source of errors. For α=0.05, the two-tailed Z critical value is 1.96 (alpha split between both tails, 0.025 each). The one-tailed Z critical value is 1.645 (all alpha in one tail). Using a one-tailed critical value in a two-tailed test makes the threshold too low, inflating Type I error.
When would I need a chi-square or F critical value?
Chi-square critical values are used in goodness-of-fit tests (e.g., testing whether a sample follows a normal distribution) and in tests of independence (e.g., whether two categorical variables are related). F critical values are used in ANOVA to test whether group means are all equal. Both are always right-tailed.