P-Value Calculator: Calculate Statistical Significance Instantly
The p-value is the most widely cited — and most commonly misunderstood — number in research and data science. This free p-value calculator computes p-values from Z-scores (using the standard normal distribution) and T-scores (using the t-distribution with any degrees of freedom) for left-tailed, right-tailed, and two-tailed hypothesis tests. Enter your test statistic, select your test type and direction, and instantly receive the p-value, CDF value, statistical decision, and an interactive distribution curve showing the shaded rejection region.
📊 Key formulas:
Z-test, two-tailed: p = 2 × (1 − Φ(|z|)) where Φ is the standard normal CDF
Z-test, right-tailed: p = 1 − Φ(z) · Left-tailed: p = Φ(z)
T-test: same structure using the t-distribution CDF with df degrees of freedom
Example: z = 1.96, two-tailed → p = 2 × (1 − Φ(1.96)) = 2 × 0.0250 = 0.0500
What Is a P-Value?
A p-value is the probability of observing a test result at least as extreme as the one actually observed, assuming the null hypothesis is true. It is not the probability that the null hypothesis is true, nor the probability of making an error — two of the most common misconceptions. The p-value is a conditional probability: P(data this extreme | H₀ is true).
When the p-value is small (below a pre-specified significance level α), it means the observed data is unlikely under H₀ — providing evidence to reject the null hypothesis. When the p-value is large, the data is consistent with H₀ and we fail to reject it. “Fail to reject” does not mean H₀ is true — it means we lack sufficient evidence to reject it.
Common Significance Levels and Their Meaning
| Significance level (α) | Z-score threshold (two-tailed) | T-score (df=30) | Common use |
| α = 0.10 | |z| > 1.645 | |t| > 1.697 | Exploratory research, social sciences |
| α = 0.05 | |z| > 1.960 | |t| > 2.042 | Most standard research, default threshold |
| α = 0.01 | |z| > 2.576 | |t| > 2.750 | Clinical trials, high-stakes decisions |
| α = 0.001 | |z| > 3.291 | |t| > 3.646 | Physics, genomics (multiple comparisons) |
Z-Score vs T-Score: When to Use Each
📏Use Z-test when
Population standard deviation (σ) is known; sample size is large (n ≥ 30); you’re testing a proportion; or comparing two proportions. The Z-distribution (standard normal) assumes the test statistic follows N(0,1) under H₀. Common in quality control, proportion testing, and large-sample studies.
📐Use T-test when
Population standard deviation is unknown (must be estimated from sample); sample size is small (n < 30); you’re comparing sample means. The t-distribution is wider than normal (heavier tails) to account for the additional uncertainty from estimating σ. T-distribution approaches normal as df increases.
◄Left-tailed test
The alternative hypothesis states the true parameter is less than the null value. The rejection region is in the left tail. H₁: μ < μ₀. Example: testing whether a new drug reduces blood pressure below the control group’s level.
↔Two-tailed test
The alternative hypothesis states the parameter differs from the null value in either direction. Rejection regions are in both tails. H₁: μ ≠ μ₀. The most commonly used test direction — appropriate when you don’t have a directional hypothesis before collecting data.
Common Misconceptions About P-Values
- “p < 0.05 means H₀ is false": A p-value below α means the data is unlikely under H₀ — not that H₀ is definitively false. Statistical significance is probabilistic, not certain.
- “p-value measures effect size”: A very small p-value in a large study may reflect a tiny, practically meaningless effect. Always report effect sizes (Cohen’s d, r², etc.) alongside p-values.
- “p = 0.06 means no effect”: A p-value of 0.06 at α = 0.05 is not dramatically different from 0.04. The α threshold is a convention, not a law of nature. Interpret p-values on a continuum.
- “p-value is the probability that H₀ is true”: This is incorrect. The p-value is P(data | H₀ is true), not P(H₀ is true | data). These require Bayesian methods.
Related Statistics Calculators
Frequently Asked Questions
What is a p-value?
A p-value is the probability of obtaining a test result at least as extreme as the observed result, given that the null hypothesis (H₀) is true. It measures how compatible the observed data is with H₀. A small p-value suggests the data is incompatible with H₀ — evidence to reject it. Important: the p-value is not the probability that H₀ is true, nor the probability of making an error. It’s P(data this extreme or more | H₀ is true).
How do you calculate a p-value?
Step 1: Calculate your test statistic (Z-score or T-score) from your data. Step 2: Determine the appropriate distribution (normal for Z-test, t-distribution for T-test with df = n − 1). Step 3: Find the cumulative probability (CDF) at your test statistic. Step 4: Compute p based on tail direction: Left-tailed: p = CDF(score); Right-tailed: p = 1 − CDF(score); Two-tailed: p = 2 × (1 − CDF(|score|)). This calculator performs all these steps automatically when you enter your score.
What does p < 0.05 mean?
p < 0.05 means that if the null hypothesis were true, there would be less than a 5% probability of observing a test result as extreme as yours by random chance. At α = 0.05, this is the conventional threshold for “statistical significance” — the result is unlikely enough under H₀ that we reject it. However, p = 0.05 is a convention, not a universal truth. The appropriate threshold depends on the study design, consequences of errors, and field norms. Some fields use α = 0.01 (clinical trials) or α = 0.001 (genomics).
Is a smaller p-value better?
“Better” depends on context. A smaller p-value represents stronger evidence against the null hypothesis. If you hypothesise an effect exists, a smaller p-value is more supportive. However, p-values don’t measure effect size, practical importance, or scientific significance. A p = 0.000001 from a massive sample might reflect a trivially small real-world effect. Always pair p-values with effect sizes (Cohen’s d, R², etc.) and confidence intervals for a complete statistical picture.
What is a Z-score?
A Z-score (or standard score) measures how many standard deviations an observation is from the population mean. Z = (X − μ) / σ, where X is the observed value, μ is the population mean, and σ is the population standard deviation. In hypothesis testing, the test statistic Z = (x̄ − μ₀) / (σ / √n) follows a standard normal distribution under H₀ when σ is known and n is large. Z-scores above |1.96| reject H₀ at α = 0.05 (two-tailed).
When should I use a T-test?
Use a t-test when: (1) the population standard deviation is unknown and must be estimated from the sample; (2) sample size is small (n < 30); (3) comparing two group means (two-sample t-test) or before/after measurements (paired t-test). The t-distribution accounts for additional uncertainty by having heavier tails than the normal distribution — this extra uncertainty comes from estimating σ from a finite sample. As degrees of freedom increase, the t-distribution converges to the normal distribution (at df ≥ 1000, they’re virtually identical).
What is a two-tailed test?
A two-tailed test checks whether the population parameter differs from the null value in either direction (larger or smaller). The alternative hypothesis is H₁: μ ≠ μ₀. The rejection region is split between both tails of the distribution, with α/2 in each tail. For α = 0.05, each tail contains 2.5% of the distribution, giving a critical Z-score of ±1.96. Use a two-tailed test when you have no prior directional hypothesis — this is the more conservative and generally recommended default.
What is statistical significance?
Statistical significance means the observed result is unlikely to have occurred by chance alone under the null hypothesis, at a pre-specified probability threshold (α). A result is “statistically significant” when p < α. It does not mean the result is large, important, or practically meaningful. A study can be statistically significant (strong evidence of a real effect) but practically insignificant (the effect is too small to matter). Conversely, a study may have a meaningful effect that fails to reach significance due to small sample size (insufficient power).
How accurate is this p-value calculator?
The Z-test calculations use the Abramowitz and Stegun approximation to the normal CDF, with an absolute error of less than 7.5 × 10⁻⁸ — accurate to at least 7 decimal places. The t-distribution calculations use the regularised incomplete beta function with Lentz’s continued fraction algorithm, accurate to approximately 7 significant figures. Results match statistical software (R, Python scipy, SPSS) to 5–6 decimal places for typical test values. For extreme values (|z| > 6, very small df), numerical precision may decrease slightly, but results remain accurate for all practical research purposes.
What is the difference between one-tailed and two-tailed p-values?
For the same test statistic, a two-tailed p-value is exactly twice the one-tailed p-value. Example: Z = 1.96. Right-tailed p = 0.0250. Left-tailed p = 0.9750. Two-tailed p = 0.0500. Choose the tail direction based on your alternative hypothesis, not based on the observed direction of your data (which would inflate Type I error). One-tailed tests are appropriate when there’s a strong prior directional hypothesis; two-tailed tests are the conservative default for most research.
Calculate your p-value instantly
Z-test, T-test, any tail direction, distribution curve — free, research-grade, accurate.
Calculate p-value ↑