P-Value

Compare Calculations

Downloads

Includes your inputs and results for this calculation, plus any additional calculations you've compared.

Converting a Z-Score Into a Probability of Chance Alone

A p-value is the probability of seeing a result at least as extreme as the one observed, purely by chance, if there were actually no real effect. Enter a z-score and choose a test type, and this calculator converts it into a p-value — the standard figure used to decide whether a result is statistically significant. If you have raw sample data instead of an already-computed z-score, find it first with the Z-Score Calculator.

A smaller p-value means the observed result would be less likely to happen by chance alone, which is stronger evidence against the assumption that nothing real is going on (the “null hypothesis”). By long-standing convention across most fields, a p-value below 0.05 is generally treated as statistically significant — but that threshold is a convention, not a law of nature, and what counts as “significant enough” can vary by field and by how costly a wrong conclusion would be.

The Formula

For a left-tailed test (is the result unusually low?):

p=Φ(z)p = \Phi(\vA{z})

For a right-tailed test (is the result unusually high?):

p=1Φ(z)p = 1 - \Phi(\vA{z})

For a two-tailed test (is the result unusual in either direction?):

p=2×(1Φ(z))p = 2 \times \left(1 - \Phi(|\vA{z}|)\right)

Where z\vA{z} is the z-score and Φ\Phi is the standard normal distribution’s cumulative distribution function — the probability that a randomly drawn value from that distribution falls at or below a given point, computed here via a well-established polynomial approximation of the error function.

Worked Example

    1. A z-score of 1.96 is the textbook reference point for a two-tailed test at the conventional 95% confidence level: p=2×(1Φ(1.96))2×(10.975)0.0500p = 2 \times (1 - \Phi(1.96)) \approx 2 \times (1 - 0.975) \approx 0.0500. 2. A p-value of about 0.05 sits right at the conventional significance threshold. 3. A larger z-score of 2.576 (the two-tailed reference point for 99% confidence) gives a smaller, more significant p-value of about 0.0100.

Key Factors to Consider

  • A p-value does NOT tell you the probability that the null hypothesis is true. This is one of the most common misinterpretations in statistics — a p-value of 0.03 means there’s a 3% chance of seeing a result this extreme (or more extreme) if the null hypothesis WERE true, not a 3% chance that the null hypothesis is actually correct.
  • A one-tailed and two-tailed test on the same z-score give different p-values, and choosing the wrong one can mislead. A one-tailed test only checks for an effect in one specific direction, while a two-tailed test checks for an effect in either direction — the choice should be decided based on the actual research question BEFORE looking at the data, not selected afterward to produce a more favorable-looking p-value.
  • Statistical significance and practical significance are two different things. With a large enough sample size, even a tiny, practically meaningless effect can produce a very small p-value — a statistically significant result is not automatically an important or large real-world effect, and effect size should be considered alongside the p-value.
  • The conventional 0.05 threshold is a widely-used convention, not a universal scientific law. Some fields (like particle physics) require far stricter thresholds before declaring a discovery, while others treat 0.05 as a reasonable standard — what counts as “significant enough” depends on the field and the cost of being wrong.

Common Mistakes

  • Deciding the test type after looking at the results. The choice between a one-tailed and two-tailed test should follow from the actual research question, decided before the data is in hand — switching to whichever test happens to produce a smaller p-value after the fact is a form of data manipulation, even if unintentional.
  • Reporting a p-value without any measure of effect size. A p-value only speaks to how unlikely a result is under the null hypothesis — it says nothing about how large or meaningful that result actually is, which is why a p-value is usually reported alongside a confidence interval or effect size, not on its own.
  • Treating p = 0.049 and p = 0.051 as fundamentally different outcomes. The 0.05 cutoff is a convention, not a bright line in nature — two results this close are practically identical in strength of evidence, even though one technically clears the threshold and the other doesn’t.
  • Running many tests and reporting only the ones that came out significant. Testing enough comparisons will eventually turn up a “significant” result by chance alone — a p-value’s meaning assumes a single, pre-planned test, not the best result cherry-picked from many.

Useful to Know

  • Starting from raw sample data instead of an already-known z-score? Z-Score Calculator converts a raw value against a known mean and standard deviation into the z-score this calculator needs.
  • Want a range of plausible values instead of a single significance verdict? Confidence Interval Calculator expresses uncertainty as an interval, a common companion to a p-value rather than a replacement for it.
  • Need summary statistics — mean, standard deviation, variance — from a raw list of numbers first? Statistics Calculator computes those directly from your data.

Source: Standard normal distribution and the Abramowitz & Stegun error function approximation.

Frequently Asked Questions

What counts as a "statistically significant" p-value?

The most common convention across many fields is a p-value below 0.05 (a 5% chance the result happened by chance alone). This is a widely-used convention, not a strict rule — some fields use stricter thresholds (like 0.01) when a wrong conclusion would be especially costly.

How do I know whether to use a one-tailed or two-tailed test?

Use a two-tailed test when you care whether a result is unusual in EITHER direction (e.g. "is this coin unfair?"). Use a one-tailed test (left- or right-tailed) only when you specifically care about ONE direction ahead of time (e.g. "is this new process FASTER?") — deciding this after seeing the data is considered poor statistical practice.

What's the difference between this and the Z-Score Calculator?

The Z-Score Calculator converts a raw value into a z-score and percentile against a known distribution. This calculator starts from an already-known z-score and answers the specific question a percentile alone does not: how likely is a result this extreme to have happened purely by chance?

Confirm Your Age

To create an account, please tell us your birth month and year.