Skip to content
Pureza

Research concepts

Statistical significance is not effect size

The word significant is often read as important, but in statistics it answers another question. Understanding a result requires its estimated magnitude, units and uncertainty. A p-value adds information under a model; it replaces none of those pieces.

Source editorial review:

What statistical testing contributes

The ASA task force statement places p-values and statistical tests among tools for assessing observations against sampling variation. It does not equate them with practical importance or propose discarding them.

Interpretation depends on design and the analytical model. A small number does not resolve selection, measurement or reporting problems. Understand which comparison was proposed and whether it was chosen before the results were observed.

Magnitude needs a scale

The Cochrane Handbook distinguishes measures in natural units from standardized measures. A concentration difference, a ratio and a standardized difference do not express the same thing. Interpretation must retain the scale and comparison group.

As a hypothetical example, two analyses may receive the same significance label although one estimates a small difference and the other a large one. The label alone cannot distinguish them. Read the estimate and determine what magnitude would meaningfully answer the experimental question.

An interval shows precision under assumptions

NIST's explanation of intervals for a mean relates their width to variability, sample size and confidence level. Two studies can have identical point estimates but very different intervals.

In frequentist interpretation, the confidence level describes the procedure's behavior over repetitions, not a probability assigned to the parameter after a specific interval is observed. The calculation also depends on assumptions: a narrow interval does not automatically correct a biased sample.

Not crossing a threshold does not demonstrate absence

Cochrane recommends presenting intervals and exact p-values without relying on binary labels. A result compatible with a broad range of effects requires a different conclusion from one that precisely excludes relevant magnitudes.

In practice, ask which values remain compatible with the data and model. If the interval is wide, uncertainty remains even if the text says no differences. If narrow, it may bound a small response without demonstrating absolute equality.

Check how many comparisons were made

The ASA highlights multiplicity and selective reporting as issues to consider from design through reporting. Highlighting one favorable comparison among many can give an incomplete picture.

To summarize an article, collect the variable, comparison, estimate, interval and analytical context. Then check whether the full set of tests is reported and how multiplicity was handled. This uses statistics without turning an isolated threshold into a biological claim.

Questions and answers

Does a smaller p-value demonstrate a larger effect?

No. Effect size is read from the estimate and its scale; the p-value also depends on precision and assumptions.

Does a narrow interval guarantee study validity?

No. It describes precision under the model used but does not correct design or measurement biases.

Sources

  1. ASA: declaración del grupo de trabajo sobre significación y replicabilidad
  2. NIST: límites de confianza para la media
  3. Cochrane: interpretación de resultados y conclusiones