ResearchOS/Wiki

Contingency tables, odds ratios, and relative risk

When your data are counts in categories rather than measured numbers, responded or did not, mutant or wild type, survived or died, you compare proportions. A contingency table lays those counts out, and a handful of tests and effect sizes tell you whether the categories are linked and how strongly. This page covers chi-square, Fisher exact, logistic regression as a first-class analysis, the odds ratio, and relative risk.

What a contingency table is

A contingency table is a grid of counts. It can be any size: two outcomes by two groups (2x2), three treatment arms by four response categories (3x4), or any R-by-C arrangement. The question is always the same: does the distribution of counts across columns depend on which row you are in, or are the rows and columns independent and any apparent pattern is just chance?

Chi-square and Fisher exact

For any R×C table the Data Hub runs the Pearson chi-square test of independence. It compares the counts you observed against the counts you would expect if the rows and columns were independent. The test is reliable when the expected count in every cell is reasonably large (a common rule of thumb is at least 5); the result reports the minimum expected count so you can check.

For a 2x2 table, the Data Hub always reports all three of the following, regardless of cell counts.

  • Chi-square (Pearson, uncorrected), the standard large-sample statistic.
  • Chi-square with Yates continuity correction, which subtracts 0.5 from each absolute deviation before squaring, giving a slightly more conservative result for the 2x2 case.
  • Fisher's exact test, which computes the probability directly from the hypergeometric distribution rather than approximating, so it stays accurate when counts are small. Fisher exact is only computed for 2x2 tables; for larger tables the chi-square is the right test.

All three report a p-value for whether the categories are associated. As elsewhere, that tells you whether there is a link, not how strong it is. For strength, you want the effect sizes below.

Odds ratios

The odds ratio is the workhorse effect size for two-by-two data. Odds are a count ratio, the number with the outcome divided by the number without it. The odds ratio compares the odds in one group against the odds in the other.

  • An odds ratio of 1 means the outcome is equally likely in both groups, no association.
  • Above 1 means higher odds in the first group. An odds ratio of 3 means the exposed group had three times the odds of the outcome.
  • Below 1 means lower odds, a protective association.

The Data Hub reports the odds ratio with its 95% confidence interval. If the interval excludes 1, the association is statistically clear, the same call the test's p-value makes, and the interval's width tells you the precision.

A 2x2 table. The Data Hub reports the chi-square (Yates-corrected and uncorrected), Fisher's exact p for small counts, the relative risk, and the odds ratio with its 95 percent confidence interval, then shows the observed counts against the counts you would expect if the two factors were unrelated.

Relative risk, and when to use which

Relative risk compares the actual probability of the outcome between groups, not the odds. A relative risk of 2 means the outcome was twice as likely in one group. It is often the more intuitive number, and it is the right one when you sampled groups and then watched for outcomes (a cohort or a trial). The odds ratio is the natural choice for case-control designs and is what logistic regression produces. When the outcome is rare the two numbers nearly coincide; when it is common the odds ratio looks more extreme than the relative risk, so do not read an odds ratio as if it were a risk ratio.

A worked example

Of 50 treated patients, 10 relapsed; of 50 controls, 25 relapsed. Fisher's exact test gives p = 0.002. The odds ratio is 0.27 (95% CI 0.11 to 0.64) and the relative risk is 0.40. You would write "relapse was less frequent on treatment, 20% versus 50% (relative risk 0.40, odds ratio 0.27, 95% CI 0.11 to 0.64, Fisher exact p = 0.002)." The interval not crossing 1 is what makes the association clear.

ResearchOS validates the chi-square, Yates-corrected chi-square, Fisher exact test, odds ratio, relative risk, and logistic regression (including the Firth fallback) against scipy and statsmodels on the transparency page.