ResearchOS/Wiki

Survival curves and hazard ratios

When your outcome is how long until something happens, time to relapse, time to death, time to a tumor reaching a size, you need methods built for time-to-event data. The reason ordinary tests fall short is that some subjects have not had the event yet when the study ends, and you cannot just drop them or pretend they never will. This page covers Kaplan-Meier curves, the log-rank test, and the Cox hazard ratio.

Why time-to-event data is special

The defining feature is censoring. A patient who is still alive at the last follow-up has not had the event, but you do not know when, or if, they will. That is real information, they survived at least this long, and survival methods use it rather than throwing it away. Averaging the times you happened to observe would be badly biased, because the longest survivors are exactly the ones still censored.

Kaplan-Meier curves

The Kaplan-Meier estimate is the survival curve itself, a stepped line showing the fraction of subjects still event-free over time. It steps down at each event and accounts for censored subjects by dropping them out of the at-risk pool without counting them as events. Reading it is intuitive, a curve that stays high means subjects are lasting longer, and the gap between two curves is the difference between groups.

The Data Hub reports the curve plus the median survivalfor each group, the time by which half the subjects have had the event, with its confidence interval. Median survival is often the cleanest single summary, "median time to relapse was 14 months in the treated arm versus 9 in the control."

A two-arm comparison. The table gives median survival for each arm, the time by which half the subjects have had the event, and the log-rank test for whether the curves differ overall. The Gehan-Breslow-Wilcoxon test below it is the same comparison weighted toward early time points, so it is more sensitive to early differences.

The log-rank test

The log-rank test asks whether two (or more) Kaplan-Meier curves differ more than chance would explain. It compares the whole curves over the entire follow-up, not just a single timepoint, so it uses all the event timing. The result is a p-value. A small one means the survival experience genuinely differs between groups.

Gehan-Breslow-Wilcoxon test

The Data Hub reports a second test alongside the log-rank: Gehan-Breslow-Wilcoxon. It uses the same observed-minus-expected machinery as the log-rank, but each event time's contribution is weighted by the total number of subjects still at risk at that moment. Because the risk set is largest early in follow-up and shrinks as subjects fail or censor, the Gehan-Breslow-Wilcoxon test gives early events more weight and late events less, making it more sensitive to curves that separate early and converge later. The log-rank weights every event time equally, so it is more sensitive to a consistent difference across the whole follow-up. When you expect or see a crossing of the two curves, neither test is ideal; read both p-values and note whether they tell the same story. Both report a chi-square statistic with groups minus 1 degrees of freedom.

Hazard ratios and the Cox model

The hazard ratio is the effect size for survival data. The hazard is the instantaneous risk of the event at any moment among those still at risk. The Cox regressionmodel estimates how a factor multiplies that risk, and reports it as a ratio.

  • A hazard ratio of 1 means no difference in risk.
  • Below 1 means lower risk. A hazard ratio of 0.6 for the treatment means treated subjects had 60% the event risk of controls at any given moment, a protective effect.
  • Above 1 means higher risk. A hazard ratio of 1.8 means 80% more risk.

The result reports the hazard ratio, its 95% confidence interval, and a p-value. Read the interval exactly as on the effect sizes page, if it does not include 1, the effect is statistically clear, and its width tells you how precisely you have pinned the risk change down. The Cox model can also adjust for other variables at once (age, stage), reporting a hazard ratio for each, much like multiple regression for survival.

The Cox result also reports two overall model summaries. The likelihood-ratio chi-square compares the fitted model to the null model (all coefficients zero), with degrees of freedom equal to the number of covariates; a significant chi-square means the covariates together improve the survival prediction. Harrell's concordance (c-index) measures how well the model ranks subjects, the probability that of a comparable pair (one fails before the other), the model assigns the higher risk to the one who failed first. A c-index of 0.5 means the model ranks no better than chance; 1.0 is perfect discrimination. It is the survival analogue of the AUC and is read the same way.

A worked example

A trial reports median time to progression of 14 months on treatment versus 9 on control, a log-rank p = 0.004, and a Cox hazard ratio of 0.62 (95% CI 0.45 to 0.85). You would write "treatment extended median time to progression from 9 to 14 months and reduced the hazard of progression by 38% (hazard ratio 0.62, 95% CI 0.45 to 0.85, log-rank p = 0.004)." The 38% comes from 1 minus 0.62, and the interval not crossing 1 is what makes it clear.

ResearchOS validates the Kaplan-Meier estimate, the log-rank test, the Gehan-Breslow-Wilcoxon test, and the Cox model (including concordance and the likelihood-ratio test) against the lifelines package and R's survival library on the transparency page.