Method validation

Compare our calculations with their references

This report compares the included ResearchOS calculations with independent references. Open a category to inspect the inputs, results and documented tolerances. The validation tests fail when a result exceeds its tolerance.

426
Comparisons checked
336
Exact match
90
Within tolerance
0
Failing
How this works

The calculations included in this report cover sequence analysis, lab calculators, statistics and phylogenetic layouts. Each included calculation is evaluated over a fixed set of test inputs and compared against an independent reference, a peer-reviewed software package (Biopython, primer3, pydna, scipy, ggtree), a published sequence, or the closed-form result of exact algebra. Reference values are pinned from the cited sources and reproducible with the listed generator scripts. The report and automated validation test use the same comparison code. A result exceeding its stated tolerance fails that test.

Where ResearchOS differs

Where a published algorithm exists, ResearchOS implements that same algorithm and the test verifies it reproduces the reference to the digit. The only non-identical cases are the ones called out here, each a known, documented difference.

  • Sequence alignment · Short homology, 180 bp block (approximation gap grows)Expected difference

    vs Biopython: ResearchOS 0.88, reference 1 (Δ 0.12 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

  • Data Hub statistics · Wilcoxon signed-rank, one-sided (greater) pExpected difference

    vs SciPy: ResearchOS 0.99015, reference 1 (Δ 0.009850468 p)

    scipy uses the EXACT signed-rank distribution at n = 6 (p = 1.0 for the greater tail here); our engine uses the normal approximation with a continuity correction (via @stdlib), giving p ~ 0.99. The roughly 0.01 gap is the documented exact-vs-asymptotic difference for small samples, not a bug. Both lead to the same decision at alpha = 0.05.

  • Data Hub statistics · Wilcoxon signed-rank, pExpected difference

    vs SciPy: ResearchOS 0.034006, reference 0.03125 (Δ 0.002756403 p)

    scipy.stats.wilcoxon uses the EXACT signed-rank distribution at this n (n = 6), giving p = 0.03125. Our engine uses the normal approximation with a continuity correction (via @stdlib), giving p ~ 0.034. The roughly 0.003 gap is the documented exact-vs-asymptotic difference for small samples, not a bug. Both lead to the same decision at alpha = 0.05.

  • Data Hub statistics · Mann-Whitney U, one-sided (greater) pExpected difference

    vs SciPy: ResearchOS 0.995943, reference 0.997672 (Δ 0.001728559 p)

    scipy and our engine both use the asymptotic normal approximation, but the 0.5 continuity correction is applied toward the mean, so on the far (upper) tail our p is about 0.002 smaller (0.9959 vs 0.9977). The near tail (less) matches to floating point. This is the documented continuity-correction direction, not a bug, and both lead to the same decision at alpha = 0.05.

  • Data Hub statistics · Wilcoxon signed-rank, one-sided (less) pExpected difference

    vs SciPy: ResearchOS 0.017003, reference 0.015625 (Δ 0.001378201 p)

    scipy uses the EXACT signed-rank distribution at n = 6 (p = 0.015625); our engine uses the normal approximation with a continuity correction, giving p ~ 0.017. The roughly 0.0014 gap is the documented exact-vs-asymptotic difference for small samples (the same one noted on the two-sided pin), not a bug.

85 further comparisons differ only by a last-digit amount that stays inside each method's documented tolerance, the kind of rounding offset two correct implementations produce. The exact per-case numbers are in each method's comparison table, and the same offsets are listed by domain below.

Show all 85 within-tolerance offsets, by domain

Published-tree reproduction · 1 offset

  • Published-tree reproduction · Craugastor frog multilocus supermatrix (47 taxa)Within tolerance

    vs the published tree: ResearchOS 2.94, reference 0 (Δ 2.94 % of published clades not recovered)

    This published tree is a topology with no branch support, so there is no support to confine differences to and the support-aware rule does not apply. Instead the case passes when it recovers at least a committed fraction of the published clades. We deliberately do NOT use symmetric Robinson-Foulds here, because a published tree with polytomies (unresolved multifurcations) makes our fully resolved ML tree look distant when it actually recovered the published groupings and resolved more; recovery measures the honest thing.

Primer melting temperature (Tm) · 1 offset

  • Primer melting temperature (Tm) · 18-mer, GC termini both ends (where tables differ most)Within tolerance

    vs primer3-py: ResearchOS 55.8882, reference 56.8985 (Δ 1.0103 C)

    primer3 uses the SantaLucia 1998 unified table and its own salt model instead of the Allawi 1997 table we share with Biopython, so a small systematic offset (largest on GC-terminal oligos) is expected, not a bug.

Validated against published results · 6 offsets

  • Validated against published results · US CDC N2 assay, mean slopeWithin tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 94.5438, reference 95 (Δ 0.4562 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

  • Validated against published results · Reported low slope (CDC N1 range)Within tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 89.5736, reference 90 (Δ 0.4264 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

  • Validated against published results · Acceptable-range lower efficiencyWithin tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 90.2522, reference 90 (Δ 0.2522 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

  • Validated against published results · Acceptable-range upper efficiencyWithin tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 110.1748, reference 110 (Δ 0.1748 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

  • Validated against published results · Ideal standard curveWithin tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 100.0805, reference 100 (Δ 0.0805 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

  • Validated against published results · Reported high efficiency (CDC N1 range)Within tolerance

    vs Published RT-qPCR standard-curve values: ResearchOS 161.0157, reference 161 (Δ 0.0157 %)

    Amplification efficiency follows from the standard-curve slope by efficiency% = (10^(-1/slope) - 1) * 100. The paper reports each efficiency to the whole percent, so our value must land within that rounding (under half a percent) of the published number.

Sequence alignment · 5 offsets

  • Sequence alignment · Short homology, 130 bp block (largest approximation gap)Within tolerance

    vs Biopython: ResearchOS 0.9177, reference 1 (Δ 0.0823 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

  • Sequence alignment · Long alignment, 400 bp homologous block in ~6 kb sequencesWithin tolerance

    vs Biopython: ResearchOS 0.9303, reference 1 (Δ 0.0697 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

  • Sequence alignment · Long alignment, 600 bp homologous block in ~9 kb sequencesWithin tolerance

    vs Biopython: ResearchOS 0.9629, reference 1 (Δ 0.0371 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

  • Sequence alignment · Short homology, 250 bp block (approximation gap grows)Within tolerance

    vs Biopython: ResearchOS 0.9737, reference 1 (Δ 0.0263 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

  • Sequence alignment · Long alignment, 800 bp homologous block in ~10 kb sequencesWithin tolerance

    vs Biopython: ResearchOS 0.9926, reference 1 (Δ 0.0074 identity)

    Our shared-region finder is a BLAST-style seed-and-extend, not an exact local aligner. It recovers the homologous block but reports identity a few percent below Biopython's exact local alignment because it includes some boundary bases, and the gap grows on shorter blocks. This is a real, expected limitation of the approximate method, not a bug.

Phylogenetic tree layout · 3 offsets

  • Phylogenetic tree layout · Candida auris global epidemiology (305 tips)Within tolerance

    vs ggtree: ResearchOS 0.993941, reference 1 (Δ 0.006059 1 - corr)

    ggtree and our renderer differ in scale, pixel sizing, and y-axis orientation, so a pixel-identical claim would be dishonest. What must agree is the topology-invariant structure both tools draw, the tip ordering and the relative branch-length depth of every node. We compare the absolute Spearman correlation of tip order (orientation-invariant) and require it within 0.02 of a perfect 1.0. A tip reordering or a depth bug would drop it well past the warn line.

  • Phylogenetic tree layout · HPV58 phylogeny with bootstrap support (90 tips)Within tolerance

    vs ggtree: ResearchOS 0.997037, reference 1 (Δ 0.002963 1 - corr)

    ggtree and our renderer differ in scale, pixel sizing, and y-axis orientation, so a pixel-identical claim would be dishonest. What must agree is the topology-invariant structure both tools draw, the tip ordering and the relative branch-length depth of every node. We compare the absolute Spearman correlation of tip order (orientation-invariant) and require it within 0.02 of a perfect 1.0. A tip reordering or a depth bug would drop it well past the warn line.

  • Phylogenetic tree layout · Human Microbiome Project tree (333 tips)Within tolerance

    vs ggtree: ResearchOS 0.99926, reference 1 (Δ 0.00074 1 - corr)

    ggtree and our renderer differ in scale, pixel sizing, and y-axis orientation, so a pixel-identical claim would be dishonest. What must agree is the topology-invariant structure both tools draw, the tip ordering and the relative branch-length depth of every node. We compare the absolute Spearman correlation of tip order (orientation-invariant) and require it within 0.02 of a perfect 1.0. A tip reordering or a depth bug would drop it well past the warn line.

Data Hub statistics · 68 offsets

  • Data Hub statistics · Cox PH, Harrell concordance (c-index)Within tolerance

    vs lifelines: ResearchOS 0.683398, reference 0.684512 (Δ 0.001114317 c)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, Wald zWithin tolerance

    vs statsmodels: ResearchOS 3.439729, reference 3.439876 (Δ 0.000146967 z)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · 5PL dose-response, asymmetry exponent SWithin tolerance

    vs SciPy: ResearchOS 1.236222, reference 1.236331 (Δ 0.000109374 S)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · 4PL dose-response, Bottom plateauWithin tolerance

    vs SciPy: ResearchOS 4.708365, reference 4.708439 (Δ 0.000073793 Bottom)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, z statisticWithin tolerance

    vs lifelines: ResearchOS -3.103032, reference -3.10298 (Δ 0.000051945 z)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, ROC AUC of fitted probabilitiesWithin tolerance

    vs SciPy: ResearchOS 0.84375, reference 0.8438 (Δ 0.00005 AUC)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, McFadden pseudo-R-squaredWithin tolerance

    vs statsmodels: ResearchOS 0.28965, reference 0.2896 (Δ 0.000049567 R2)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Global fit, shared Bottom plateau (least_squares)Within tolerance

    vs SciPy: ResearchOS -0.076117, reference -0.07607039616151008 (Δ 0.000046198 Bottom)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, slope Wald pWithin tolerance

    vs statsmodels: ResearchOS 0.024845, reference 0.0248 (Δ 0.000044527 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, coefficient (log hazard ratio, Treatment vs Control)Within tolerance

    vs lifelines: ResearchOS -1.370846, reference -1.370812 (Δ 0.000033886 coef)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Tukey HSD A vs B, mean difference (magnitude)Within tolerance

    vs statsmodels: ResearchOS 1.063333, reference 1.0633 (Δ 0.000033333 diff)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, 95% CI upper boundWithin tolerance

    vs statsmodels: ResearchOS 1.896844, reference 1.896814 (Δ 0.000029908 CI)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, 95% CI lower boundWithin tolerance

    vs statsmodels: ResearchOS 0.519823, reference 0.519852 (Δ 0.000029242 CI)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Tukey HSD A vs C, adjusted pWithin tolerance

    vs statsmodels: ResearchOS 0.002374, reference 0.0024 (Δ 0.000025855 p)

    Both compute the studentized-range adjusted p. statsmodels evaluates the range distribution via psturng (a tabulated approximation) and ours via a numeric integral, so the adjusted p can differ in the third decimal. The A vs B and B vs C comparisons are pinned only on the mean difference because statsmodels clamps their adjusted p to exactly 0.

  • Data Hub statistics · Cox PH, hazard ratio 95% CI upperWithin tolerance

    vs lifelines: ResearchOS 0.603517, reference 0.603534 (Δ 0.000016537 HR)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, between-subgroup variance (sigma_u^2)Within tolerance

    vs statsmodels: ResearchOS 0.180417, reference 0.180401 (Δ 0.000015715 variance)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, group difference SEWithin tolerance

    vs statsmodels: ResearchOS 0.351287, reference 0.351272 (Δ 0.000015361 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, slope standard errorWithin tolerance

    vs statsmodels: ResearchOS 0.246414, reference 0.2464 (Δ 0.000014397 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, intercept (b0)Within tolerance

    vs statsmodels: ResearchOS -2.271012, reference -2.271 (Δ 0.000012463 b0)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, hazard ratio exp(coef)Within tolerance

    vs lifelines: ResearchOS 0.253892, reference 0.253901 (Δ 0.000008895 HR)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear mixed model, between-subject variance (sigma_u^2)Within tolerance

    vs statsmodels: ResearchOS 0.086333, reference 0.086325 (Δ 0.000008286 variance)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Logistic regression, slope (b1)Within tolerance

    vs statsmodels: ResearchOS 0.552907, reference 0.5529 (Δ 0.000007477 b1)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear mixed model, intercept SEWithin tolerance

    vs statsmodels: ResearchOS 0.124276, reference 0.12427 (Δ 0.000005648 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, hazard ratio 95% CI lowerWithin tolerance

    vs lifelines: ResearchOS 0.106809, reference 0.106814 (Δ 0.000004827 HR)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · 4PL dose-response, Hill slopeWithin tolerance

    vs SciPy: ResearchOS 0.930921, reference 0.930926 (Δ 0.000004707 Hill)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Global fit, shared Hill slopeWithin tolerance

    vs SciPy: ResearchOS 1.014551, reference 1.0145544554504538 (Δ 0.000003608 Hill)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, coefficient standard errorWithin tolerance

    vs lifelines: ResearchOS 0.441776, reference 0.441773 (Δ 0.000003272 se)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear mixed model, condition Q SEWithin tolerance

    vs statsmodels: ResearchOS 0.045947, reference 0.045948 (Δ 0.000001165 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear mixed model, condition R SEWithin tolerance

    vs statsmodels: ResearchOS 0.045947, reference 0.045948 (Δ 0.000001165 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear mixed model, residual variance (sigma_e^2)Within tolerance

    vs statsmodels: ResearchOS 0.006333, reference 0.006334 (Δ 6.65e-7 variance)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, two-sided pWithin tolerance

    vs lifelines: ResearchOS 0.001915, reference 0.001916 (Δ 5.1e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Normal QQ plot, reference-line slope (probplot least-squares fit)Within tolerance

    vs SciPy: ResearchOS 0.289656, reference 0.289656 (Δ 4.9e-7 slope)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Paired t-test, pWithin tolerance

    vs SciPy: ResearchOS 0.0002, reference 0.0002 (Δ 4.83e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested one-way ANOVA, within-subgroup variance (residual)Within tolerance

    vs SciPy: ResearchOS 0.018056, reference 0.018056 (Δ 4.44e-7 variance)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Mann-Whitney U, one-sided (less) pWithin tolerance

    vs SciPy: ResearchOS 0.004057, reference 0.004057 (Δ 4.41e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Log-rank test, pWithin tolerance

    vs lifelines: ResearchOS 0.000893, reference 0.000893 (Δ 4.21e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · One-way ANOVA, Holm-Sidak A vs C adjusted pWithin tolerance

    vs statsmodels: ResearchOS 0.000876, reference 0.000876 (Δ 3.95e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Brown-Forsythe (median-centered), WWithin tolerance

    vs SciPy: ResearchOS 0.072115, reference 0.072115 (Δ 3.85e-7 W)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Residual plot, first residual (statsmodels OLS resid[0])Within tolerance

    vs statsmodels: ResearchOS 0.066667, reference 0.066667 (Δ 3.33e-7 resid)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Residual plot, last residual (statsmodels OLS resid[-1])Within tolerance

    vs statsmodels: ResearchOS 0.083333, reference 0.083333 (Δ 3.33e-7 resid)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Fisher exact test (2x2), two-sided pWithin tolerance

    vs SciPy: ResearchOS 0.000112, reference 0.000112 (Δ 3.19e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested one-way ANOVA, pWithin tolerance

    vs SciPy: ResearchOS 0.034137, reference 0.034137 (Δ 2.99e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested t-test, two-sided pWithin tolerance

    vs statsmodels: ResearchOS 0.000582, reference 0.000582 (Δ 2.97e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Kruskal-Wallis, pWithin tolerance

    vs SciPy: ResearchOS 0.000744, reference 0.000744 (Δ 2.93e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Linear regression, interceptWithin tolerance

    vs SciPy: ResearchOS 0.035714, reference 0.035714 (Δ 2.86e-7 intercept)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Shapiro-Wilk, pWithin tolerance

    vs SciPy: ResearchOS 0.234207, reference 0.234207 (Δ 2.65e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Friedman, pWithin tolerance

    vs SciPy: ResearchOS 0.002479, reference 0.002479 (Δ 2.48e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Nested one-way ANOVA, between-subgroup variance componentWithin tolerance

    vs SciPy: ResearchOS 0.172222, reference 0.172222 (Δ 2.22e-7 variance)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Chi-square test (2x2, Yates-corrected), pWithin tolerance

    vs SciPy: ResearchOS 0.000141, reference 0.000141 (Δ 1.89e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · One-way ANOVA, Bonferroni A vs C adjusted pWithin tolerance

    vs statsmodels: ResearchOS 0.002629, reference 0.002629 (Δ 1.84e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Two-way ANOVA, interaction pWithin tolerance

    vs statsmodels: ResearchOS 0.017043, reference 0.017043 (Δ 1.82e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Chi-square test (2x3, uncorrected), pWithin tolerance

    vs SciPy: ResearchOS 0.010343, reference 0.010343 (Δ 1.73e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Multiple regression, x1 slope standard errorWithin tolerance

    vs statsmodels: ResearchOS 0.095008, reference 0.095008 (Δ 1.71e-7 SE)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Gehan-Breslow-Wilcoxon test, pWithin tolerance

    vs lifelines: ResearchOS 0.001193, reference 0.001193 (Δ 1.5e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · One-way ANOVA, Sidak A vs C adjusted pWithin tolerance

    vs statsmodels: ResearchOS 0.002627, reference 0.002627 (Δ 1.2e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Mann-Whitney U, p (asymptotic, continuity-corrected)Within tolerance

    vs SciPy: ResearchOS 0.008113, reference 0.008113 (Δ 1.17e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Student unpaired t-test, pWithin tolerance

    vs SciPy: ResearchOS 0.000111, reference 0.000111 (Δ 1.03e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · From-stats Student t-test, pWithin tolerance

    vs SciPy: ResearchOS 0.000111, reference 0.000111 (Δ 1.03e-7 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Bootstrap BCa jackknife acceleration of the mean on a fixed sampleWithin tolerance

    vs SciPy: ResearchOS 0.03867, reference 0.03867 (Δ 9.8e-8 a)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Cox PH, likelihood-ratio pWithin tolerance

    vs lifelines: ResearchOS 0.000772, reference 0.000772 (Δ 8.1e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Two-way ANOVA, Time pWithin tolerance

    vs statsmodels: ResearchOS 0.000057, reference 0.0000571 (Δ 4.4e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Paired t-test, one-sided (less) pWithin tolerance

    vs SciPy: ResearchOS 0.0001, reference 0.0000998 (Δ 4.1e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Welch unpaired t-test, one-sided (less) pWithin tolerance

    vs SciPy: ResearchOS 0.000048, reference 0.0000476 (Δ 4e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · From-stats Welch t-test, one-sided (less) pWithin tolerance

    vs SciPy: ResearchOS 0.000048, reference 0.0000476 (Δ 4e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Welch unpaired t-test, pWithin tolerance

    vs SciPy: ResearchOS 0.000095, reference 0.0000953 (Δ 1.9e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · From-stats Welch t-test, pWithin tolerance

    vs SciPy: ResearchOS 0.000095, reference 0.0000953 (Δ 1.9e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Chi-square test (2x2, uncorrected), pWithin tolerance

    vs SciPy: ResearchOS 0.000056, reference 0.0000558 (Δ 1.4e-8 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

  • Data Hub statistics · Repeated-measures ANOVA, Huynh-Feldt corrected pWithin tolerance

    vs Pingouin: ResearchOS 0.000001, reference 0.00000116 (Δ 5e-9 p)

    ResearchOS computes this statistic by the same definition as the reference tool, so it must agree to numerical precision on this fixed dataset. Any drift beyond the tolerance is an engine regression.

Protein parameters · 1 offset

  • Protein parameters · Amyloid-beta N-terminus (acidic, low aliphatic)Within tolerance

    vs Biopython: ResearchOS 23.4313, reference 23.43125 (Δ 0.0001 )

    ResearchOS ports the Biopython ProtParam algorithm with its verbatim constant tables, so this value must match Biopython to floating-point precision.

Separately, the melting-temperature method shows the simpler Wallace and GC-percent rules as context. Those are different methods, not a target to match, so they diverge from nearest-neighbor by several degrees and are labelled as context, not counted toward the totals.

Primer melting temperature (Tm)

Within documented tolerance

Melting temperature is computed from nearest-neighbor thermodynamics (Allawi and SantaLucia 1997 parameters with the SantaLucia 1998 entropy salt correction) given the primer sequence, monovalent and divalent ion concentrations, and oligo concentration. Values are compared against Biopython Tm_NN under identical parameters and against primer3. The simpler Wallace and GC-percent rules are shown as context, to make the spread between Tm methods visible.

Tested module frontend/src/lib/calculators/tm-nn.ts

18-mer, GC termini both ends (where tables differ most) ResearchOS 55.8882 C ≈ primer3-py 56.8985 C (Δ 1.0103 C)

exact agreement (y = x)1111343457577979Reference value (C)ResearchOS value (C)15-mer, ~50% GC vs biopython: ours 46.5572 C, reference 46.5572 C15-mer, ~50% GC vs primer3: ours 46.5572 C, reference 46.5572 C16-mer, very low GC, AT termini vs biopython: ours 26.4155 C, reference 26.4155 C16-mer, very low GC, AT termini vs primer3: ours 26.4155 C, reference 26.4155 C16-mer, very high GC, GC termini vs biopython: ours 71.1344 C, reference 71.1344 C16-mer, very high GC, GC termini vs primer3: ours 71.1344 C, reference 71.1344 C25-mer, realistic primer vs biopython: ours 57.9787 C, reference 57.9787 C25-mer, realistic primer vs primer3: ours 57.9787 C, reference 57.9787 C28-mer, Biopython reference oligo vs biopython: ours 61.924 C, reference 61.924 C28-mer, Biopython reference oligo vs primer3: ours 61.924 C, reference 61.924 C40-mer, high GC vs biopython: ours 84.5607 C, reference 84.5607 C40-mer, high GC vs primer3: ours 84.5607 C, reference 84.5607 C18-mer, GC termini both ends (where tables differ most) vs biopython: ours 55.8882 C, reference 55.8882 C18-mer, GC termini both ends (where tables differ most) vs primer3: ours 55.8882 C, reference 56.8985 C28-mer, full PCR buffer (Na + Mg + dNTP) vs biopython: ours 66.3099 C, reference 66.3099 C28-mer, full PCR buffer (Na + Mg + dNTP) vs primer3: ours 66.3099 C, reference 66.3099 C8-mer palindrome (EcoRV-like), self-complementary vs biopython: ours 6.185 C, reference 6.185 C12-mer palindrome, self-complementary vs biopython: ours 31.8009 C, reference 31.8009 C17-mer M13 forward primer vs biopython: ours 51.6742 C, reference 51.6742 C17-mer M13 forward primer vs primer3: ours 51.6742 C, reference 51.6742 C20-mer T7 promoter primer vs biopython: ours 46.7766 C, reference 46.7766 C20-mer T7 promoter primer vs primer3: ours 46.7766 C, reference 46.7766 C22-mer, very high GC vs biopython: ours 71.6958 C, reference 71.6958 C22-mer, very high GC vs primer3: ours 71.6958 C, reference 71.6958 C20-mer, AT-rich vs biopython: ours 42.3846 C, reference 42.3846 C20-mer, AT-rich vs primer3: ours 42.3846 C, reference 42.3846 C30-mer, mixed composition vs biopython: ours 62.7326 C, reference 62.7326 C30-mer, mixed composition vs primer3: ours 62.7326 C, reference 62.7326 C24-mer M13 reverse primer vs biopython: ours 55.0381 C, reference 55.0381 C24-mer M13 reverse primer vs primer3: ours 55.0381 C, reference 55.0381 C13-mer, short oligo vs biopython: ours 42.1761 C, reference 42.1761 C13-mer, short oligo vs primer3: ours 42.1761 C, reference 42.1761 C45-mer, long oligo vs biopython: ours 70.2062 C, reference 70.2062 C45-mer, long oligo vs primer3: ours 70.2062 C, reference 70.2062 C24-mer M13 reverse, with 50 mM K+ added vs biopython: ours 56.6147 C, reference 56.6147 C24-mer M13 reverse, with 50 mM K+ added vs primer3: ours 56.6147 C, reference 56.6147 C24-mer M13 reverse, full PCR buffer (Mg + dNTP) vs biopython: ours 59.1141 C, reference 59.1141 C24-mer M13 reverse, full PCR buffer (Mg + dNTP) vs primer3: ours 59.1141 C, reference 59.1141 C10-mer NdeI-centered palindrome, self-complementary vs biopython: ours 29.3784 C, reference 29.3784 C8-mer XhoI palindrome, self-complementary vs biopython: ours 19.2883 C, reference 19.2883 C
Biopythonprimer3-py

Plus 12 cross-method context comparisons, shown for reference and not counted toward the totals.

References (4)
  • Biopythonv1.83

    Bio.SeqUtils.MeltingTemp.Tm_NN

    Allawi & SantaLucia 1997 (DNA_NN3), SantaLucia 1998 salt correction

    Reproduce: frontend/scripts/gen-tm-golden.py docs

  • primer3-pyv2.0.3

    primer3.calc_tm (tm_method='santalucia', salt_corrections_method='santalucia')

    SantaLucia 1998 unified nearest-neighbor table

    Reproduce: frontend/scripts/gen-tm-golden.py docs

  • Wallace rule (2+4)vn/a

    Bio.SeqUtils.MeltingTemp.Tm_Wallace

    Wallace et al. 1979, 4*GC + 2*AT; valid only for short oligos

    Reproduce: frontend/scripts/gen-tm-golden.py docs

  • GC% rulevn/a

    Bio.SeqUtils.MeltingTemp.Tm_GC

    Marmur-Doty / empirical GC-percent formula with salt term

    Reproduce: frontend/scripts/gen-tm-golden.py docs