CDD Bench
This benchmark evaluates whether Convexia's custom model can prospectively rank clinical trial readout success better than base rates, simple baselines, and common diligence heuristics. The benchmark focuses on trial-level prediction and assesses performance through discrimination, calibration, and portfolio lift. The key finding is that the model appears most valuable as a diligence-ranking layer: it helps concentrate higher-probability opportunities toward the top of the review queue while still requiring human review for sparse, ambiguous, or distribution-shifted cases. The benchmark also defines the main limits of the benchmark, including public-data constraints, label ambiguity, sponsor disclosure bias, and the risk that strong predictive performance does not automatically translate into clinical, regulatory, or commercial success.
Benchmark at a glance
The headline result, the size and validity of the benchmark cohort, and the boundary between primary and supplemental claims.
4,636
resolved prediction events
36.0%
observed base rate
0.731
ranking under class imbalance
0.812
success versus failure separation
0.160
probability error
Cohort construction and temporal validity
How the benchmark population narrows from screened records to primary scored predictions, supplemental simulations, pending outcomes, and exclusions.
Cohort construction
Headline performance
The headline result with primary and supplemental cohorts kept clearly separated across the three major metrics.
Primary versus supplemental performance
Baseline comparison
How the model compares with tested baselines across AUPRC, AUROC, and Brier score, without overstating the claim.
Baseline comparison — AUPRC
Baseline comparison across AUPRC, AUROC, and Brier
Calibration and rank enrichment
Whether predicted probabilities match observed success rates, and whether the model concentrates successes at the top of the ranked list.
Calibration reliability
Rank enrichment by decile
Threshold operating points
How queue size, precision, recall, and false positive rate change across operating thresholds.
Threshold operating points
Robustness and stress tests
Which assumptions move the benchmark and where the model is weakest.
Robustness and stress tests
Subgroup diagnostics
Where the model performs better or worse, and the main error modes that need human review and monitoring.