Benchmark report2026-06-18

CDD Bench

This benchmark evaluates whether Convexia's custom model can prospectively rank clinical trial readout success better than base rates, simple baselines, and common diligence heuristics. The benchmark focuses on trial-level prediction and assesses performance through discrimination, calibration, and portfolio lift. The key finding is that the model appears most valuable as a diligence-ranking layer: it helps concentrate higher-probability opportunities toward the top of the review queue while still requiring human review for sparse, ambiguous, or distribution-shifted cases. The benchmark also defines the main limits of the benchmark, including public-data constraints, label ambiguity, sponsor disclosure bias, and the risk that strong predictive performance does not automatically translate into clinical, regulatory, or commercial success.

01

Benchmark at a glance

The headline result, the size and validity of the benchmark cohort, and the boundary between primary and supplemental claims.

PrimaryPrimary metric cards
Primary cohort

4,636

resolved prediction events

Success rate

36.0%

observed base rate

AUPRC

0.731

ranking under class imbalance

AUROC

0.812

success versus failure separation

Brier score

0.160

probability error

02

Cohort construction and temporal validity

How the benchmark population narrows from screened records to primary scored predictions, supplemental simulations, pending outcomes, and exclusions.

Figure 1

Cohort construction

03

Headline performance

The headline result with primary and supplemental cohorts kept clearly separated across the three major metrics.

Figure 2

Primary versus supplemental performance

04

Baseline comparison

How the model compares with tested baselines across AUPRC, AUROC, and Brier score, without overstating the claim.

Figure 3

Baseline comparison — AUPRC

Figure 4

Baseline comparison across AUPRC, AUROC, and Brier

05

Calibration and rank enrichment

Whether predicted probabilities match observed success rates, and whether the model concentrates successes at the top of the ranked list.

Figure 5

Calibration reliability

Figure 6

Rank enrichment by decile

06

Threshold operating points

How queue size, precision, recall, and false positive rate change across operating thresholds.

Figure 7

Threshold operating points

07

Robustness and stress tests

Which assumptions move the benchmark and where the model is weakest.

Figure 8

Robustness and stress tests

08

Subgroup diagnostics

Where the model performs better or worse, and the main error modes that need human review and monitoring.

Figure 9

Subgroup diagnostic heatmap

Figure 10

Error review tags