Skip to content
Methods

Benchmark protocol

OpenDeco defines the question, eligible evidence, model identity, evaluation mode and analysis plan before a result becomes a benchmark release.

Scope and eligibility

A release defines which domains, datasets, model implementations and configurations are eligible. Exclusions are recorded with reasons so the analysis cannot quietly narrow itself around a preferred result.

Evaluation modes

Replay, compliance and counterfactual analyses answer different questions. The mode is labelled explicitly, and a counterfactual schedule is not presented as though it had an observed outcome.

Calibration overlap

Known use of an evaluation dataset in model development is recorded before scoring. Results on overlapping data can remain visible, but they are not described as independent validation.

Metrics

Metrics are chosen for the scientific question and versioned with the protocol. No single league-table number is used to conceal calibration, discrimination, burden, uncertainty or failure-region behaviour.

Failure regions

Aggregate performance is accompanied by inspection of exposure regions where errors or model divergence concentrate. Those regions are research objects. They are not hidden as inconvenient outliers.

Sensitivity

Configuration sweeps and analytical choices are tested where they can materially change the conclusion. A finding that disappears under reasonable settings is reported as fragile.

External and leave-study-out validation

When multiple independent study sources exist, OpenDeco uses held-out or leave-study-out analysis where scientifically appropriate. This is distinguished from random record splitting within one historical programme.

Release

A public release is built from frozen inputs and immutable run manifests, checked for provenance and restricted-data leakage, rerun cleanly and assigned stable version metadata.