Skip to the content.

Experiments and evidence status

Current status

The repository does not currently publish a learned-system benchmark. Earlier NDCG, recall, and micro-F1 values were withdrawn after a repository audit found that the available code and artifacts could not establish the reported protocol. The values must not be copied into papers, posts, repository metadata, or application material.

The repair separates two official Amazon ESCI populations:

The ranking evaluator now requires the bi-encoder to select candidates before reranking, computes ideal DCG from the complete qrels, and assigns zero to missing queries. Hard-negative mining and the second training pass now load the explicit warmed checkpoint instead of silently restarting from the base model.

These repairs improve evaluation integrity; they do not create new benchmark evidence. A result becomes publishable only through the bundle below.

Minimum rerun plan

1. Freeze the inputs

2. Establish Task 1 ranking baselines

3. Establish Task 2 classification

4. Run learned systems repeatedly

5. Invite independent reproduction

See reproducibility.md and the result-bundle contract.