The root cause: GPQA publishes its eval set in a single split labeled 'train' on Hugging Face, and Soofi's pipeline sele...

The root cause: GPQA publishes its eval set in a single split labeled 'train' on Hugging Face, and Soofi's pipeline selected by name rather than meaning. An audit found the same failure across four benchmarks. Future datasets now get allowlist validation and n-gram screening. https://www.implicator.ai/soofi-gpqa-contamination-scores-withdrawn/ #ai #benchmarking

Read Original

Related