Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions...

Germany's Soofi consortium withdrew benchmark scores after discovering its training data contained paraphrased questions from the GPQA evaluation set. The 11.1-point gain on GPQA-Diamond can no longer stand as evidence of model capability. A researcher spotted the contamination; the team confirmed it within a week. https://www.implicator.ai/soofi-gpqa-contamination-scores-withdrawn/ #ai #evaluation #openscience

Read Original

Related

Mastodon discussion 30m ago

Как делать скилы правильно и собрать из них маркетплейсНачалось всё с подборки промтов cursor-vibe-prompts , которую я т...

Как делать скилы правильно и собрать из них маркетплейсНачалось всё с подборки промтов cursor-vibe-prompts , которую я таскал из проекта в проект копипастой, и в какой-то момент ме...