SocialGrid Benchmark Shows LLMs Fail at Deception, Score Below 60% on PlanningResearchers introduced SocialGrid, a multi...

SocialGrid Benchmark Shows LLMs Fail at Deception, Score Below 60% on PlanningResearchers introduced SocialGrid, a multi-agent benchmark inspired by Among Us. It shows state-of-the-art LLMs fail at deception detection and task planning, scoring below 60% accuracy.https://gentic.news/article/socialgrid-benchmark-shows-llms#AI #ArtificialIntelligence #Tech

Read Original

Related