/// AI HUB
Dashboard News Models Tools Papers Repos Videos Companies Trending
Login

#Benchmark

721 articles tagged with Benchmark

Latest Trending
Dev.to tutorial Apr 7

I published my benchmark scores. Your turn.

Back in March I released agent-egress-bench, a test corpus for evaluating security tools that sit...

Benchmark
12
Mastodon discussion Apr 7

I just read that Meta has an internal leaderboard for LLM token usage, and I really can't imagine a less thoughtful, les...

I just read that Meta has an internal leaderboard for LLM token usage, and I really can't imagine a less thoughtful, less useful status metric. Even assuming that the tokens are be...

LLM Benchmark
38
Mastodon discussion Apr 7

A 600-run benchmark by Yusuke Endoh tested #ClaudeCode across 13 #ProgrammingLanguages by implementing a simplified Git....

A 600-run benchmark by Yusuke Endoh tested #ClaudeCode across 13 #ProgrammingLanguages by implementing a simplified Git.Key takeaways:⇨ Ruby, Python & JavaScript were the fastest, ...

Benchmark
24
Papers with Code paper Apr 7

Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents

Large language models are increasingly deployed as autonomous agents executing multi-step workflows in real-world software environments. However, existing agent benchmarks suffer f...

Benchmark
21
Papers with Code paper Apr 7

MedConclusion: A Benchmark for Biomedical Conclusion Generation from Structured Abstracts

Large language models (LLMs) are widely explored for reasoning-intensive research tasks, yet resources for testing whether they can infer scientific conclusions from structured bio...

Benchmark
21
Mastodon discussion Apr 7

Microsoft udostępnia trzy autorskie modele MAI w ramach platformy Foundry. Wśród nowości znalazł się lider benchmarków t...

Microsoft udostępnia trzy autorskie modele MAI w ramach platformy Foundry. Wśród nowości znalazł się lider benchmarków transkrypcji oraz narzędzia do błyskawicznej syntezy mowy i f...

Microsoft Benchmark
18
Mastodon discussion Apr 6

📰 LTX-Video 2.3 Benchmark: 11 Models Tested for AI Video GenerationLTX-Video 2.3 emerges as a leading AI video generatio...

📰 LTX-Video 2.3 Benchmark: 11 Models Tested for AI Video GenerationLTX-Video 2.3 emerges as a leading AI video generation framework, with 11 distinct models tested across hardware ...

Benchmark
9
Mastodon discussion Apr 6

📰 Gemma 4 Leads 2026 AI Leaderboard with $0.20 Inference Cost and 100% Survival RateGemma 4, a 31-billion-parameter mode...

📰 Gemma 4 Leads 2026 AI Leaderboard with $0.20 Inference Cost and 100% Survival RateGemma 4, a 31-billion-parameter model, has shattered benchmarks by achieving 1,144% median ROI a...

OpenAI Google Benchmark
9
Mastodon discussion Apr 6

📰 AI Memory Benchmark 2026: Chinese Teen Coders Break Records with 98.7% Referential Resolution Acc...A new AI memory be...

📰 AI Memory Benchmark 2026: Chinese Teen Coders Break Records with 98.7% Referential Resolution Acc...A new AI memory benchmark, pioneered by teenage developers from China, has ach...

Benchmark
18
Mastodon discussion Apr 6

Is AI just a benchmark race? Google’s Gemini 3.1 Pro says yes, but 1 million medical images processed in Galicia say the...

Is AI just a benchmark race? Google’s Gemini 3.1 Pro says yes, but 1 million medical images processed in Galicia say there’s a much deeper story. From doctor's offices to insurance...

Google Benchmark
18
Mastodon discussion Apr 6

Google DeepMind just scored 85% on ARC-AGI-2 — the most difficult general reasoning benchmark in AI.Previous best? 54%.T...

Google DeepMind just scored 85% on ARC-AGI-2 — the most difficult general reasoning benchmark in AI.Previous best? 54%.This kind of leap doesn't happen often. General reasoning is ...

Google Benchmark
18
Mastodon discussion Apr 6

new writeup: why eval pipelines matter more than model selection.teams with good evals ship 3x faster. teams without the...

new writeup: why eval pipelines matter more than model selection.teams with good evals ship 3x faster. teams without them are flying blind, afraid to change anything because they c...

Benchmark
30
Papers with Code paper Apr 6

DeonticBench: A Benchmark for Reasoning over Rules

Reasoning with complex, context-specific rules remains challenging for large language models (LLMs). In legal and policy settings, this manifests as deontic reasoning: reasoning ab...

Benchmark
21
Mastodon discussion Apr 5

Title: P2: P2: P2: P2: How LLM sees own adaptive thinking and evolution..5. Metacognition and Self-Reflection- Self-Eval...

Title: P2: P2: P2: P2: How LLM sees own adaptive thinking and evolution..5. Metacognition and Self-Reflection- Self-Evaluation :: Assesses own performance, outputs, and reasoning.-...

LLM Benchmark
18
Dev.to tutorial Apr 5

Stop Vibing, Start Eval-ing: EDD for AI-Native Engineers

When I was doing traditional development, I had TDD. I wrote a test, it passed or failed, done. But...

Benchmark
12
Mastodon discussion Apr 5

NEW: Why Most AI Benchmarks Are Lying to You (And What to Measure Instead)MMLU is saturated. Contamination is rampant. C...

NEW: Why Most AI Benchmarks Are Lying to You (And What to Measure Instead)MMLU is saturated. Contamination is rampant. Companies cherry-pick evals.What actually predicts production...

Benchmark
18
Dev.to tutorial Apr 3

Which AI models are actually "brain-like"? I built an open-source benchmark to measure it

Meta released TRIBE v2 last week - a foundation model that predicts fMRI brain activation from video,...

Open Source Benchmark
12
Mastodon discussion Apr 3

#Gemma4 is the first model to pass the techno vibe benchmark. Thank you @Google #vibes #ai #google #llm #agi #techno #mu...

#Gemma4 is the first model to pass the techno vibe benchmark. Thank you @Google #vibes #ai #google #llm #agi #techno #music

Google LLM Benchmark
27
Mastodon discussion Apr 3

How creative are AI scientists? A new benchmark evaluates idea generation across originality, feasibility, and flexibili...

How creative are AI scientists? A new benchmark evaluates idea generation across originality, feasibility, and flexibility, revealing gaps between reasoning skills and scientific c...

Benchmark
24
Mastodon discussion Apr 3

📰 Qwen 3.6 Beats GPT-4o and Claude 3.5 in China’s Blind AI Coding Benchmark (2026)Qwen 3.6 has emerged as China's leadin...

📰 Qwen 3.6 Beats GPT-4o and Claude 3.5 in China’s Blind AI Coding Benchmark (2026)Qwen 3.6 has emerged as China's leading AI programming model after dominating a global blind bench...

Anthropic Multimodal Benchmark
9
Papers with Code paper Apr 3

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across ...

Benchmark
21
Papers with Code paper Apr 3

GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers

The autonomous discovery of bugs remains a significant challenge in modern software development. Compared to code generation, the complexity of dynamic runtime environments makes b...

Benchmark
21
Mastodon discussion Apr 2

📰 MLPerf Inference v6.0 2026: Nvidia, AMD, and Intel Break Records Amid Benchmark ChallengesMLPerf Inference v6.0 introd...

📰 MLPerf Inference v6.0 2026: Nvidia, AMD, and Intel Break Records Amid Benchmark ChallengesMLPerf Inference v6.0 introduces multimodal and video models, with Nvidia, AMD, and Inte...

NVIDIA Benchmark
9
Mastodon discussion Apr 2

Ghost hosting operator tested 16 coding models on custom TypeScript benchmark built from real commits. Open-weights Qwen...

Ghost hosting operator tested 16 coding models on custom TypeScript benchmark built from real commits. Open-weights Qwen3-Coder-Next scored 94.8% vs Claude Opus's 98.8% - gap close...

Anthropic Benchmark
18
« Previous Page 27 of 31 (721 items) Next »
AI Hub // AI Intelligence Platform // LIVE FEED // Impressum // Datenschutz © 2026
0 new articles available