ArXiv paper Mar 26

RenoBench: A Citation Parsing Benchmark

Accurate parsing of citations is necessary for machine-readable scholarly infrastructure. But, despite sustained interest in this problem, existing evaluation techniques are often ...

Papers with Code paper Mar 26

Robust Reasoning Benchmark

While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their underlying reasoning processes remain highly overfit to standard textual form...