We audited our own pattern-embedding evaluation and found 53% of held-out samples had same-symbol training neighbors within 20 days. Here's what we changed — and why agent developers should demand this kind of rigor from any historical-pattern API.
Eval Integrity: How We Found the Leakage and Why Our Baseline Lied