
Hugging Face has launched BenchMIRT, a new method for analyzing LLM benchmarks at the prompt level. BenchMIRT uses multidimensional Item Response Theory to separate different capabilities, such as safety and general reasoning, that contribute to a model's performance. This approach helps clarify what benchmarks are truly measuring and can predict model performance on unseen questions with high accuracy. The tool provides a more detailed understanding of model capabilities, moving beyond simple aggregate scores.
Read original