TY - UNPB
T1 - The Unreasonable Benchmark
AU - Wang, Ziyang
N1 - © 2026 F. Perez-Cruz et al..
PY - 2026/7/3
Y1 - 2026/7/3
N2 - Effective benchmarking is instrumental for the systematic evaluation of Large Language Models (LLMs) performance, providing vital insights in the path towards advancements in Artificial Intelligence (AI). Existing benchmarks, such as Humanity's Last Exam and FrontierMath, stress-test the boundaries of the model's performance. However, they provide limited understanding of a model's reliability in routine, low-difficulty scenarios. It is, however, these contexts, which, despite their simplicity, frequently expose persistent model weaknesses. To address this, we present the Dataset, a complementary benchmark aiming to address this oversight via systematic (and continuous) evaluation of LLMs' performance in basic reasoning and everyday tasks. It is the end result of a crowdsourcing effort seeking to ensure a diversity of perspectives and topic coverage. Moreover, the dataset is designed to be dynamic in nature; it incorporates new items as emerging failure modes are identified while retiring resolved items, thereby maintaining relevance over time. Similar to previous efforts, such as TruthfulQA, the primary objective of this new benchmark is to identify persistent shortcomings and support the development of robust and reliable LLMs. Ultimately, this process should converge upon a set of unreasonable issues, at which point we can be confident that the LLMs will indeed exhibit impressive and resilient capabilities. The set of (multimodal) questions is available at https://huggingface.co/datasets/unreasonablebenchmark/unreasonable-benchmark , under a Creative Commons Attribution-ShareAlike (CC BY-SA) license.
AB - Effective benchmarking is instrumental for the systematic evaluation of Large Language Models (LLMs) performance, providing vital insights in the path towards advancements in Artificial Intelligence (AI). Existing benchmarks, such as Humanity's Last Exam and FrontierMath, stress-test the boundaries of the model's performance. However, they provide limited understanding of a model's reliability in routine, low-difficulty scenarios. It is, however, these contexts, which, despite their simplicity, frequently expose persistent model weaknesses. To address this, we present the Dataset, a complementary benchmark aiming to address this oversight via systematic (and continuous) evaluation of LLMs' performance in basic reasoning and everyday tasks. It is the end result of a crowdsourcing effort seeking to ensure a diversity of perspectives and topic coverage. Moreover, the dataset is designed to be dynamic in nature; it incorporates new items as emerging failure modes are identified while retiring resolved items, thereby maintaining relevance over time. Similar to previous efforts, such as TruthfulQA, the primary objective of this new benchmark is to identify persistent shortcomings and support the development of robust and reliable LLMs. Ultimately, this process should converge upon a set of unreasonable issues, at which point we can be confident that the LLMs will indeed exhibit impressive and resilient capabilities. The set of (multimodal) questions is available at https://huggingface.co/datasets/unreasonablebenchmark/unreasonable-benchmark , under a Creative Commons Attribution-ShareAlike (CC BY-SA) license.
KW - benchmarks
KW - LLMs
KW - nreasonable benchmarks
UR - https://openreview.net/forum?id=EkWAAfbGrh
M3 - Preprint
T3 - Journal of Data-centric Machine Learning Research
BT - The Unreasonable Benchmark
ER -