2026 年 33 巻 3 号 p. 1555-1590
LLMs excel at well-known table-based question answering (Table QA) benchmarks, yet their ability to reason over multi-entity time-series tables, common in economics and the social sciences, remains underexplored. We present METSQA, a benchmark of approximately 46,000 QA pairs constructed from synthetic and public multi-entity time-series tables. METSQA varies across five table factors: size, shape, value orientation, timestamp temporal order, and numeric scale (digits). The benchmark encompasses ten numerical operations and their combinations, enabling fine-grained, factorized evaluation. Using established prompting protocols, we evaluate seven proprietary models and three open-weight models. Our results show that, within the factor ranges tested, accuracy is most sensitive to table size and also declines with larger numeric values, while temporal factors have more localized, model-dependent effects. Aggregation is the most challenging operation, while performance on other operations varies across models. Beyond serving as a fair testbed, METSQA points to concrete research directions, including scale robust planning and improved aggregation reasoning for multi-entity time-series QA.