SciTables : A Dataset and Evaluation Framework for Complex Table-to-Text Generation
Abstract
Generating coherent and factually grounded text from structured data is a core challenge in natural language generation, with appli- cations in scientific communication, medical documentation, and automated reporting. Existing datasets primarily focus on open- domain or simplified table formats, limiting progress in more com- plex, high-stakes domains. We present SciTables, a new dataset and evaluation framework for scientific table-to-text generation, addressing the gap in existing resources that focus largely on open-domain or simplified tables. Our dataset is constructed from Computer Science papers on arXiv (2017–2023) and features complex tables rich in numeric, symbolic, and mathematical content paired with naturally occurring textual descriptions. We develop a scalable, semi-automated pipeline to extract, clean, and align tables with their associated text, preserving domain-specific language while minimizing annotation cost. The resulting benchmark poses realistic challenges for current models and supports evaluation beyond semantic similarity, including factual accuracy, relevance, and multiple forms of reasoning. We conduct extensive experiments with state-of-the-art generation models and show that while current models achieve strong semantic alignment with reference descriptions, they struggle with higher-order reasoning, aggregation, and factual grounding as table complexity increases. Our work provides a realistic and scalable benchmark for advancing faithful, informative, and reasoning-aware table-to-text generation in scientific domains.
Assigned reviewers
No reviewers assigned yet.
Candidates from the panel ranked by taxonomy affinity
| # | Reviewer | Match | Load | Why |
|---|