Haoyu Dong 0001

dblp:180/5089-1 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0003-0692-2228ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 18 since 2021Databases, data management, data science and information retrieval · 6 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided Search
abstract
Large language models (LLMs) have shown great promise in automating data science workflows. However, existing models still struggle with multi-step reasoning and tool use, limiting their effectiveness on complex data analysis tasks. To address this limitation, we propose a scalable pipeline that extracts high-quality, tool-based data analysis tasks and their executable multi-step solutions from real-world Jupyter notebooks and associated data files. Using this pipeline, we introduce NbQA, a large-scale dataset of standardized task–solution pairs that reflect authentic tool-use patterns in practical data science scenarios. To further enhance the multi-step reasoning capabilities, we present Jupiter, a framework that formulates data analysis as a search problem and applies Monte Carlo Tree Search (MCTS) to generate diverse solution trajectories for value model learning. During inference, Jupiter combines the value model and node visit counts to efficiently collect executable multi-step plans with minimal search steps. Experimental results show that Qwen2.5-7B and 14B-Instruct models on NbQA solve 77.82% and 86.38% of tasks on InfiAgent-DABench, respectively—matching or surpassing GPT-4o and advanced agent frameworks. Further evaluations demonstrate improved generalization and stronger tool-use reasoning across diverse multi-step reasoning tasks.
Shuocheng Li, Silin Du, Wenxuan Zeng, Mengyu Zhou, Yeye He, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001
AAAI8
2026 SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large Spreadsheets
abstract
Understanding and reasoning over complex spreadsheets remain fundamental challenges for large language models (LLMs), which often struggle with intricate structures and rely solely on neural computation. In this work, we propose SheetBrain, a neuro-symbolic dual-workflow agent framework for precise and interpretable reasoning over tabular data. SheetBrain consists of an understanding module that produces a comprehensive overview of the spreadsheet, including structural summaries and query-specific analyses to guide execution; an execution module that integrates a Python sandbox with preloaded table-processing libraries and an Excel helper toolkit for effective data manipulation; and a validation module that verifies the correctness of reasoning and answers, triggering re-execution if necessary. We evaluate SheetBrain on multiple public QA and manipulation benchmarks, and introduce SheetBench, a new benchmark targeting large, multi-table, and structurally complex spreadsheets. Experimental results show that SheetBrain significantly improves reasoning performance on both existing benchmarks and the more challenging scenarios presented in SheetBench.
Jiayuan Su, Mengyu Zhou, Huaxing Zeng, Mengni Jia, Haoyu Dong 0001, Xiaojun Ma 0001, Shi Han, Dongmei Zhang 0001
AAAI7
2026 Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
abstract
Hanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong, Xiaojun Ma, Shi Han, Dongmei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong 0001, Xiaojun Ma 0001, Shi Han, Dongmei Zhang 0001
ACL (1)5
2025 RelationalCoder: Rethinking Complex Tables via Programmatic Relational Transformation
abstract
Semi-structured tables, with their varied layouts and formatting artifacts, remain a major obstacle for automated data processing and analytics.To address these challenges, we propose RELATIONALCODER, which uniformly converts semi-structured tables into relational data, enabling smooth integration with the rich ecosystem of data processing and analytics tools.By leveraging SQL code, RELATIONAL-CODER prevents schema errors and markedly improves normalization quality across multiple relational tables.
Haoyu Dong 0001, Yue Hu 0002, Huailiang Peng, Yanan Cao 0001
ACL (1)1
2025 TableLoRA: Low-rank Adaptation on Table Structure Understanding for Large Language Models
abstract
Tabular data are crucial in many fields and their understanding by large language models (LLMs) under high parameter efficiency paradigm is important. However, directly applying parameter-efficient fine-tuning (PEFT) techniques to tabular tasks presents significant challenges, particularly in terms of better table serialization and the representation of two-dimensional structured information within a one-dimensional sequence. To address this, we propose TableLoRA, a module designed to improve LLMs’ understanding of table structure during PEFT. It incorporates special tokens for serializing tables with special token encoder and uses 2D LoRA to encode low-rank information on cell positions. Experiments on four tabular-related datasets demonstrate that TableLoRA consistently outperforms vanilla LoRA and surpasses various table encoding methods tested in control experiments. These findings reveal that TableLoRA, as a table-specific LoRA, enhances the ability of LLMs to process tabular data effectively, especially in low-parameter settings, demonstrating its potential as a robust solution for handling table-related tasks.
Mengyu Zhou, Yeye He, Haoyu Dong 0001, Shi Han, Zejian Yuan, Dongmei Zhang 0001
ACL (1)5
2025 Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuning
abstract
Language models such as GPT and Llama have shown remarkable ability on diverse natural language tasks, yet their performance on complex table tasks (e.g., NL-to-Code, data cleaning, etc.) continues to be suboptimal.To improve their performance, task-specific fine-tuning is often needed, which, however, require expensive human labeling and is prone to over-fitting.In this work, we propose TABLE-SPECIALIST, a self-trained fine-tuning paradigm specifically designed for table tasks.Our insight is that for each table task, there often exist two dual versions of the same task, one generative and one classification in nature.Leveraging their duality, we propose a Generator-Validator paradigm to iteratively generate-then-validate training data from language models, to finetune stronger TABLE-SPECIALIST models that can specialize in a given task, without using manually-labeled data.
Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001, Surajit Chaudhuri
EMNLP4
2025 MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
abstract
Tables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally require expert-level users like data engineers, data analysts, and database administrators to operate. Although LLMs have shown remarkable progress in working with tables (e.g., in spreadsheet and database copilot scenarios), comprehensive benchmarking of such capabilities remains limited. In contrast to an extensive and growing list of NLP benchmarks, evaluations of table-related tasks are scarce, and narrowly focus on tasks like NL-to-SQL and Table-QA, overlooking the broader spectrum of real-world tasks that professional users face. This gap limits our understanding and model progress in this important area.In this work, we introduce MMTU, a large-scale benchmark with over 28K questions across 25 real-world table tasks, designed to comprehensively evaluate models ability to understand, reason, and manipulate real tables at the expert-level. These tasks are drawn from decades’ worth of computer science research on tabular data, with a focus on complex table tasks faced by professional users. We show that MMTU require a combination of skills -- including table understanding, reasoning, and coding -- that remain challenging for today's frontier models, where even frontier reasoning models like OpenAI GPT-5 and DeepSeek R1 score only around 69% and 57% respectively, suggesting significant room for improvement. We highlight key findings in our evaluation using MMTU and hope this benchmark drives further advances in understanding and developing foundation models for structured data processing and analysis.Our code and data are available at https://github.com/MMTU-Benchmark/MMTU and https://huggingface.co/datasets/MMTU-benchmark/MMTU.
Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong 0001, Shi Han, Lingjiao Chen, Dongmei Zhang 0001, Surajit Chaudhuri, H. V. Jagadish
NeurIPS4
2025 Reasoning and Retrieval for Complex Semi-structured Tables via Reinforced Relational Data Transformation
abstract
We introduce TabFormer, a framework that normalizes diverse semi-structured tables into relational data via large language models to facilitate various table retrieval and reasoning tasks. Our approach employs a chain-of-thought methodology, transforming one or multiple tables through a sequence of soft operations. Compared to existing operators that are sensitive and brittle to human-induced artifacts in real-world tables, soft operators are designed with greater flexibility to accommodate diverse formatting variations.
Haoyu Dong 0001, Yue Hu 0002, Yanan Cao 0001
SIGIR1
2024 KET-QA: A Dataset for Knowledge Enhanced Table Question Answering
abstract
Due to the concise and structured nature of tables, the knowledge contained therein may be incomplete or missing, posing a significant challenge for table question answering (TableQA) systems. However, most existing datasets either overlook the challenge of missing knowledge in TableQA or only utilize unstructured text as supplementary information for tables. In this paper, we propose to use a knowledge base (KB) as the external knowledge source for TableQA and construct a dataset KET-QA with fine-grained gold evidence annotation. Each table in the dataset corresponds to a sub-graph of the entire KB, and every question requires the integration of information from both the table and the sub-graph to be answered. To extract pertinent information from the vast knowledge sub-graph and apply it to TableQA, we design a retriever-reasoner structured pipeline model. Experimental results demonstrate that our model consistently achieves remarkable relative performance improvements ranging from 1.9 to 6.5 times on EM scores across three distinct settings (fine-tuning, zero-shot, and few-shot), in comparison with solely relying on table information. However, even the best model achieves a 60.23% EM score, which still lags behind the human-level performance, highlighting the challenging nature of KET-QA for the question-answering community.
Mengkang Hu, Haoyu Dong 0001, Ping Luo 0002, Shi Han, Dongmei Zhang 0001
LREC/COLING2
2024 Encoding Spreadsheets for Large Language Models
abstract
Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Haoyu Dong 0001, Yuzhang Tian, Junyu Xiong, Mengyu Zhou, José Cambronero, Yeye He, Shi Han, Dongmei Zhang 0001
EMNLP1
2024 OpenTE: Open-Structure Table Extraction From Text
abstract
This paper presents an Open-Structure Table Extraction (OpenTE) task, which aims to extract a table with intrinsic semantic, calculational, and hierarchical structure from unstructured text. We devise a novel Identification-Extraction-Grounding (IEG) framework for language models (LMs) comprising three chaining steps: (1) identifying semantic and calculational relationships among columns, (2) extracting structured data from unstructured text, and (3) aligning extracted data with the source text and the table structure with a separate discrete grounding model. Experiment results suggest that OpenTE presents a significant challenge for state-of-the-art LMs and demonstrate that the IEG framework achieves superior performance on both datasets, with over 9% F1 improvements in the few-shot setting for GPT-3.5&4 and other large language models (LLMs) and over 4.9% F1 enhancements in the fine-tuning setting for open-source BART. We’ll release the dataset to facilitate future research.
Haoyu Dong 0001, Mengkang Hu, Qinyu Xu, Yue Hu 0002
ICASSP1
2024 Large Language Models for Tabular Data: Progresses and Future Directions
abstract
Tables contain a significant portion of the world's structured information. The ability to efficiently and accurately understand, process, reason about, analyze, and generate tabular data is critical for achieving Artificial General Intelligence (AGI) systems. However, despite their prevalence and importance, tables present unique challenges due to their structured nature and the diverse semantics embedded within them. Textual content, numerical values, visual formats, and even formulas in tables carry rich semantic information that is often underutilized due to the complexity of accurately interpreting and integrating. Fortunately, the advent of Large Language Models (LLMs) has opened new frontiers in natural language processing (NLP) and machine learning (ML), showing remarkable success in understanding and generating text, code, etc. Applying these advanced models to the domain of tabular data holds the promise of significant breakthroughs in how we process and leverage structured information. Therefore, this tutorial aims to provide a comprehensive study of the advances, challenges, and opportunities in leveraging cutting-edge LLMs for tabular data. By introducing methods of prompting or training cutting-edge LLMs for table interpreting, processing, reasoning, analytics, and generation, we aim to equip researchers and practitioners with the knowledge and tools needed to unlock the full potential of LLMs for tabular data in their domains.
Haoyu Dong 0001, Zhiruo Wang 0001
SIGIR1
2024 TTC-QuAli: A Text-Table-Chart Dataset for Multimodal Quantity Alignment
abstract
In modern documents, numerical information is often presented using multimodal formats such as text, tables, and charts. However, the heterogeneity of these sources poses a challenge for machines attempting to jointly read and understand the numerical semantics conveyed through text, tables, and charts. In this paper, we introduce a multimodal dataset called Text-Table-Chart Quantity Alignment (TTC-QuAli). This dataset is designed to facilitate a new task that involves linking related quantities across text, tables, and charts. TTC-QuAli is a comprehensive dataset that contains 4,498 quantities in text, aligned with 1,086 chart images and 1,503 tables from real-world statistical reports. It is the first dataset to provide high-quality annotations for linking quantities across multiple modalities, and it includes challenging composite (aggregated/calculated) quantity linking. To address the challenge of bridging representation gaps between different modalities and capturing their shared contextual semantic meaning, we introduce ConTTC, a novel transformer-based cross-modal contrastive learning architecture. This is the first architecture to jointly model text, tables, and charts, and contrastive learning is employed for multimodal quantity linking towards unified representation learning. Our experiments demonstrate that TTC-QuAli presents a significant challenge for existing baselines and serves as a valuable benchmark for future research. Experiment results show that ConTTC significantly outperforms all baseline methods.
Haoyu Dong 0001, Anda Zhou, Yue Hu 0002
WSDM1
2023 SheetPT: Spreadsheet Pre-training Based on Hierarchical Attention Network
abstract
Spreadsheets are an important and unique type of business document for data storage, analysis and presentation. The distinction between spreadsheets and most other types of digital documents lies in that spreadsheets provide users with high flexibility of data organization on the grid. Existing related techniques mainly focus on the tabular data and are incompetent in understanding the entire sheet. On the one hand, spreadsheets have no explicit separation across tabular data and other information, leaving a gap for the deployment of such techniques. On the other hand, pervasive data dependence and semantic relations across the sheet require comprehensive modeling of all the information rather than only the tables. In this paper, we propose SheetPT, the first pre-training technique on spreadsheets to enable effective representation learning under this scenario. For computational effectiveness and efficiency, we propose the coherent chunk, an intermediate semantic unit of sheet structure; and we accordingly devise a hierarchical attention-based architecture to capture contextual information across different structural granularities. Three pre-training objectives are also designed to ensure sufficient training against millions of spreadsheets. Two representative downstream tasks, formula prediction and sheet structure recognition are utilized to evaluate its capability and the prominent results reveal its superiority over existing state-of-the-art methods.
Ran Jia, Qiyu Li 0001, Xiaoyuan Jin, Lun Du, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001
AAAI6
2023 HermEs: Interactive Spreadsheet Formula Prediction via Hierarchical Formulet Expansion
abstract
Wanrong He, Haoyu Dong, Yihuai Gao, Zhichao Fan, Xingzhuo Guo, Zhitao Hou, Xiao Lv, Ran Jia, Shi Han, Dongmei Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Wanrong He, Haoyu Dong 0001, Yihuai Gao, Zhichao Fan, Xingzhuo Guo, Zhitao Hou, Ran Jia, Shi Han, Dongmei Zhang 0001
ACL (1)2
2023 Personalized Educational Video Evaluation Combining Student's Cognitive and Teaching Style
abstract
AI-powered technologies, like ChatGPT and learning analytic technologies, have encouraged the sharing of online teaching resources and the transformation of teaching methods and learning pathways. However, the mixed resources and the result-oriented video evaluation repeatedly let students fall into an information trap and only appeal to students' attention to unsuitable resources. Inspired by human-computer interaction, a novel online video assessment LPSA(Linguistic- Presentative-scientific-Artistic) is proposed, integrated by cognitive style and teaching style, to realize more precise learning detection and teaching quality assessment. The LPSA evaluation consists of a four-level classification and eight secondary indexes quantified by machine learning algorithms. By automatically searching for an appropriate threshold of all secondary indexes, a real-time video assessment system is developed to certify its technical feasibility and pedagogical availability. The results show that the proposed AI-assisted assessment could realize practical pedagogical recommendations and real-time supervision.
Jinta Weng, Haoyu Dong 0001, Yue Hu 0002, Hao Wu 0066, Heyan Huang
SMC2
2022 FORTAP: Using Formulas for Numerical-Reasoning-Aware Table Pretraining
abstract
Zhoujun Cheng, Haoyu Dong, Ran Jia, Pengfei Wu, Shi Han, Fan Cheng, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zhoujun Cheng, Haoyu Dong 0001, Ran Jia, Pengfei Wu 0006, Shi Han, Fan Cheng 0002, Dongmei Zhang 0001
ACL (1)2
2022 HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation
abstract
Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zhoujun Cheng, Haoyu Dong 0001, Zhiruo Wang 0001, Ran Jia, Yan Gao 0002, Shi Han, Jian-Guang Lou, Dongmei Zhang 0001
ACL (1)2
2022 PLOG: Table-to-Logic Pretraining for Logical Table-to-Text Generation
abstract
Logical table-to-text generation is a task that involves generating logically faithful sentences from tables, which requires models to derive logical-level facts from table records via logical inference.It raises a new challenge on the logical-level content planning of table-to-text models.However, directly learning the logical inference knowledge from table-text pairs is very difficult for neural models because of the ambiguity of natural language and the scarcity of parallel data.Hence even large-scale pretrained language models present low logical fidelity on logical table-to-text.In this work, we propose a Pretrained Logical Form Generator (PLOG) framework to improve generation fidelity.Specifically, PLOG is first pretrained on a table-to-logical-form generation (table-to-logic) task, then finetuned on downstream table-to-text tasks.The logical forms are formally defined with unambiguous semantics.Hence we can collect a large amount of accurate logical forms from tables without human annotation.In addition, PLOG can learn logical inference from table-logic pairs much more reliably than from table-text pairs.To evaluate our model, we further collect a controlled logical table-to-text dataset CONTLOG based on an existing dataset.On two benchmarks, LOGICNLG and CONTLOG, PLOG outperforms strong baselines by a large margin on logical fidelity, demonstrating the effectiveness of table-to-logic pretraining.
Ao Liu 0008, Haoyu Dong 0001, Naoaki Okazaki, Shi Han, Dongmei Zhang 0001
EMNLP2
2022 TaCube: Pre-computing Data Cubes for Answering Numerical-Reasoning Questions over Tabular Data
abstract
Existing auto-regressive pre-trained language models (PLMs) like T5 and BART, have been well applied to table question answering by UNIFIEDSKG and TAPEX, respectively, and demonstrated state-of-the-art results on multiple benchmarks.However, auto-regressive PLMs are challenged by recent emerging numerical reasoning datasets, such as TAT-QA, due to the error-prone implicit calculation.In this paper, we present TACUBE, to precompute aggregation/arithmetic results for the table in advance, so that they are handy and readily available for PLMs to answer numerical reasoning questions.TACUBE systematically and comprehensively covers a collection of computational operations over table segments.By simply concatenating TACUBE to the input sequence of PLMs, it shows significant experimental effectiveness.TACUBE promotes the F1 score from 49.6% to 66.2% on TAT-QA and achieves new state-of-the-art results on WikiTQ (59.6% denotation accuracy).TACUBE 's improvements on numerical reasoning cases are even more notable: on TAT-QA, TACUBE promotes the exact match accuracy of BART-large by 39.6% on sum, 52.5% on average, 36.6% on subtraction and 22.2% on division.We believe that TACUBE is a general and portable pre-computation solution that can be potentially integrated into various numerical reasoning frameworks.Data and code will be available at https://github.com/ microsoft/TaCube.
Mengkang Hu, Haoyu Dong 0001, Zhoujun Cheng, Fan Cheng 0002, Shi Han, Dongmei Zhang 0001
EMNLP3
2022 Table Pre-training: A Survey on Model Architectures, Pre-training Objectives, and Downstream Tasks
abstract
Following the success of pre-training techniques in the natural language domain, a flurry of table pre-training frameworks have been proposed and have achieved new state-of-the-arts on various downstream tasks such as table question answering, table type recognition, column relation classification, table search, and formula prediction. Various model architectures have been explored to best capture the characteristics of (semi-)structured tables, especially specially-designed attention mechanisms. Moreover, to fully leverage the supervision signals in unlabeled tables, diverse pre-training objectives have been designed and evaluated, for example, denoising cell values, predicting numerical relationships, and learning a neural SQL executor. This survey aims to provide a comprehensive review of model designs, pre-training objectives, and downstream tasks for table pre-training, and we further share our thoughts on existing challenges and future opportunities.
Haoyu Dong 0001, Zhoujun Cheng, Mengyu Zhou, Anda Zhou, Ao Liu 0008, Shi Han, Dongmei Zhang 0001
IJCAI1
2021 Semantic table structure identification in spreadsheets
abstract
Spreadsheets are widely used in various business tasks, and contain amounts of valuable data. However, spreadsheet tables are usually organized in a semi-structured way, and contain complicated semantic structures, e.g., header types and relations among headers. Lack of documented semantic table structures, existing data analysis and error detection tools can hardly understand spreadsheet tables. Therefore, identifying semantic table structures in spreadsheet tables is of great importance, and can greatly promote various analysis tasks on spreadsheets.
Haoyu Dong 0001, Wensheng Dou, Shi Han, Dongmei Zhang 0001, Jun Wei 0001, Dan Ye 0004
ISSTA3
2021 TUTA: Tree-based Transformers for Generally Structured Table Pre-training
abstract
We propose TUTA, a unified pre-training architecture for understanding generally structured tables. Noticing that understanding a table requires spatial, hierarchical, and semantic information, we enhance transformers with three novel structure-aware mechanisms. First, we devise a unified tree-based structure, called a bi-dimensional coordinate tree, to describe both the spatial and hierarchical information of generally structured tables. Upon this, we propose tree-based attention and position embedding to better capture the spatial and hierarchical information. Moreover, we devise three progressive pre-training objectives to enable representations at the token, cell, and table levels. We pre-train TUTA on a wide range of unlabeled web and spreadsheet tables and fine-tune it on two critical tasks in the field of table structure understanding: cell type classification and table type classification. Experiments show that TUTA is highly effective, achieving state-of-the-art on five widely-studied datasets.
Zhiruo Wang 0001, Haoyu Dong 0001, Ran Jia, Jia Li 0012, Zhiyi Fu, Shi Han, Dongmei Zhang 0001
KDD2
2020 Neural Formatting for Spreadsheet Tables
abstract
Spreadsheets are popular and widely used for data presentation and management, where users create tables in various structures to organize and present data. Table formatting is an important yet tedious task for better exhibiting table structures and data relationships. However, without the aid of intelligent tools, manual formatting remains a tedious and time-consuming task. In this paper, we propose CellGAN, a neural formatting model for learning and recommending formats of spreadsheet tables. Based on a novel conditional generative adversarial network (cGAN) architecture, CellGAN learns table formatting from real-world spreadsheet tables in a self-supervised fashion without requiring human labeling. In CellGAN we devise two mechanisms, row/column-wise pooling and local refinement network, to address challenges from the spreadsheet domain. We evaluate the effectiveness of CellGAN against real-world datasets using both quantitative metrics and human perception studies. The results indicate remarkable performance gains over rule-based methods, graphical models or direct application of the state-of-the-art cGANs used in visual synthesis tasks. Neural Formatting is the first step towards auto-formatting for spreadsheet tables with promising results.
Haoyu Dong 0001, Zhouyu Fu, Shi Han, Dongmei Zhang 0001
CIKM1
2020 Learning Formatting Style Transfer and Structure Extraction for Spreadsheet Tables with a Hybrid Neural Network Architecture
abstract
Table formatting is a typical task for spreadsheet users to better exhibit table structures and data relationships. But quickly and effectively formatting tables is a challenge for users. Lots of manual operations are needed, especially for complex tables. In this paper, we propose techniques for table formatting style transfer, i.e., to automatically format a target table according to the style of a reference table. Considering the latent many-to-many mappings between table structures and formats, we propose CellNet, which is a novel end-to-end, multi-task model leveraging conditional Generative Adversarial Networks (cGANs) with three key components to (1) model and recognize table structures; (2) encode formatting styles; (3) learn and apply the latent mapping based on recognized table structure and encoded style, respectively. Moreover, we build up a spreadsheet table corpus containing 5,226 tables with high-quality formats and 784 tables with human-labeled structures. Our evaluation shows that CellNet is highly effective according to both quantitative metrics and human perception studies by comparing with heuristic-based and other learning-based methods.
Haoyu Dong 0001, Jiong Yang 0002, Shi Han, Dongmei Zhang 0001
CIKM1
2019 TableSense: Spreadsheet Table Detection with Convolutional Neural Networks
abstract
Spreadsheet table detection is the task of detecting all tables on a given sheet and locating their respective ranges. Automatic table detection is a key enabling technique and an initial step in spreadsheet data intelligence. However, the detection task is challenged by the diversity of table structures and table layouts on the spreadsheet. Considering the analogy between a cell matrix as spreadsheet and a pixel matrix as image, and encouraged by the successful application of Convolutional Neural Networks (CNN) in computer vision, we have developed TableSense, a novel end-to-end framework for spreadsheet table detection. First, we devise an effective cell featurization scheme to better leverage the rich information in each cell; second, we develop an enhanced convolutional neural network model for table detection to meet the domain-specific requirement on precise table boundary detection; third, we propose an effective uncertainty metric to guide an active learning based smart sampling algorithm, which enables the efficient build-up of a training dataset with 22,176 tables on 10,220 sheets with broad coverage of diverse table structures and layouts. Our evaluation shows that TableSense is highly effective with 91.3% recall and 86.5% precision in EoB-2 metric, a significant improvement over both the current detection algorithm that are used in commodity spreadsheet tools and state-of-the-art convolutional neural networks in computer vision.
Haoyu Dong 0001, Shi Han, Zhouyu Fu, Dongmei Zhang 0001
AAAI1