EDBT 2026 Demo / reviewers in the wild / expert
Mengyu Zhou
dblp:181/1084
· DBLP profile ↗
22ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-0322-7513ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Computer networks · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Jupiter: Enhancing LLM Data Analysis Capabilities via Notebook and Inference-Time Value-Guided SearchabstractLarge language models (LLMs) have shown great promise in automating data science workflows. However, existing models still struggle with multi-step reasoning and tool use, limiting their effectiveness on complex data analysis tasks. To address this limitation, we propose a scalable pipeline that extracts high-quality, tool-based data analysis tasks and their executable multi-step solutions from real-world Jupyter notebooks and associated data files. Using this pipeline, we introduce NbQA, a large-scale dataset of standardized task–solution pairs that reflect authentic tool-use patterns in practical data science scenarios. To further enhance the multi-step reasoning capabilities, we present Jupiter, a framework that formulates data analysis as a search problem and applies Monte Carlo Tree Search (MCTS) to generate diverse solution trajectories for value model learning. During inference, Jupiter combines the value model and node visit counts to efficiently collect executable multi-step plans with minimal search steps. Experimental results show that Qwen2.5-7B and 14B-Instruct models on NbQA solve 77.82% and 86.38% of tasks on InfiAgent-DABench, respectively—matching or surpassing GPT-4o and advanced agent frameworks. Further evaluations demonstrate improved generalization and stronger tool-use reasoning across diverse multi-step reasoning tasks. Shuocheng Li, Silin Du, Wenxuan Zeng, Mengyu Zhou, Yeye He, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001 |
AAAI | 6 |
| 2026 | SheetBrain: A Neuro-Symbolic Agent for Accurate Reasoning over Complex and Large SpreadsheetsabstractUnderstanding and reasoning over complex spreadsheets remain fundamental challenges for large language models (LLMs), which often struggle with intricate structures and rely solely on neural computation. In this work, we propose SheetBrain, a neuro-symbolic dual-workflow agent framework for precise and interpretable reasoning over tabular data. SheetBrain consists of an understanding module that produces a comprehensive overview of the spreadsheet, including structural summaries and query-specific analyses to guide execution; an execution module that integrates a Python sandbox with preloaded table-processing libraries and an Excel helper toolkit for effective data manipulation; and a validation module that verifies the correctness of reasoning and answers, triggering re-execution if necessary. We evaluate SheetBrain on multiple public QA and manipulation benchmarks, and introduce SheetBench, a new benchmark targeting large, multi-table, and structurally complex spreadsheets. Experimental results show that SheetBrain significantly improves reasoning performance on both existing benchmarks and the more challenging scenarios presented in SheetBench. Jiayuan Su, Mengyu Zhou, Huaxing Zeng, Mengni Jia, Haoyu Dong 0001, Xiaojun Ma 0001, Shi Han, Dongmei Zhang 0001 |
AAAI | 3 |
| 2026 | MARCH: Multi-Agent Reinforced Check for HallucinationabstractZhuo Li, Yupeng Zhang, Pengyu Cheng, Jiajun Song, Mengyu Zhou, Hao Li, Shujie Hu, Yu Qin, Erchao.zec, Xiaoxi Jiang, Guanjunjiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Pengyu Cheng, Jiajun Song, Mengyu Zhou, Shujie Hu, Erchao Zhao, Xiaoxi Jiang, Guanjun Jiang |
ACL (1) | 5 |
| 2026 | Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement LearningabstractHanbing Liu, Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong, Xiaojun Ma, Shi Han, Dongmei Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lang Cao, Yuanyi Ren, Mengyu Zhou, Haoyu Dong 0001, Xiaojun Ma 0001, Shi Han, Dongmei Zhang 0001 |
ACL (1) | 4 |
| 2025 | TableLoRA: Low-rank Adaptation on Table Structure Understanding for Large Language ModelsabstractTabular data are crucial in many fields and their understanding by large language models (LLMs) under high parameter efficiency paradigm is important. However, directly applying parameter-efficient fine-tuning (PEFT) techniques to tabular tasks presents significant challenges, particularly in terms of better table serialization and the representation of two-dimensional structured information within a one-dimensional sequence. To address this, we propose TableLoRA, a module designed to improve LLMs’ understanding of table structure during PEFT. It incorporates special tokens for serializing tables with special token encoder and uses 2D LoRA to encode low-rank information on cell positions. Experiments on four tabular-related datasets demonstrate that TableLoRA consistently outperforms vanilla LoRA and surpasses various table encoding methods tested in control experiments. These findings reveal that TableLoRA, as a table-specific LoRA, enhances the ability of LLMs to process tabular data effectively, especially in low-parameter settings, demonstrating its potential as a robust solution for handling table-related tasks. Mengyu Zhou, Yeye He, Haoyu Dong 0001, Shi Han, Zejian Yuan, Dongmei Zhang 0001 |
ACL (1) | 3 |
| 2025 | Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuningabstractLanguage models such as GPT and Llama have shown remarkable ability on diverse natural language tasks, yet their performance on complex table tasks (e.g., NL-to-Code, data cleaning, etc.) continues to be suboptimal.To improve their performance, task-specific fine-tuning is often needed, which, however, require expensive human labeling and is prone to over-fitting.In this work, we propose TABLE-SPECIALIST, a self-trained fine-tuning paradigm specifically designed for table tasks.Our insight is that for each table task, there often exist two dual versions of the same task, one generative and one classification in nature.Leveraging their duality, we propose a Generator-Validator paradigm to iteratively generate-then-validate training data from language models, to finetune stronger TABLE-SPECIALIST models that can specialize in a given task, without using manually-labeled data. Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong 0001, Shi Han, Dongmei Zhang 0001, Surajit Chaudhuri |
EMNLP | 3 |
| 2025 | MMTU: A Massive Multi-Task Table Understanding and Reasoning BenchmarkabstractTables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally require expert-level users like data engineers, data analysts, and database administrators to operate. Although LLMs have shown remarkable progress in working with tables (e.g., in spreadsheet and database copilot scenarios), comprehensive benchmarking of such capabilities remains limited. In contrast to an extensive and growing list of NLP benchmarks, evaluations of table-related tasks are scarce, and narrowly focus on tasks like NL-to-SQL and Table-QA, overlooking the broader spectrum of real-world tasks that professional users face. This gap limits our understanding and model progress in this important area.In this work, we introduce MMTU, a large-scale benchmark with over 28K questions across 25 real-world table tasks, designed to comprehensively evaluate models ability to understand, reason, and manipulate real tables at the expert-level. These tasks are drawn from decades’ worth of computer science research on tabular data, with a focus on complex table tasks faced by professional users. We show that MMTU require a combination of skills -- including table understanding, reasoning, and coding -- that remain challenging for today's frontier models, where even frontier reasoning models like OpenAI GPT-5 and DeepSeek R1 score only around 69% and 57% respectively, suggesting significant room for improvement. We highlight key findings in our evaluation using MMTU and hope this benchmark drives further advances in understanding and developing foundation models for structured data processing and analysis.Our code and data are available at https://github.com/MMTU-Benchmark/MMTU and https://huggingface.co/datasets/MMTU-benchmark/MMTU. Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong 0001, Shi Han, Lingjiao Chen, Dongmei Zhang 0001, Surajit Chaudhuri, H. V. Jagadish |
NeurIPS | 3 |
| 2024 | Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesabstractTabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities. Mengyu Zhou, Xinrun Xu, Xiaojun Ma 0001, Rui Ding 0001, Lun Du, Yan Gao 0002, Ran Jia, Xu Chen 0022, Shi Han, Zejian Yuan, Dongmei Zhang 0001 |
AAAI | 2 |
| 2024 | Encoding Spreadsheets for Large Language ModelsabstractHaoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haoyu Dong 0001, Yuzhang Tian, Junyu Xiong, Mengyu Zhou, José Cambronero, Yeye He, Shi Han, Dongmei Zhang 0001 |
EMNLP | 5 |
| 2024 | CoCoST: Automatic Complex Code Generation with Online Searching and Correctness TestingabstractLarge Language Models have revolutionized code generation ability by converting natural language descriptions into executable code.However, generating complex code within realworld scenarios remains challenging due to intricate structures, subtle bugs, understanding of advanced data types, and lack of supplementary contents.To address these challenges, we introduce the CoCoST framework, which enhances complex code generation by online searching for more information with planned queries and correctness testing for code refinement.Moreover, CoCoST serializes the complex inputs and outputs to improve comprehension and generates test cases to ensure the adaptability for real-world applications.CoCoST is validated through rigorous experiments on the DS-1000 and ClassEval datasets.Experimental results show that CoCoST substantially improves the quality of complex code generation, highlighting its potential to enhance the practicality of LLMs in generating complex code. Jiaru Zou, Mengyu Zhou, Shi Han, Zejian Yuan, Dongmei Zhang 0001 |
EMNLP | 4 |
| 2024 | Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical StudyabstractLarge language models (LLMs) are becoming attractive as few-shot reasoners to solve Natural Language (NL)-related tasks. However, there is still much to learn about how well LLMs understand structured data, such as tables. Although tables can be used as input to LLMs with serialization, there is a lack of comprehensive studies that examine whether LLMs can truly comprehend such data. In this paper, we try to understand this by designing a benchmark to evaluate the structural understanding capabilities (SUC) of LLMs. The benchmark we create includes seven tasks, each with its own unique challenges, \eg, cell lookup, row retrieval, and size detection. We perform a series of evaluations on GPT-3.5 and GPT-4. We find that performance varied depending on several input choices, including table input format, content order, role prompting, and partition marks. Drawing from the insights gained through the benchmark evaluations, we proposeself-augmentation for effective structural prompting, such as critical value / range identification using internal knowledge of LLMs. When combined with carefully chosen input choices, these structural prompting methods lead to promising improvements in LLM performance on a variety of tabular tasks, \eg, TabFact(\uparrow2.31%), HybridQA(\uparrow2.13%), SQA(\uparrow2.72%), Feverous(\uparrow0.84%), and ToTTo(\uparrow5.68%). We believe that our open-source (please find code and data at https://github.com/microsoft/TableProvider) benchmark and proposed prompting methods can serve as a simple yet generic selection for future research. Yuan Sui 0001, Mengyu Zhou, Mingjie Zhou, Shi Han, Dongmei Zhang 0001 |
WSDM | 2 |
| 2023 | CASR: Generating Complex Sequences with Autoregressive Self-Boost Refinement
Mengyu Zhou, Shi Han, Xiu Li 0001, Dongmei Zhang 0001 |
ICLR | 2 |
| 2022 | FormLM: Recommending Creation Ideas for Online Forms by Modelling Semantic and Structural InformationabstractOnline forms are widely used to collect data from human and have a multi-billion market.Many software products provide online services for creating semi-structured forms where questions and descriptions are organized by predefined structures.However, the design and creation process of forms is still tedious and requires expert knowledge.To assist form designers, in this work we present FormLM to model online forms (by enhancing pre-trained language model with form structural information) and recommend form creation ideas (including question / options recommendations and block type suggestion).For model training and evaluation, we collect the first public online form dataset with 62K online forms.Experiment results show that FormLM significantly outperforms general-purpose language models on all tasks, with an improvement by 4.71 on Question Recommendation and 10.6 on Block Type Suggestion in terms of ROUGE-1 and Macro-F1, respectively. Yijia Shao, Mengyu Zhou, Yifan Zhong, Shi Han, Gideon Huang, Dongmei Zhang 0001 |
EMNLP | 2 |
| 2022 | Towards Robust Numerical Question Answering: Diagnosing Numerical Capabilities of NLP SystemsabstractNumerical Question Answering is the task of answering questions that require numerical capabilities.Previous works introduce general adversarial attacks to Numerical Question Answering, while not systematically exploring numerical capabilities specific to the topic.In this paper, we propose to conduct numerical capability diagnosis on a series of Numerical Question Answering systems and datasets.A series of numerical capabilities are highlighted, and corresponding dataset perturbations are designed.Empirical results indicate that existing systems are severely challenged by these perturbations.E.g., Graph2Tree experienced a 53.83% absolute accuracy drop against the "Extra" perturbation on ASDiv-a, and BART experienced 13.80% accuracy drop against the "Language" perturbation on the numerical subset of DROP.As a counteracting approach, we also investigate the effectiveness of applying perturbations as data augmentation to relieve systems' lack of robust numerical capabilities.With experiment analysis and empirical studies, it is demonstrated that Numerical Question Answering with robust numerical capabilities is still to a large extent an open question.We discuss future directions of Numerical Question Answering and summarize guidelines on future dataset collection and system design. Mengyu Zhou, Shi Han, Dongmei Zhang 0001 |
EMNLP | 2 |
| 2022 | Table Pre-training: A Survey on Model Architectures, Pre-training Objectives, and Downstream TasksabstractFollowing the success of pre-training techniques in the natural language domain, a flurry of table pre-training frameworks have been proposed and have achieved new state-of-the-arts on various downstream tasks such as table question answering, table type recognition, column relation classification, table search, and formula prediction. Various model architectures have been explored to best capture the characteristics of (semi-)structured tables, especially specially-designed attention mechanisms. Moreover, to fully leverage the supervision signals in unlabeled tables, diverse pre-training objectives have been designed and evaluated, for example, denoising cell values, predicting numerical relationships, and learning a neural SQL executor. This survey aims to provide a comprehensive review of model designs, pre-training objectives, and downstream tasks for table pre-training, and we further share our thoughts on existing challenges and future opportunities. Haoyu Dong 0001, Zhoujun Cheng, Mengyu Zhou, Anda Zhou, Ao Liu 0008, Shi Han, Dongmei Zhang 0001 |
IJCAI | 4 |
| 2022 | MultiVision: Designing Analytical Dashboards with Deep Learning Based RecommendationabstractWe contribute a deep-learning-based method that assists in designing analytical dashboards for analyzing a data table. Given a data table, data workers usually need to experience a tedious and time-consuming process to select meaningful combinations of data columns for creating charts. This process is further complicated by the needs of creating dashboards composed of multiple views that unveil different perspectives of data. Existing automated approaches for recommending multiple-view visualizations mainly build on manually crafted design rules, producing sub-optimal or irrelevant suggestions. To address this gap, we present a deep learning approach for selecting data columns and recommending multiple charts. More importantly, we integrate the deep learning models into a mixed-initiative system. Our model could make recommendations given optional user-input selections of data columns. The model, in turn, learns from provenance data of authoring logs in an offline manner. We compare our deep learning model with existing methods for visualization recommendation and conduct a user study to evaluate the usefulness of the system. Aoyu Wu, Yun Wang 0012, Mengyu Zhou, Huamin Qu, Dongmei Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Table2Charts: Recommending Charts by Learning Shared Table RepresentationsabstractIt is common for people to create different types of charts to explore a multi-dimensional dataset (table). However, to recommend commonly composed charts in real world, one should take the challenges of efficiency, imbalanced data and table context into consideration. In this paper, we propose Table2Charts framework which learns common patterns from a large corpus of (table, charts) pairs. Based on deep Q-learning with copying mechanism and heuristic searching, Table2Charts does table-to-sequence generation, where each sequence follows a chart template. On a large spreadsheet corpus with 165k tables and 266k charts, we show that Table2Charts could learn a shared representation of table fields so that recommendation tasks on different chart types could mutually enhance each other. Table2Charts outperforms other chart recommendation systems in both multi-type task (with doubled recall numbers [email protected]=0.61 and [email protected]=0.43) and human evaluations. Mengyu Zhou, Qingtao Li, Yuejiang Li, Shi Han, Daxin Jiang, Dongmei Zhang 0001 |
KDD | 1 |
| 2020 | Table2Analysis: Modeling and Recommendation of Common Analysis Patterns for Multi-Dimensional DataabstractGiven a table of multi-dimensional data, what analyses would human create to extract information from it? From scientific exploration to business intelligence (BI), this is a key problem to solve towards automation of knowledge discovery and decision making. In this paper, we propose Table2Analysis to learn commonly conducted analysis patterns from large amount of (table, analysis) pairs, and recommend analyses for any given table even not seen before. Multi-dimensional data as input challenges existing model architectures and training techniques to fulfill the task. Based on deep Q-learning with heuristic search, Table2Analysis does table to sequence generation, with each sequence encoding an analysis. Table2Analysis has 0.78 recall at top-5 and 0.65 recall at top-1 in our evaluation against a large scale spreadsheet corpus on the PivotTable recommendation task. Mengyu Zhou, Pengxin Ji, Dongmei Zhang 0001 |
AAAI | 1 |
| 2018 | The Frame Latency of Personalized Livestreaming Can Be Significantly Slowed Down by WiFiabstractThe popular personalized livestreaming (PL) in China, arguably the largest PL market in the world, is more monetized than PL in US and hence demands much lower interactive latencies to ensure a good quality of user experience. However, our pilot experiment shows that the video frame latency, dominant component of PL's interactive latency, can be significantly slowed down by WiFi, the primary Internet access method for PL. Understanding and further improving the frame latency over WiFi, however, have difficulties in 1) measuring end-to-end latency; 2) parsing encrypted PL's traffic and 3) modeling complex relationships between WiFi radio factors and the latency. To tackle these challenges, we design and prototype Latency Doctor (LTDr), a practical system which aims to model and optimize PL's video frame latency over WiFi. We deploy LTDr in our campus and obtain several key observations based on 13.9M video frames extracted from 12K individual views on three leading PLs in China. We observe that 40% frame latencies over WiFi hop are more than 30ms, and channel utilization should be less than 64% for low latency. Then we build a predictive model based on the dataset using the machine learning methodologies. Two real cases show that the median frame latencies are decreased by LTDr from 130ms to 22ms, and 50ms to 12ms respectively over WiFi networks. Guoshun Nan, Xiuquan Qiao, Jiting Wang, Zeyan Li 0001, Jiahao Bu, Changhua Pei, Mengyu Zhou, Dan Pei |
IPCCC | 7 |
| 2017 | MinHash hierarchy for privacy preserving trajectory sensing and queryabstractIn this work, we study privacy preserving trajectory sensing and query when n mobile entities (e.g., mobile devices or vehicles) move in an environment of m checkpoints (e.g, WiFi or cellular towers). The checkpoints detect the appearances of mobile entities in the proximity, meanwhile, employ the MinHash signatures to record the set of mobile entities passing by. We build on the checkpoints a distributed data structure named the MinHash hierarchy, with which one can efficiently answer queries regarding popular paths and other traffic patterns. The MinHash hierarchy has a total of near linear storage, linear construction cost, and logarithmic update cost. The cost of a popular path query is logarithmic in the number of checkpoints. Further, the MinHash signature provides privacy protection using a model inspired by the differential privacy model. We evaluated our algorithm using a large mobility data set and compared with previous works to demonstrate its utilities and performances. Jiaxin Ding 0001, Chien-Chun Ni, Mengyu Zhou, Jie Gao 0001 |
IPSN | 3 |
| 2016 | EDUM: classroom education measurements via large-scale WiFi networksabstractBehavior in classroom-based courses is hard to measure at large-scale. In this paper, we propose the EDUM (EDUcation Measurement) system to help characterize educational behavior through data collected from WLANs (WiFi networks) on campuses. EDUM characterizes students' punctuality (attendances, late arrivals, and early departures) for lectures using longitudinal WLAN data, and further characterizes the attractiveness of lectures using mobile phone's interactive states at minute-scale granularity. EDUM is easy to deploy and extensible for new types of data. We deploy EDUM at Tsinghua University where ~700 volunteer students' data are measured during a 9-week period by ~2,800 APs and two popular mobile apps. Our results show that EDUM makes it possible to obtain large-scale observations on punctuality, distraction and study performance, and quantitatively confirm or disprove numerous assumptions about educational behavior. Mengyu Zhou, Minghua Ma, Yangkun Zhang, Kaixin Sui, Dan Pei, Thomas Moscibroda |
UbiComp | 1 |
| 2016 | Characterizing and Improving WiFi Latency in Large-Scale Operational NetworksabstractWiFi latency is a key factor impacting the user experience of modern mobile applications, but it has not been well studied at large scale. In this paper, we design and deploy WiFiSeer, a framework to measure and characterize WiFi latency at large scale. WiFiSeer comprises a systematic methodology for modeling the complex relationships between WiFi latency and a diverse set of WiFi performance metrics, device characteristics, and environmental factors. WiFiSeer was deployed on Tsinghua campus to conduct a WiFi latency measurement study of unprecedented scale with more than 47,000 unique user devices. We observe that WiFi latency follows a long tail distribution and the 90th (99th) percentile is around 20 ms (250 ms). Furthermore, our measurement results quantitatively confirm some anecdotal perceptions about impacting factors and disapprove others. We deploy three practical solutions for improving WiFi latency in Tsinghua, and the results show significantly improved WiFi latencies. In particular, over 1,000 devices use our AP selection service based on a predictive WiFi latency model for 2.5 months, and 72% of their latencies are reduced by over half after they re-associate to the suggested APs. Kaixin Sui, Mengyu Zhou, Minghua Ma, Dan Pei, Youjian Zhao, Zimu Li, Thomas Moscibroda |
MobiSys | 2 |