EDBT 2026 Demo / reviewers in the wild / expert
Yan Gao 0002
dblp:46/3479-2
· DBLP profile ↗
20ranked-venue papers
0as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?abstractConversational Tabular Data Analysis, a collaboration between humans and machines, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic conversational logs for tabular data analysis hinder comprehensive quantitative evaluation of Large Language Models (LLMs) in this task. To mitigate this issue, we introduce CoTA, a new benchmark to evaluate LLMs on conversational tabular data analysis. CoTA contains 1013 conversations, covering 4 practical scenarios: Normal, Action, Private, and Private Action. Notably, CoTA is constructed by an economical multi-agent environment, Decision Company, with few human efforts. This environment ensures efficiency and scalability of generating new conversational data. Our comprehensive study, conducted by data analysis experts, demonstrates that Decision Company is capable of producing diverse and high-quality data, laying the groundwork for efficient data annotation. We evaluate popular and advanced LLMs in CoTA, which highlights the challenges of conversational tabular data analysis. Furthermore, we propose Adaptive Conversation Reflection (ACR), a self-generated reflection strategy that guides LLMs to learn from successful histories. Experiments demonstrate that ACR can evolve LLMs into effective conversational data analysis agents, achieving a relative performance improvement of up to 35.14%. Jinyang Li 0003, Nan Huo, Yan Gao 0002, Yingxiu Zhao, Ge Qu, Bowen Qin, Yurong Wu, Xiaodong Li 0009, Chenhao Ma 0001, Jian-Guang Lou, Reynold Cheng |
ICML | 3 |
| 2024 | Text2Analysis: A Benchmark of Table Question Answering with Advanced Data Analysis and Unclear QueriesabstractTabular data analysis is crucial in various fields, and large language models show promise in this area. However, current research mostly focuses on rudimentary tasks like Text2SQL and TableQA, neglecting advanced analysis like forecasting and chart generation. To address this gap, we developed the Text2Analysis benchmark, incorporating advanced analysis tasks that go beyond the SQL-compatible operations and require more in-depth analysis. We also develop five innovative and effective annotation methods, harnessing the capabilities of large language models to enhance data quality and quantity. Additionally, we include unclear queries that resemble real-world user questions to test how well models can understand and tackle such challenges. Finally, we collect 2249 query-result pairs with 347 tables. We evaluate five state-of-the-art models using three different metrics and the results show that our benchmark presents introduces considerable challenge in the field of tabular data analysis, paving the way for more advanced research opportunities. Mengyu Zhou, Xinrun Xu, Xiaojun Ma 0001, Rui Ding 0001, Lun Du, Yan Gao 0002, Ran Jia, Xu Chen 0022, Shi Han, Zejian Yuan, Dongmei Zhang 0001 |
AAAI | 7 |
| 2024 | AMPO: Automatic Multi-Branched Prompt OptimizationabstractSheng Yang, Yurong Wu, Yan Gao, Zineng Zhou, Bin Benjamin Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Yuan Fang, Yunsong Li, Junyan Chen, Linjun Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yurong Wu, Yan Gao 0002, Zineng Zhou, Bin B. Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Linjun Yang |
EMNLP | 3 |
| 2024 | E⁵: Zero-shot Hierarchical Table Analysis using Augmented LLMs via Explain, Extract, Execute, Exhibit and ExtrapolateabstractZhehao Zhang, Yan Gao, Jian-Guang Lou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhehao Zhang 0001, Yan Gao 0002, Jian-Guang Lou |
NAACL-HLT | 2 |
| 2023 | MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic ParsingabstractText-to-SQL semantic parsing is an important NLP task, which facilitates the interaction between users and the database. Much recent progress in text-to-SQL has been driven by large-scale datasets, but most of them are centered on English. In this work, we present MultiSpider, the largest multilingual text-to-SQL semantic parsing dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese). Upon MultiSpider we further identify the lexical and structural challenges of text-to-SQL (caused by specific language properties and dialect sayings) and their intensity across different languages. Experimental results under various settings (zero-shot, monolingual and multilingual) reveal a 6.1% absolute drop in accuracy in non-English languages. Qualitative and quantitative analyses are conducted to understand the reason for the performance drop of each language. Besides the dataset, we also propose a simple schema augmentation framework SAVe (Schema-Augmentation-with-Verification), which significantly boosts the overall performance by about 1.8% and closes the 29.5% performance gap across languages. Longxu Dou, Yan Gao 0002, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Jian-Guang Lou |
AAAI | 2 |
| 2023 | Uncovering and Categorizing Social Biases in Text-to-SQLabstractContent Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.Large pre-trained language models are acknowledged to carry social biases towards different demographics, which can further amplify existing stereotypes in our society and cause even more harm.Text-to-SQL is an important task, models of which are mainly adopted by authoritative institutions, where unfair decisions may lead to catastrophic consequences.However, existing Text-to-SQL models are trained on clean, neutral datasets, such as Spider and WikiSQL.This, to some extent, cover up social bias in models under ideal conditions, which nevertheless may emerge in real application scenarios.In this work, we aim to uncover and categorize social biases in Text-to-SQL models.We summarize the categories of social biases that may occur in structured data for Text-to-SQL models.We build test benchmarks and reveal that models with similar task accuracy can contain social biases at very different rates.We show how to take advantage of our methodology to uncover and assess social biases in the downstream Text-to-SQL task 1 . Yan Liu 0002, Yan Gao 0002, Xiaokang Chen, Elliott Ash, Jian-Guang Lou |
ACL (1) | 2 |
| 2023 | CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular DataabstractLarge language models (LLMs) show powerful reasoning abilities on various text-based tasks.However, their reasoning capability on structured data such as tables has not been systematically explored.In this work, we first establish a comprehensive taxonomy of reasoning and operation types for tabular data analysis.Then, we construct a complex reasoning QA dataset over tabular data, named CRT-QA (Complex Reasoning QA over Tabular data), with the following unique features: ( 1) it is the first Table QA dataset with multi-step operation and informal reasoning; (2) it contains fine-grained annotations on questions' directness, composition types of sub-questions, and human reasoning paths which can be used to conduct a thorough investigation on LLMs' reasoning ability; (3) it contains a collection of unanswerable and indeterminate questions that commonly arise in real-world situations.We further introduce an efficient and effective tool-augmented method, named ARC (Autoexemplar-guided Reasoning with Code), to use external tools such as Pandas to solve table reasoning tasks without handcrafted demonstrations.The experiment results show that CRT-QA presents a strong challenge for baseline methods and ARC achieves the best result.The dataset and code are available at https://github.com/zzh-SJTU/CRT-QA. Zhehao Zhang 0001, Xitao Li, Yan Gao 0002, Jian-Guang Lou |
EMNLP | 3 |
| 2023 | Uncovering and Quantifying Social Biases in Code GenerationabstractWith the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successfully uncover social biases in code generation models. To quantify the severity of social biases in generated code, we develop a dataset along with three metrics to evaluate the overall social bias and fine-grained unfairness across different demographics. Experimental results on three pre-trained code generation models (Codex, InCoder, and CodeGen) with varying sizes, reveal severe social biases. Moreover, we conduct analysis to provide useful insights for further choice of code generation models with low social bias. Yan Liu 0002, Xiaokang Chen, Yan Gao 0002, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Tsung-Yi Ho |
NeurIPS | 3 |
| 2022 | HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language GenerationabstractZhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Zhoujun Cheng, Haoyu Dong 0001, Zhiruo Wang 0001, Ran Jia, Yan Gao 0002, Shi Han, Jian-Guang Lou, Dongmei Zhang 0001 |
ACL (1) | 6 |
| 2022 | Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table PerturbationabstractThe robustness of Text-to-SQL parsers against adversarial perturbations plays a crucial role in delivering highly reliable applications.Previous studies along this line primarily focused on perturbations in the natural language question side, neglecting the variability of tables.Motivated by this, we propose the Adversarial Table Perturbation (ATP) as a new attacking paradigm to measure the robustness of Textto-SQL models.Following this proposition, we curate ADVETA, the first robustness evaluation benchmark featuring natural and realistic ATPs.All tested state-of-the-art models experience dramatic performance drops on ADVETA, revealing models' vulnerability in real-world practices.To defend against ATP, we build a systematic adversarial training example generation framework tailored for better contextualization of tabular data.Experiments show that our approach not only brings the best robustness improvement against tableside perturbations but also substantially empowers models against NL-side perturbations.We release our benchmark and code at: https://github.com/microsoft/ContextualSP. Xinyu Pi, Yan Gao 0002, Zhoujun Li 0001, Jian-Guang Lou |
ACL (1) | 3 |
| 2022 | Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic KnowledgeabstractLongxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Longxu Dou, Yan Gao 0002, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou |
EMNLP | 2 |
| 2022 | Reasoning Like Program ExecutorsabstractReasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors. Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen |
EMNLP | 7 |
| 2022 | LogiGAN: Learning Logical Reasoning via Adversarial Pre-trainingabstractWe present LogiGAN, an unsupervised adversarial pre-training framework for improving logical reasoning abilities of language models. Upon automatic identification of logical reasoning phenomena in massive text corpus via detection heuristics, we train language models to predict the masked-out logical statements. Inspired by the facilitation effect of reflective thinking in human learning, we analogically simulate the learning-thinking process with an adversarial Generator-Verifier architecture to assist logic learning. LogiGAN implements a novel sequential GAN approach that (a) circumvents the non-differentiable challenge of the sequential GAN by leveraging the Generator as a sentence-level generative likelihood scorer with a learning objective of reaching scoring consensus with the Verifier; (b) is computationally feasible for large-scale pre-training with arbitrary target length. Both base and large size language models pre-trained with LogiGAN demonstrate obvious performance improvement on 12 datasets requiring general reasoning abilities, revealing the fundamental role of logic in broad reasoning, as well as the effectiveness of LogiGAN. Ablation studies on LogiGAN components reveal the relative orthogonality between linguistic and logic abilities and suggest that reflective thinking's facilitation effect might also generalize to machine learning. Xinyu Pi, Wanjun Zhong, Yan Gao 0002, Nan Duan 0001, Jian-Guang Lou |
NeurIPS | 3 |
| 2022 | How to manage a task-oriented virtual assistant software project: an experience reportabstractTask-oriented virtual assistants are software systems that provide users with a natural language interface to complete domain-specific tasks. With the recent technological advances in natural language processing and machine learning, an increasing number of task-oriented virtual assistants have been developed. However, due to the well-known complexity and difficulties of the natural language understanding problem, it is challenging to manage a task-oriented virtual assistant software project. Meanwhile, the management and experience related to the development of virtual assistants are hardly studied or shared in the research community or industry, to the best of our knowledge. To bridge this knowledge gap, in this paper, we share our experience and the lessons that we have learned at managing a task-oriented virtual assistant software project at Microsoft. We believe that our practices and the lessons learned can provide a useful reference for other researchers and practitioners who aim to develop a virtual assistant system. Finally, we have developed a requirement management tool, named SpecSpace, which can facilitate the management of virtual assistant projects. Shuyue Li, Yan Gao 0002, Jian-Guang Lou, Dejian Yang, Ting Liu 0002 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2021 | Translating Headers of Tabular Data: A Pilot Study of Schema TranslationabstractSchema translation is the task of automatically translating headers of tabular data from one language to another.High-quality schema translation plays an important role in crosslingual table searching, understanding and analysis.Despite its importance, schema translation is not well studied in the community, and state-of-the-art neural machine translation models cannot work well on this task because of two intrinsic differences between plain text and tabular data: morphological difference and context difference.To facilitate the research study, we construct the first parallel dataset for schema translation, which consists of 3,158 tables with 11,979 headers written in 6 different languages, including English, Chinese, French, German, Spanish, and Japanese.Also, we propose the first schema translation model called CAST, which is a header-to-header neural machine translation model augmented with schema context.Specifically, we model a target header and its context as a directed graph to represent their entity types and relations.Then CAST encodes the graph with a relational-aware transformer and uses another transformer to decode the header in the target language.Experiments on our dataset demonstrate that CAST significantly outperforms state-of-the-art neural machine translation models.Our dataset will be released at https://github.com/microsoft/ContextualSP. Kunrui Zhu, Yan Gao 0002, Jian-Guang Lou |
EMNLP (1) | 2 |
| 2021 | Keep the Structure: A Latent Shift-Reduce Parser for Semantic ParsingabstractTraditional end-to-end semantic parsing models treat a natural language utterance as a holonomic structure. However, hierarchical structures exist in natural languages, which also align with the hierarchical structures of logical forms. In this paper, we propose a latent shift-reduce parser, called LASP, which decomposes both natural language queries and logical form expressions according to their hierarchical structures and finds local alignment between them to enhance semantic parsing. LASP consists of a base parser and a shift-reduce splitter. The splitter dynamically separates an NL query into several spans. The base parser converts the relevant simple spans into logical forms, which are further combined to obtain the final logical form. We conducted empirical studies on two datasets across different domains and different types of logical forms. The results demonstrate that the proposed method significantly improves the performance of semantic parsing, especially on unseen scenarios. Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001 |
IJCAI | 4 |
| 2020 | "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQLabstractIn Natural Language Interfaces to Databases systems, the text-to-SQL technique allows users to query databases by using natural language questions. Though significant progress in this area has been made recently, most parsers may fall short when they are deployed in real systems. One main reason stems from the difficulty of fully understanding the users' natural language questions. In this paper, we include human in the loop and present a novel parser-independent interactive approach (PIIA) that interacts with users using multi-choice questions and can easily work with arbitrary parsers. Experiments were conducted on two cross-domain datasets, the WikiSQL and the more complex Spider, with five state-of-the-art parsers. These demonstrated that PIIA is capable of enhancing the text-to-SQL performance with limited interaction turns by using both simulation and human evaluation. Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001 |
EMNLP (1) | 4 |
| 2020 | RECPARSER: A Recursive Semantic Parsing Framework for Text-to-SQL TaskabstractNeural semantic parsers usually fail to parse long and complicated utterances into nested SQL queries, due to the large search space. In this paper, we propose a novel recursive semantic parsing framework called RECPARSER to generate the nested SQL query layer-by-layer. It decomposes the complicated nested SQL query generation problem into several progressive non-nested SQL query generation problems. Furthermore, we propose a novel Question Decomposer module to explicitly encourage RECPARSER to focus on different components of an utterance when predicting SQL queries of different layers. Experiments on the Spider dataset show that our approach is more effective compared to the previous works at predicting the nested SQL queries. In addition, we achieve an overall accuracy that is comparable with state-of-the-art approaches. Yan Gao 0002, Bei Chen 0008, Qian Liu 0033, Jian-Guang Lou, Fei Teng 0001, Dongmei Zhang 0001 |
IJCAI | 2 |
| 2020 | Compositional Generalization by Learning Analytical ExpressionsabstractCompositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in such a capability. Inspired by work in cognition which argues compositionality can be captured by variable slots with symbolic functions, we present a refreshing view that connects a memory-augmented neural model with analytical expressions, to achieve compositional generalization. Our model consists of two cooperative neural modules, Composer and Solver, fitting well with the cognitive argument while being able to be trained in an end-to-end manner via a hierarchical reinforcement learning algorithm. Experiments on the well-known benchmark SCAN demonstrate that our model seizes a great ability of compositional generalization, solving all challenges addressed by previous works with 100% accuracies. Qian Liu 0033, Shengnan An, Jian-Guang Lou, Bei Chen 0008, Zeqi Lin, Yan Gao 0002, Nanning Zheng 0001, Dongmei Zhang 0001 |
NeurIPS | 6 |
| 2019 | Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate RepresentationabstractWe present a neural approach called IRNet for complex and cross-domain Text-to-SQL.IR-Net aims to address two challenges: 1) the mismatch between intents expressed in natural language (NL) and the implementation details in SQL; 2) the challenge in predicting columns caused by the large number of outof-domain words.Instead of end-to-end synthesizing a SQL query, IRNet decomposes the synthesis process into three phases.In the first phase, IRNet performs a schema linking over a question and a database schema.Then, IRNet adopts a grammar-based neural model to synthesize a SemQL query which is an intermediate representation that we design to bridge NL and SQL.Finally, IRNet deterministically infers a SQL query from the synthesized SemQL query with domain knowledge.On the challenging Text-to-SQL benchmark Spider, IRNet achieves 46.7% accuracy, obtaining 19.5% absolute improvement over previous state-of-the-art approaches.At the time of writing, IRNet achieves the first position on the Spider leaderboard. * Equal Contributions. Work done during an internship at MSRA.NL: Show the names of students who have a grade higher than 5 and have at least 2 friends. SQL: SELECT T1.name FROM friend AS T1 JOIN highschooler AS T2ON T1.student_id = T2.idWHERE T2 Zecheng Zhan, Yan Gao 0002, Jian-Guang Lou, Ting Liu 0002, Dongmei Zhang 0001 |
ACL (1) | 3 |