Jian-Guang Lou

dblp:37/1917 · DBLP profile ↗
← Back
97ranked-venue papers
5as first author
53since 2021 · last 2026
0000-0001-8496-033XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 1 first-author · 45 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-author · 12 since 2021Software engineering, systems software and programming languages · 15 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 1 first-authorSecurity and privacy · 2Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 HyFunc: Accelerating LLM-based Function Calls for Agentic AI through Hybrid-Model Cascade and Dynamic Templating
abstract
While agentic AI systems rely on LLMs to translate user intent into structured function calls, this process is fraught with computational redundancy, leading to high inference latency that hinders real-time applications. This paper identifies and addresses three key redundancies: (1) the redundant processing of a large library of function descriptions for every request; (2) the redundant use of a large, slow model to generate an entire, often predictable, token sequence; and (3) the redundant generation of fixed, boilerplate parameter syntax. We introduce HyFunc, a novel framework that systematically eliminates these inefficiencies. HyFunc employs a hybrid-model cascade where a large model distills user intent into a single ''soft token.'' This token guides a lightweight retriever to select relevant functions and directs a smaller, prefix-tuned model to generate the final call, thus avoiding redundant context processing and full-sequence generation by the large model. To eliminate syntactic redundancy, our ''dynamic templating'' technique injects boilerplate parameter syntax on-the-fly within an extended vLLM engine. To avoid potential limitations in generalization, we evaluate HyFunc on an unseen benchmark dataset, BFCL. Experimental results demonstrate that HyFunc achieves an excellent balance between efficiency and performance. It achieves an inference latency of 0.828 seconds, outperforming all baseline models, and reaches a performance of 80.1%, surpassing all models with a comparable parameter scale. These results suggest that HyFunc offers a more efficient paradigm for agentic AI. Our code is publicly available at https://github.com/MrBlankness/HyFunc.
Weibin Liao, Jian-Guang Lou, Haoyi Xiong
KDD (1)2
2026 Evaluating LLM-based Agents for Multi-turn Conversations: A Survey
abstract
This survey examines evaluation methods for large language model (LLM)-based agents in multi-turn conversational settings. Using a PRISMA-inspired framework, we systematically reviewed nearly 250 scholarly sources, capturing the state-of-the-art from various venues of publication, and establishing a solid foundation for our analysis. Our study offers a structured approach by developing two interrelated taxonomy systems: one that defines what to evaluate and another that explains how to evaluate . The first taxonomy identifies key components of LLM-based agents for multi-turn conversations and their evaluation dimensions, including task completion, response quality, user experience, memory and context retention, as well as planning and tool integration. These components ensure that the performance of conversational agents is assessed in a holistic and meaningful manner. The second taxonomy system focuses on the evaluation methodologies. It categorizes approaches into annotation-based evaluations, automated metrics, hybrid strategies that combine human assessments with quantitative measures, and self-judging methods utilizing LLMs. This framework not only captures traditional metrics derived from language understanding, such as BLEU and ROUGE scores, but also incorporates advanced techniques that reflect the dynamic, interactive nature of multi-turn dialogues. Together, these frameworks summarize the current status quo, expose limitations in traditional practices, and provide a structured blueprint for improvement. Based on the summarization of existing studies, we identify several challenges and propose future directions, including the development of scalable, real-time evaluation pipelines, enhanced privacy-preserving mechanisms, and robust metrics that capture dynamic multi-turn interactions. Our contributions bridge historical insights with modern practices, paving the way for next-generation, reliably evaluated conversational AI systems and offering a comprehensive guide for researchers and practitioners.
Shengyue Guan, Jindong Wang 0001, Jiang Bian 0003, Bin B. Zhu, Jian-Guang Lou, Haoyi Xiong
ACM Trans. Intell. Syst. Technol.5
2025 WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
abstract
Large language models (LLMs), such as GPT-4, have shown remarkable performance in natural language processing (NLP) tasks, including challenging mathematical reasoning. However, most existing open-source models are only pre-trained on large-scale internet data and without math-related optimization. In this paper, we present WizardMath, which enhances the mathematical reasoning abilities of LLMs, by applying our proposed Reinforcement Learning from Evol-Instruct Feedback (RLEIF) method to the domain of math. Through extensive experiments on two mathematical reasoning benchmarks, namely GSM8k and MATH, we reveal the extraordinary capabilities of our model. Remarkably, WizardMath-Mistral 7B surpasses all other open-source LLMs by a substantial margin. Furthermore, WizardMath 70B even outperforms ChatGPT-3.5, Claude Instant, Gemini Pro and Mistral Medium. Additionally, our preliminary exploration highlights the pivotal role of instruction evolution and process supervision in achieving exceptional math performance.
Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Jian-Guang Lou, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, Yansong Tang, Dongmei Zhang 0001
ICLR5
2025 Are Large Language Models Ready for Multi-Turn Tabular Data Analysis?
abstract
Conversational Tabular Data Analysis, a collaboration between humans and machines, enables real-time data exploration for informed decision-making. The challenges and costs of collecting realistic conversational logs for tabular data analysis hinder comprehensive quantitative evaluation of Large Language Models (LLMs) in this task. To mitigate this issue, we introduce CoTA, a new benchmark to evaluate LLMs on conversational tabular data analysis. CoTA contains 1013 conversations, covering 4 practical scenarios: Normal, Action, Private, and Private Action. Notably, CoTA is constructed by an economical multi-agent environment, Decision Company, with few human efforts. This environment ensures efficiency and scalability of generating new conversational data. Our comprehensive study, conducted by data analysis experts, demonstrates that Decision Company is capable of producing diverse and high-quality data, laying the groundwork for efficient data annotation. We evaluate popular and advanced LLMs in CoTA, which highlights the challenges of conversational tabular data analysis. Furthermore, we propose Adaptive Conversation Reflection (ACR), a self-generated reflection strategy that guides LLMs to learn from successful histories. Experiments demonstrate that ACR can evolve LLMs into effective conversational data analysis agents, achieving a relative performance improvement of up to 35.14%.
Jinyang Li 0003, Nan Huo, Yan Gao 0002, Yingxiu Zhao, Ge Qu, Bowen Qin, Yurong Wu, Xiaodong Li 0009, Chenhao Ma 0001, Jian-Guang Lou, Reynold Cheng
ICML11
2025 AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
Mengkang Hu, Pu Zhao 0004, Can Xu 0002, Qingfeng Sun, Jian-Guang Lou, Qingwei Lin, Ping Luo 0002, Saravan Rajmohan
KDD (1)5
2025 Illustration Layout Generation for Slide Enhancement with Pixel-based Diffusion Model
Zhaoyun Jiang, Shakie Liu, Ting Liu 0002, Jian-Guang Lou, Dongmei Zhang 0001
ACM Multimedia6
2025 Explainable automated debugging via large language model-driven scientific debugging
abstract
Abstract Automated debugging techniques have the potential to reduce developer effort in debugging. However, while developers want rationales for the provided automatic debugging results, existing techniques are ill-suited to provide them, as their deduction process differs significantly froof human developers. Inspired by the way developers interact with code when debugging, we propose Automated Scientific Debugging ( AutoSD ), a technique that prompts large language models to automatically generate hypotheses, uses debuggers to interact with buggy code, and thus automatically reach conclusions prior to patch generation. In doing so, we aim to produce explanations of how a specific patch has been generated, with the hope that these explanations will lead to enhanced developer decision-making. Our empirical analysis on three program repair benchmarks shows that AutoSD performs competitively with other program repair baselines, and that it can indicate when it is confident in its results. Furthermore, we perform a human study with 20 participants to evaluate AutoSD -generated explanations. Participants with access to explanations judged patch correctness more accurately in five out of six real-world bugs studied. Furthermore, 70% of participants answered that they wanted explanations when using repair tools, and 55% answered that they were satisfied with the Scientific Debugging presentation.
Sungmin Kang, Bei Chen 0008, Shin Yoo, Jian-Guang Lou
Empir. Softw. Eng.4
2024 Re-Reading Improves Reasoning in Large Language Models
abstract
To enhance the reasoning capabilities of offthe-shelf Large Language Models (LLMs), we introduce a simple, yet general and effective prompting method, RE2, i.e., Re-Reading the question as input.Unlike most thoughteliciting prompting methods, such as Chain-of-Thought (CoT), which aim to elicit the reasoning process in the output, RE2 shifts the focus to the input by processing questions twice, thereby enhancing the understanding process.Consequently, RE2 demonstrates strong generality and compatibility with most thoughteliciting prompting methods, including CoT.Crucially, RE2 facilitates a "bidirectional" encoding in unidirectional decoder-only LLMs because the first pass could provide global information for the second pass.We begin with a preliminary empirical study as the foundation of RE2, illustrating its potential to enable "bidirectional" attention mechanisms.We then evaluate RE2 on extensive reasoning benchmarks across 14 datasets, spanning 112 experiments, to validate its effectiveness and generality.Our findings indicate that, with the exception of a few scenarios on vanilla ChatGPT, RE2 consistently enhances the reasoning performance of LLMs through a simple re-reading strategy.Further analyses reveal RE2's adaptability, showing how it can be effectively integrated with different LLMs, thought-eliciting prompting, and ensemble strategies. 1
Chongyang Tao, Tao Shen 0001, Can Xu 0002, Guodong Long, Jian-Guang Lou, Shuai Ma 0001
EMNLP7
2024 AMPO: Automatic Multi-Branched Prompt Optimization
abstract
Sheng Yang, Yurong Wu, Yan Gao, Zineng Zhou, Bin Benjamin Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Yuan Fang, Yunsong Li, Junyan Chen, Linjun Yang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yurong Wu, Yan Gao 0002, Zineng Zhou, Bin B. Zhu, Xiaodi Sun, Jian-Guang Lou, Zhiming Ding, Anbang Hu, Linjun Yang
EMNLP7
2024 Automatic Instruction Evolving for Large Language Models
abstract
Fine-tuning large pre-trained language models with Evol-Instruct has achieved encouraging results across a wide range of tasks.However, designing effective evolving methods for instruction evolution requires substantial human expertise.This paper proposes Auto Evol-Instruct, an end-to-end framework that evolves instruction datasets using large language models without any human effort.The framework automatically analyzes and summarizes suitable evolutionary strategies for the given instruction data and iteratively improves the evolving method based on issues exposed during the instruction evolution process.Our extensive experiments demonstrate that the best method optimized by Auto Evol-Instruct outperforms human-designed methods on various benchmarks, including MT-Bench, AlpacaEval, GSM8K, and HumanEval.
Can Xu 0002, Yingxiu Zhao, Jian-Guang Lou, Weizhu Chen
EMNLP4
2024 IconDM: Text-Guided Icon Set Expansion Using Diffusion Models
abstract
Icons are ubiquitous visual elements in graphic design, yet their creation is often complex and time-consuming. To resolve this problem, we draw inspiration from the booming text-to-image field and propose Text-Guided Icon Set Expansion, a novel task that helps users design high-quality icons using textual descriptions. Besides, users can control the style consistency of the created icons by inputting a few hand-crafted icons as style reference. Despite its practicality, the task poses two unique challenges. (i) Abstract Concept Visualization. Abstract concepts like technology and health are frequently encountered in icon creation, but their visualization is not straightforward and requires a grounding process that translates them into physical, easy-to-depict objects. (ii) Fine-grained Style Transfer. Unlike ordinary images, icons exhibit richer fine-grained stylistic elements, including tones, line widths, shapes, shadow effects, etc., which puts higher demands on capturing and preserving detailed styles during icon generation.
Zhaoyun Jiang, Shizhao Sun, Ting Liu 0002, Zijiang Yang 0006, Jian-Guang Lou, Dongmei Zhang 0001
ACM Multimedia7
2024 E⁵: Zero-shot Hierarchical Table Analysis using Augmented LLMs via Explain, Extract, Execute, Exhibit and Extrapolate
abstract
Zhehao Zhang, Yan Gao, Jian-Guang Lou. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhehao Zhang 0001, Yan Gao 0002, Jian-Guang Lou
NAACL-HLT3
2024 Make Your LLM Fully Utilize the Context
abstract
While many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the *lost-in-the-middle* challenge. We hypothesize that it stems from insufficient explicit supervision during the long-context training, which fails to emphasize that any position in a long context can hold crucial information. Based on this intuition, our study presents **information-intensive (IN2) training**, a purely data-driven solution to overcome lost-in-the-middle. Specifically, IN2 training leverages a synthesized long-context question-answer dataset, where the answer requires (1) **fine-grained information awareness** on a short segment (~128 tokens) within a synthesized long context (4K-32K tokens), and (2) the **integration and reasoning** of information from two or more short segments. Through applying this information-intensive training on Mistral-7B, we present **FILM-7B** (FIll-in-the-Middle). To thoroughly assess the ability of FILM-7B for utilizing long contexts, we design three probing tasks that encompass various context styles (document, code, and structured-data context) and information retrieval patterns (forward, backward, and bi-directional retrieval). The probing results demonstrate that FILM-7B can robustly retrieve information from different positions in its 32K context window. Beyond these probing tasks, FILM-7B significantly improves the performance on real-world long-context tasks (e.g., 23.5->26.9 F1 score on NarrativeQA), while maintaining a comparable performance on short-context tasks (e.g., 59.3->59.2 accuracy on MMLU).
Shengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng 0001, Jian-Guang Lou, Weizhu Chen
NeurIPS5
2024 WizardArena: Post-training Large Language Models via Simulated Offline Chatbot Arena
abstract
Recent work demonstrates that, post-training large language models with open-domain instruction following data have achieved colossal success. Simultaneously, human Chatbot Arena has emerged as one of the most reasonable benchmarks for model evaluation and developmental guidance. However, the processes of manually curating high-quality training data and utilizing online human evaluation platforms are both expensive and limited. To mitigate the manual and temporal costs associated with post-training, this paper introduces a Simulated Chatbot Arena named WizardArena, which is fully based on and powered by open-source LLMs. For evaluation scenario, WizardArena can efficiently predict accurate performance rankings among different models based on offline test set. For training scenario, we simulate arena battles among various state-of-the-art models on a large scale of instruction data, subsequently leveraging the battle results to constantly enhance target model in both the supervised fine-tuning and reinforcement learning . Experimental results demonstrate that our WizardArena aligns closely with the online human arena rankings, and our models trained on offline extensive battle data exhibit significant performance improvements during SFT, DPO, and PPO stages.
Qingfeng Sun, Can Xu 0002, Pu Zhao 0004, Qingwei Lin, Jian-Guang Lou, Shifeng Chen, Yansong Tang, Weizhu Chen
NeurIPS6
2023 MultiSpider: Towards Benchmarking Multilingual Text-to-SQL Semantic Parsing
abstract
Text-to-SQL semantic parsing is an important NLP task, which facilitates the interaction between users and the database. Much recent progress in text-to-SQL has been driven by large-scale datasets, but most of them are centered on English. In this work, we present MultiSpider, the largest multilingual text-to-SQL semantic parsing dataset which covers seven languages (English, German, French, Spanish, Japanese, Chinese, and Vietnamese). Upon MultiSpider we further identify the lexical and structural challenges of text-to-SQL (caused by specific language properties and dialect sayings) and their intensity across different languages. Experimental results under various settings (zero-shot, monolingual and multilingual) reveal a 6.1% absolute drop in accuracy in non-English languages. Qualitative and quantitative analyses are conducted to understand the reason for the performance drop of each language. Besides the dataset, we also propose a simple schema augmentation framework SAVe (Schema-Augmentation-with-Verification), which significantly boosts the overall performance by about 1.8% and closes the 29.5% performance gap across languages.
Longxu Dou, Yan Gao 0002, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Jian-Guang Lou
AAAI7
2023 How Do In-Context Examples Affect Compositional Generalization?
abstract
Shengnan An, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Shengnan An, Zeqi Lin, Qiang Fu 0015, Bei Chen 0008, Nanning Zheng 0001, Jian-Guang Lou, Dongmei Zhang 0001
ACL (1)6
2023 Making Language Models Better Reasoners with Step-Aware Verifier
abstract
Yifei Li, Zeqi Lin, Shizhuo Zhang, Qiang Fu, Bei Chen, Jian-Guang Lou, Weizhu Chen. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yifei Li 0005, Zeqi Lin, Shizhuo Zhang, Qiang Fu 0015, Bei Chen 0008, Jian-Guang Lou, Weizhu Chen
ACL (1)6
2023 Uncovering and Categorizing Social Biases in Text-to-SQL
abstract
Content Warning: This work contains examples that potentially implicate stereotypes, associations, and other harms that could be offensive to individuals in certain social groups.Large pre-trained language models are acknowledged to carry social biases towards different demographics, which can further amplify existing stereotypes in our society and cause even more harm.Text-to-SQL is an important task, models of which are mainly adopted by authoritative institutions, where unfair decisions may lead to catastrophic consequences.However, existing Text-to-SQL models are trained on clean, neutral datasets, such as Spider and WikiSQL.This, to some extent, cover up social bias in models under ideal conditions, which nevertheless may emerge in real application scenarios.In this work, we aim to uncover and categorize social biases in Text-to-SQL models.We summarize the categories of social biases that may occur in structured data for Text-to-SQL models.We build test benchmarks and reveal that models with similar task accuracy can contain social biases at very different rates.We show how to take advantage of our methodology to uncover and assess social biases in the downstream Text-to-SQL task 1 .
Yan Liu 0002, Yan Gao 0002, Xiaokang Chen, Elliott Ash, Jian-Guang Lou
ACL (1)6
2023 Large Language Models Meet NL2Code: A Survey
abstract
Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Wang Yongji, Jian-Guang Lou. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Daoguang Zan, Bei Chen 0008, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang 0002, Jian-Guang Lou
ACL (1)8
2023 Hadamard Adapter: An Extreme Parameter-Efficient Adapter Tuning Method for Pre-trained Language Models
abstract
Recent years, Pre-trained Language models (PLMs) have swept into various fields of artificial intelligence and achieved great success. However, most PLMs, such as T5 and GPT3, have a huge amount of parameters, fine-tuning them is often expensive and time consuming, and storing them takes up a lot of space. Therefore, it is necessary to adopt a parameter-efficient approach to reduce parameters of PLMs in fine-tuning without compromising their performance in downstream tasks. In this paper, we design a novel adapter which only acts on self-attention outputs in PLMs. This adapter adopts element-wise linear transformation using Hadamard product, hence named as Hadamard adapter, requires the fewest parameters compared to previous parameter-efficient adapters. In addition, we also summarize some tuning patterns for Hadamard adapter shared by various downstream tasks, expecting to provide some guidance for further parameter reduction with shared adapters in future studies. The experiments conducted on the widely-used GLUE benchmark with several SOTA PLMs prove that the Hadamard adapter achieves competitive performance with only 0.033% parameters compared with full fine-tuning, and it has the fewest parameters compared with other adapters. Moreover, we further find that there is also some redundant layers in the Hadamard adapter which can be removed to achieve more parameter efficiency with only 0.022% parameters.
Yuyan Chen, Qiang Fu 0015, Ge Fan, Lun Du, Jian-Guang Lou, Shi Han, Dongmei Zhang 0001, Zhixu Li, Yanghua Xiao
CIKM5
2023 LayoutFormer++: Conditional Graphic Layout Generation via Constraint Serialization and Decoding Space Restriction
abstract
Conditional graphic layout generation, which generates realistic layouts according to user constraints, is a challenging task that has not been well-studied yet. First, there is limited discussion about how to handle diverse user constraints flexibly and uniformly. Second, to make the layouts conform to user constraints, existing work often sacrifices generation quality significantly. In this work, we propose LayoutFormer++ to tackle the above problems. First, to flexibly handle diverse constraints, we propose a constraint serialization scheme, which represents different user constraints as sequences of tokens with a predefined format. Then, we formulate conditional layout generation as a sequence-to-sequence transformation, and leverage encoder-decoder framework with Transformer as the basic architecture. Furthermore, to make the layout better meet user requirements without harming quality, we propose a decoding space restriction strategy. Specifically, we prune the predicted distribution by ignoring the options that definitely violate user constraints and likely result in low-quality layouts, and make the model samples from the restricted distribution. Experiments demonstrate that LayoutFormer++ outperforms existing approaches on all the tasks in terms of both better generation quality and less constraint violation.
Zhaoyun Jiang, Shizhao Sun, Huayu Deng, Zhongkai Wu, Vuksan Mijovic, Zijiang Yang 0006, Jian-Guang Lou, Dongmei Zhang 0001
CVPR8
2023 Skill-Based Few-Shot Selection for In-Context Learning
abstract
In-context learning is the paradigm that adapts large language models to downstream tasks by providing a few examples.Few-shot selectionselecting appropriate examples for each test instance separately-is important for in-context learning.In this paper, we propose SKILL-KNN, a skill-based few-shot selection method for in-context learning.The key advantages of SKILL-KNN include: (1) it addresses the problem that existing methods based on pre-trained embeddings can be easily biased by surface natural language features that are not important for the target task; (2) it does not require training or fine-tuning of any models, making it suitable for frequently expanding or changing example banks.The key insight is to optimize the inputs fed into the embedding model, rather than tuning the model itself.Technically, SKILL-KNN generates the skill-based descriptions for each test case and candidate example by utilizing a pre-processing few-shot prompting, thus eliminating unimportant surface features.Experimental results across five cross-domain semantic parsing datasets and six backbone models show that SKILL-KNN significantly outperforms existing methods.* Work done during the internship at Microsoft.DB Schema: employee (name, age, city, …) … Question: Which cities do more than one employee under age 30 come from?Prompting-Based Rewriting Off-the-Shelf Embedding Model Skill-Based Descriptions Input Query Candidate 1:This task requires the greater-than and less-than constraints.Candidate 2:This task requires to apply two constraints on one selected column.Candidate … … Example Bank Skill-Based Selection DB Schema: cinema (name, capacity, location, …) … Question: Find the locations that have more than one movie theater with capacity above 300.SQL Query: SELECT … WHERE capacity > 300 GROUP BY location HAVING count(*) > 1 DB Schema: endowment (school, donator, amount, …) … Question : Find the number of schools that have more than one donator whose donation amount is less than 8
Shengnan An, Zeqi Lin, Qiang Fu 0015, Bei Chen 0008, Nanning Zheng 0001, Weizhu Chen, Jian-Guang Lou
EMNLP8
2023 RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation
abstract
Fengji Zhang, Bei Chen, Yue Zhang, Jacky Keung, Jin Liu, Daoguang Zan, Yi Mao, Jian-Guang Lou, Weizhu Chen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Fengji Zhang, Bei Chen 0008, Jacky W. Keung, Jin Liu 0016, Daoguang Zan, Jian-Guang Lou, Weizhu Chen
EMNLP8
2023 CRT-QA: A Dataset of Complex Reasoning Question Answering over Tabular Data
abstract
Large language models (LLMs) show powerful reasoning abilities on various text-based tasks.However, their reasoning capability on structured data such as tables has not been systematically explored.In this work, we first establish a comprehensive taxonomy of reasoning and operation types for tabular data analysis.Then, we construct a complex reasoning QA dataset over tabular data, named CRT-QA (Complex Reasoning QA over Tabular data), with the following unique features: ( 1) it is the first Table QA dataset with multi-step operation and informal reasoning; (2) it contains fine-grained annotations on questions' directness, composition types of sub-questions, and human reasoning paths which can be used to conduct a thorough investigation on LLMs' reasoning ability; (3) it contains a collection of unanswerable and indeterminate questions that commonly arise in real-world situations.We further introduce an efficient and effective tool-augmented method, named ARC (Autoexemplar-guided Reasoning with Code), to use external tools such as Pandas to solve table reasoning tasks without handcrafted demonstrations.The experiment results show that CRT-QA presents a strong challenge for baseline methods and ARC achieves the best result.The dataset and code are available at https://github.com/zzh-SJTU/CRT-QA.
Zhehao Zhang 0001, Xitao Li, Yan Gao 0002, Jian-Guang Lou
EMNLP4
2023 Question Answering as Programming for Solving Time-Sensitive Questions
abstract
Question answering plays a pivotal role in human daily life because it involves our acquisition of knowledge about the world.However, due to the dynamic and ever-changing nature of real-world facts, the answer can be completely different when the time constraint in the question changes.Recently, Large Language Models (LLMs) have shown remarkable intelligence in question answering, while our experiments reveal that the aforementioned problems still pose a significant challenge to existing LLMs.This can be attributed to the LLMs' inability to perform rigorous reasoning based on surfacelevel text semantics.To overcome this limitation, rather than requiring LLMs to directly answer the question, we propose a novel approach where we reframe the Question Answering task as Programming (QAaP).Concretely, by leveraging modern LLMs' superior capability in understanding both natural language and programming language, we endeavor to harness LLMs to represent diversely expressed text as wellstructured code and select the best matching answer from multiple candidates through programming.We evaluate our QAaP framework on several time-sensitive question answering datasets and achieve decent improvement, up to 14.5% over strong baselines.1
Cheng Yang 0007, Bei Chen 0008, Siheng Li, Jian-Guang Lou, Yujiu Yang 0001
EMNLP5
2023 LayoutDiffusion: Improving Graphic Layout Generation by Discrete Diffusion Probabilistic Models
abstract
Creating graphic layouts is a fundamental step in graphic designs. In this work, we present a novel generative model named LayoutDiffusion for automatic layout generation. As layout is typically represented as a sequence of discrete tokens, LayoutDiffusion models layout generation as a discrete denoising diffusion process. It learns to reverse a mild forward process, in which layouts become increasingly chaotic with the growth of forward steps and layouts in the neighboring steps do not differ too much. Designing such a mild forward process is however very challenging as layout has both categorical attributes and ordinal attributes. To tackle the challenge, we summarize three critical factors for achieving a mild forward process for the layout, i.e., legality, coordinate proximity and type disruption. Based on the factors, we propose a block-wise transition matrix coupled with a piece-wise linear noise schedule. Experiments on RICO and PubLayNet datasets show that LayoutDiffusion outperforms state-of-the-art approaches significantly. Moreover, it enables two conditional layout generation tasks in a plug-and-play manner without re-training and achieves better performance than existing methods. Project page: https://layoutdiffusion.github.io.
Junyi Zhang 0004, Shizhao Sun, Jian-Guang Lou, Dongmei Zhang 0001
ICCV4
2023 A Parse-Then-Place Approach for Generating Graphic Layouts from Textual Descriptions
abstract
Creating layouts is a fundamental step in graphic design. In this work, we propose to use text as the guidance to create graphic layouts, i.e., Text-to-Layout, aiming to lower the design barriers. Text-to-Layout is a challenging task, because it needs to consider the implicit, combined, and incomplete layout constraints from text, each of which has not been studied in previous work. To address this, we present a two-stage approach, named parse-then-place. The approach introduces an intermediate representation (IR) between text and layout to represent diverse layout constraints. With IR, Text-to-Layout is decomposed into a parse stage and a place stage. The parse stage takes a textual description as input and generates an IR, in which the implicit constraints from the text are transformed into explicit ones. The place stage generates layouts based on the IR. To model combined and incomplete constraints, we use a Transformer-based layout generation model and carefully design a way to represent constraints and layouts as sequences. Besides, we adopt the pretrain-then-finetune strategy to boost the performance of the layout generation model with large-scale unlabeled layouts. To evaluate our approach, we construct two Text-to-Layout datasets and conduct experiments on them. Quantitative results, qualitative analysis, and user studies demonstrate our approach’s effectiveness.
Shizhao Sun, Weijiang Xu, Ting Liu 0002, Jian-Guang Lou, Dongmei Zhang 0001
ICCV6
2023 Does Deep Learning Learn to Abstract? A Systematic Probing Framework
Shengnan An, Zeqi Lin, Bei Chen 0008, Qiang Fu 0015, Nanning Zheng 0001, Jian-Guang Lou
ICLR6
2023 CodeT: Code Generation with Generated Tests
Bei Chen 0008, Fengji Zhang, Daoguang Zan, Zeqi Lin, Jian-Guang Lou, Weizhu Chen
ICLR6
2023 LayoutPrompter: Awaken the Design Ability of Large Language Models
abstract
Conditional graphic layout generation, which automatically maps user constraints to high-quality layouts, has attracted widespread attention today. Although recent works have achieved promising performance, the lack of versatility and data efficiency hinders their practical applications. In this work, we propose LayoutPrompter, which leverages large language models (LLMs) to address the above problems through in-context learning. LayoutPrompter is made up of three key components, namely input-output serialization, dynamic exemplar selection and layout ranking. Specifically, the input-output serialization component meticulously designs the input and output formats for each layout generation task. Dynamic exemplar selection is responsible for selecting the most helpful prompting exemplars for a given input. And a layout ranker is used to pick the highest quality layout from multiple outputs of LLMs. We conduct experiments on all existing layout generation tasks using four public datasets. Despite the simplicity of our approach, experimental results show that LayoutPrompter can compete with or even outperform state-of-the-art approaches on these tasks without any model training or fine-tuning. This demonstrates the effectiveness of this versatile and training-free approach. In addition, the ablation studies show that LayoutPrompter is significantly superior to the training-based baseline in a low-data regime, further indicating the data efficiency of LayoutPrompter. Our project is available at https://github.com/microsoft/LayoutGeneration/tree/main/LayoutPrompter.
Shizhao Sun, Zijiang Yang 0006, Jian-Guang Lou, Dongmei Zhang 0001
NeurIPS5
2023 Uncovering and Quantifying Social Biases in Code Generation
abstract
With the popularity of automatic code generation tools, such as Copilot, the study of the potential hazards of these tools is gaining importance. In this work, we explore the social bias problem in pre-trained code generation models. We propose a new paradigm to construct code prompts and successfully uncover social biases in code generation models. To quantify the severity of social biases in generated code, we develop a dataset along with three metrics to evaluate the overall social bias and fine-grained unfairness across different demographics. Experimental results on three pre-trained code generation models (Codex, InCoder, and CodeGen) with varying sizes, reveal severe social biases. Moreover, we conduct analysis to provide useful insights for further choice of code generation models with low social bias.
Yan Liu 0002, Xiaokang Chen, Yan Gao 0002, Fengji Zhang, Daoguang Zan, Jian-Guang Lou, Tsung-Yi Ho
NeurIPS7
2022 Coarse-to-Fine Generative Modeling for Graphic Layouts
abstract
Even though graphic layout generation has attracted growing attention recently, it is still challenging to synthesis realistic and diverse layouts, due to the complicated element relationships and varied element arrangements. In this work, we seek to improve the performance of layout generation by incorporating the concept of regions, which consist of a smaller number of elements and appears like a simple layout, into the generation process. Specifically, we leverage Variational Autoencoder (VAE) as the overall architecture and decompose the decoding process into two stages. The first stage predicts representations for regions, and the second stage fills in the detailed position for each element within the region based on the predicted region representation. Compared to prior studies that merely abstract the layout into a list of elements and generate all the element positions in one go, our approach has at least two advantages. First, by the two-stage decoding, our approach decouples the complex layout generation task into several simple layout generation tasks, which reduces the problem difficulty. Second, the predicted regions can help the model roughly know what the graphic layout looks like and serve as global context to improve the generation of detailed element positions. Qualitative and quantitative experiments demonstrate that our approach significantly outperforms the existing methods, especially on the complex graphic layouts.
Zhaoyun Jiang, Shizhao Sun, Jihua Zhu, Jian-Guang Lou, Dongmei Zhang 0001
AAAI4
2022 HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation
abstract
Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Zhoujun Cheng, Haoyu Dong 0001, Zhiruo Wang 0001, Ran Jia, Yan Gao 0002, Shi Han, Jian-Guang Lou, Dongmei Zhang 0001
ACL (1)8
2022 Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table Perturbation
abstract
The robustness of Text-to-SQL parsers against adversarial perturbations plays a crucial role in delivering highly reliable applications.Previous studies along this line primarily focused on perturbations in the natural language question side, neglecting the variability of tables.Motivated by this, we propose the Adversarial Table Perturbation (ATP) as a new attacking paradigm to measure the robustness of Textto-SQL models.Following this proposition, we curate ADVETA, the first robustness evaluation benchmark featuring natural and realistic ATPs.All tested state-of-the-art models experience dramatic performance drops on ADVETA, revealing models' vulnerability in real-world practices.To defend against ATP, we build a systematic adversarial training example generation framework tailored for better contextualization of tabular data.Experiments show that our approach not only brings the best robustness improvement against tableside perturbations but also substantially empowers models against NL-side perturbations.We release our benchmark and code at: https://github.com/microsoft/ContextualSP.
Xinyu Pi, Yan Gao 0002, Zhoujun Li 0001, Jian-Guang Lou
ACL (1)6
2022 GL-CLeF: A Global-Local Contrastive Learning Framework for Cross-lingual Spoken Language Understanding
abstract
Libo Qin, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Libo Qin 0001, Qiguang Chen, Tianbao Xie, Qixin Li, Jian-Guang Lou, Wanxiang Che, Min-Yen Kan
ACL (1)5
2022 AdapterShare: Task Correlation Modeling with Adapter Differentiation
abstract
Thanks to the development of pre-trained language models, multitask learning (MTL) methods have achieved great success in natural language understanding.However, current MTL methods pay more attention to task selection or model design to fuse as much knowledge as possible, while the intrinsic task correlation is often neglected.It is important to learn sharing strategies among multiple tasks rather than sharing everything.In this paper, we propose AdapterShare, an adapter differentiation method to explicitly model task correlation among multiple tasks.AdapterShare is automatically learned based on the gradients on tiny held-out validation data.Compared to single-task learning and fully shared MTL methods, our proposed method obtains obvious performance improvements.Compared to the existing MTL method AdapterFusion, AdapterShare achieves an absolute average improvement of 1.90 points on five dialogue understanding tasks and 2.33 points on NLU tasks.Our implementation is available at https:// github.com/microsoft/ContextualSP.
Zhi Chen 0006, Bei Chen 0008, Lu Chen 0002, Kai Yu 0004, Jian-Guang Lou
EMNLP5
2022 Towards Knowledge-Intensive Text-to-SQL Semantic Parsing with Formulaic Knowledge
abstract
Longxu Dou, Yan Gao, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Longxu Dou, Yan Gao 0002, Xuqi Liu, Mingyang Pan, Dingzirui Wang, Wanxiang Che, Dechen Zhan, Min-Yen Kan, Jian-Guang Lou
EMNLP9
2022 Exploring the Secrets Behind the Learning Difficulty of Meaning Representations for Semantic Parsing
abstract
Previous research has shown that the design of Meaning Representation (MR) greatly influences model performance of a neural semantic parser.Therefore, designing a good MR is a long-term goal for semantic parsing.However, it is still an art as there is no quantitative indicator that can tell us which MR among a set of candidates may have the best final model performance.In practice, in order to select an MR, researchers often have to go through the whole training-testing process for all MR candidates, and the process often costs a lot.In this paper, we propose a data-aware metric called ISS (denoting incremental structural stability) of MRs, and demonstrate that ISS is highly correlated with model performance.The finding shows that ISS can be used as an indicator for designing MRs to avoid the costly training-testing process.
Zhenwen Li, Qian Liu 0033, Jian-Guang Lou, Tao Xie 0001
EMNLP4
2022 Reasoning Like Program Executors
abstract
Reasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.
Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen
EMNLP8
2022 TAPEX: Table Pre-training via Learning a Neural SQL Executor
Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou
ICLR7
2022 Nufix: Escape From NuGet Dependency Maze
abstract
Developers usually suffer from dependency maze (DM) issues, i.e., package dependency constraints are violated when a project's platform or dependencies are changed. This problem is especially serious in .NET ecosystem due to its fragmented platforms (e.g., .NET Framework, .NET Core, and .NET Standard). Fixing DM issues is challenging due to the complexity of dependency constraints: multiple DM issues often occur in one project; solving one DM issue usually causes another DM issue cropping up; the exponential search space of possible dependency combinations is also a barrier.
Zhenming Li, Ying Wang 0038, Zeqi Lin, Shing-Chi Cheung, Jian-Guang Lou
ICSE5
2022 CERT: Continual Pre-training on Sketches for Library-oriented Code Generation
abstract
Code generation is a longstanding challenge, aiming to generate a code snippet based on a natural language description. Usually, expensive text-code paired data is essential for training a code generation model. Recently, thanks to the success of pre-training techniques, large language models are trained on large unlabelled code corpora and perform well in generating code. In this paper, we investigate how to leverage an unlabelled code corpus to train a model for library-oriented code generation. Since it is a common practice for programmers to reuse third-party libraries, in which case the text-code paired data are harder to obtain due to the huge number of libraries. We observe that library-oriented code snippets are more likely to share similar code sketches. Hence, we present CERT with two steps: a sketcher generates the sketch, then a generator fills the details in the sketch. Both the sketcher and generator are continually pre-trained upon a base model using unlabelled data. Also, we carefully craft two benchmarks to evaluate library-oriented code generation named PandasEval and NumpyEval. Experimental results have shown the impressive performance of CERT. For example, it surpasses the base model by an absolute 15.67% improvement in terms of pass@1 on PandasEval. Our work is available at https://github.com/microsoft/PyCodeGPT.
Daoguang Zan, Bei Chen 0008, Dejian Yang, Zeqi Lin, Bei Guan, Yongji Wang 0002, Weizhu Chen, Jian-Guang Lou
IJCAI9
2022 LogiGAN: Learning Logical Reasoning via Adversarial Pre-training
abstract
We present LogiGAN, an unsupervised adversarial pre-training framework for improving logical reasoning abilities of language models. Upon automatic identification of logical reasoning phenomena in massive text corpus via detection heuristics, we train language models to predict the masked-out logical statements. Inspired by the facilitation effect of reflective thinking in human learning, we analogically simulate the learning-thinking process with an adversarial Generator-Verifier architecture to assist logic learning. LogiGAN implements a novel sequential GAN approach that (a) circumvents the non-differentiable challenge of the sequential GAN by leveraging the Generator as a sentence-level generative likelihood scorer with a learning objective of reaching scoring consensus with the Verifier; (b) is computationally feasible for large-scale pre-training with arbitrary target length. Both base and large size language models pre-trained with LogiGAN demonstrate obvious performance improvement on 12 datasets requiring general reasoning abilities, revealing the fundamental role of logic in broad reasoning, as well as the effectiveness of LogiGAN. Ablation studies on LogiGAN components reveal the relative orthogonality between linguistic and logic abilities and suggest that reflective thinking's facilitation effect might also generalize to machine learning.
Xinyu Pi, Wanjun Zhong, Yan Gao 0002, Nan Duan 0001, Jian-Guang Lou
NeurIPS5
2022 UniDU: Towards A Unified Generative Dialogue Understanding Framework
abstract
With the development of pre-trained language models, remarkable success has been witnessed in dialogue understanding (DU).However, current DU approaches usually employ independent models for each distinct DU task without considering shared knowledge across different DU tasks.In this paper, we propose a unified generative dialogue understanding framework, named UniDU, to achieve effective information exchange across diverse DU tasks.Here, we reformulate all DU tasks into a unified promptbased generative model paradigm.More importantly, a novel model-agnostic multi-task training strategy (MATS) is introduced to dynamically adapt the weights of diverse tasks for best knowledge sharing during training, based on the nature and available data of each task.Experiments on ten DU datasets covering five fundamental DU tasks show that the proposed UniDU framework largely outperforms task-specific well-designed methods on all tasks.MATS also reveals the knowledgesharing structure of these tasks.Finally, UniDU obtains promising performance in the unseen dialogue domain, showing the great potential for generalization.
Zhi Chen 0006, Lu Chen 0002, Bei Chen 0008, Libo Qin 0001, Yuncong Liu, Su Zhu, Jian-Guang Lou, Kai Yu 0004
SIGDIAL7
2022 How to manage a task-oriented virtual assistant software project: an experience report
abstract
Task-oriented virtual assistants are software systems that provide users with a natural language interface to complete domain-specific tasks. With the recent technological advances in natural language processing and machine learning, an increasing number of task-oriented virtual assistants have been developed. However, due to the well-known complexity and difficulties of the natural language understanding problem, it is challenging to manage a task-oriented virtual assistant software project. Meanwhile, the management and experience related to the development of virtual assistants are hardly studied or shared in the research community or industry, to the best of our knowledge. To bridge this knowledge gap, in this paper, we share our experience and the lessons that we have learned at managing a task-oriented virtual assistant software project at Microsoft. We believe that our practices and the lessons learned can provide a useful reference for other researchers and practitioners who aim to develop a virtual assistant system. Finally, we have developed a requirement management tool, named SpecSpace, which can facilitate the management of virtual assistant projects.
Shuyue Li, Yan Gao 0002, Jian-Guang Lou, Dejian Yang, Ting Liu 0002
Frontiers Inf. Technol. Electron. Eng.4
2021 Iterative Utterance Segmentation for Neural Semantic Parsing
abstract
Neural semantic parsers usually fail to parse long and complex utterances into correct meaning representations, due to the lack of exploiting the principle of compositionality. To address this issue, we present a novel framework for boosting neural semantic parsers via iterative utterance segmentation. Given an input utterance, our framework iterates between two neural modules: a segmenter for segmenting a span from the utterance, and a parser for mapping the span into a partial meaning representation. Then, these intermediate parsing results are composed into the final meaning representation. One key advantage is that this framework does not require any handcraft templates or additional labeled data for utterance segmentation: we achieve this through proposing a novel training method, in which the parser provides pseudo supervision for the segmenter. Experiments on Geo, ComplexWebQuestions and Formulas show that our framework can consistently improve performances of neural semantic parsers in different domains. On data splits that require compositional generalization, our framework brings significant accuracy gains: Geo 63.1~81.2, Formulas 59.7~72.7, ComplexWebQuestions 27.1~56.3.
Yinuo Guo, Zeqi Lin, Jian-Guang Lou, Dongmei Zhang 0001
AAAI3
2021 Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization
abstract
Human intelligence exhibits compositional generalization (i.e., the capacity to understand and produce unseen combinations of seen components), but current neural seq2seq models lack such ability. In this paper, we revisit iterative back-translation, a simple yet effective semi-supervised method, to investigate whether and how it can improve compositional generalization. In this work: (1) We first empirically show that iterative back-translation substantially improves the performance on compositional generalization benchmarks (CFQ and SCAN). (2) To understand why iterative back-translation is useful, we carefully examine the performance gains and find that iterative back-translation can increasingly correct errors in pseudo-parallel data. (3) To further encourage this mechanism, we propose curriculum iterative back-translation, which better improves the quality of pseudo-parallel data, thus further improving the performance.
Yinuo Guo, Hualei Zhu, Zeqi Lin, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001
AAAI5
2021 Chase: A Large-Scale and Pragmatic Chinese Dataset for Cross-Database Context-Dependent Text-to-SQL
abstract
Jiaqi Guo, Ziliang Si, Yu Wang, Qian Liu, Ming Fan, Jian-Guang Lou, Zijiang Yang, Ting Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ziliang Si, Yu Wang 0093, Qian Liu 0033, Ming Fan 0002, Jian-Guang Lou, Zijiang Yang 0006, Ting Liu 0002
ACL/IJCNLP (1)6
2021 Translating Headers of Tabular Data: A Pilot Study of Schema Translation
abstract
Schema translation is the task of automatically translating headers of tabular data from one language to another.High-quality schema translation plays an important role in crosslingual table searching, understanding and analysis.Despite its importance, schema translation is not well studied in the community, and state-of-the-art neural machine translation models cannot work well on this task because of two intrinsic differences between plain text and tabular data: morphological difference and context difference.To facilitate the research study, we construct the first parallel dataset for schema translation, which consists of 3,158 tables with 11,979 headers written in 6 different languages, including English, Chinese, French, German, Spanish, and Japanese.Also, we propose the first schema translation model called CAST, which is a header-to-header neural machine translation model augmented with schema context.Specifically, we model a target header and its context as a directed graph to represent their entity types and relations.Then CAST encodes the graph with a relational-aware transformer and uses another transformer to decode the header in the target language.Experiments on our dataset demonstrate that CAST significantly outperforms state-of-the-art neural machine translation models.Our dataset will be released at https://github.com/microsoft/ContextualSP.
Kunrui Zhu, Yan Gao 0002, Jian-Guang Lou
EMNLP (1)4
2021 Keep the Structure: A Latent Shift-Reduce Parser for Semantic Parsing
abstract
Traditional end-to-end semantic parsing models treat a natural language utterance as a holonomic structure. However, hierarchical structures exist in natural languages, which also align with the hierarchical structures of logical forms. In this paper, we propose a latent shift-reduce parser, called LASP, which decomposes both natural language queries and logical form expressions according to their hierarchical structures and finds local alignment between them to enhance semantic parsing. LASP consists of a base parser and a shift-reduce splitter. The splitter dynamically separates an NL query into several spans. The base parser converts the relevant simple spans into logical forms, which are further combined to obtain the final logical form. We conducted empirical studies on two datasets across different domains and different types of logical forms. The results demonstrate that the proposed method significantly improves the performance of semantic parsing, especially on unseen scenarios.
Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001
IJCAI5
2021 Can Neural Clone Detection Generalize to Unseen Functionalitiesƒ
abstract
Many recently proposed code clone detectors exploit neural networks to capture latent semantics of source code, thus achieving impressive results for detecting semantic clones. These neural clone detectors rely on the availability of large amounts of labeled training data. We identify a key oversight in the current evaluation methodology for neural clone detection: cross-functionality generalization (i.e., detecting semantic clones of which the functionalities are unseen in training). Specifically, we focus on this question: do neural clone detectors truly learn the ability to detect semantic clones, or they just learn how to model specific functionalities in training data while cannot generalize to realistic unseen functionalitiesƒ This paper investigates how the generalizability can be evaluated and improved.Our contributions are 3-folds: (1) We propose an evaluation methodology that can systematically measure the cross-functionality generalizability of neural clone detection. Based on this evaluation methodology, an empirical study is conducted and the results indicate that current neural clone detectors cannot generalize well as expected. (2) We conduct empirical analysis to understand key factors that can impact the generalizability. We investigate 3 factors: training data diversity, vocabulary, and locality. Results show that the performance loss on unseen functionalities can be reduced through addressing the out-of-vocabulary problem and increasing training data diversity. (3) We propose a human-in-the-loop mechanism that help adapt neural clone detectors to new code repositories containing lots of unseen functionalities. It improves annotation efficiency with the combination of transfer learning and active learning. Experimental results show that it reduces the amount of annotations by about 88%. Our code and data are publicly available1.
Chenyao Liu, Zeqi Lin, Jian-Guang Lou, Lijie Wen 0001, Dongmei Zhang 0001
ASE3
2021 AnaSearch: Extract, Retrieve and Visualize Structured Results from Unstructured Text for Analytical Queries
abstract
Modern search engines retrieve results mainly based on the keyword matching techniques, and thus fail to answer analytical queries like "apps with more than 1 billion monthly active users" or "population growth of the US from 2015 to 2019", which requires numerical reasoning or aggregating results from multiple web pages. Such analytical queries are very common in the data analysis area, the expected results would be structured tables or charts. In most cases, these structured results are not available or accessible, they scatter in various text sources. In this work, we build AnaSearch, a search system to support analytical queries, and return structured results that can be visualized in the form of tables or charts. We collect and build structured quantitative data from the unstructured text on the web automatically. With AnaSearch, data analysts could easily derive insights for decision making with keyword or natural language queries. Specifically, we build AnaSearch under the COVID-19 news data, which makes it easy to compare with manually collected structured data.
Tongliang Li, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001, Dongmei Zhang 0001
WSDM3
2021 Retrieve-Then-Adapt: Example-based Automatic Generation for Proportion-related Infographics
abstract
Infographic is a data visualization technique which combines graphic and textual descriptions in an aesthetic and effective manner. Creating infographics is a difficult and time-consuming process which often requires significant attempts and adjustments even for experienced designers, not to mention novice users with limited design expertise. Recently, a few approaches have been proposed to automate the creation process by applying predefined blueprints to user information. However, predefined blueprints are often hard to create, hence limited in volume and diversity. In contrast, good infogrpahics have been created by professionals and accumulated on the Internet rapidly. These online examples often represent a wide variety of design styles, and serve as exemplars or inspiration to people who like to create their own infographics. Based on these observations, we propose to generate infographics by automatically imitating examples. We present a two-stage approach, namely retrieve-then-adapt. In the retrieval stage, we index online examples by their visual elements. For a given user information, we transform it to a concrete query by sampling from a learned distribution about visual elements, and then find appropriate examples in our example library based on the similarity between example indexes and the query. For a retrieved example, we generate an initial drafts by replacing its content with user information. However, in many cases, user information cannot be perfectly fitted to retrieved examples. Therefore, we further introduce an adaption stage. Specifically, we propose a MCMC-like approach and leverage recursive neural networks to help adjust the initial draft and improve its visual appearance iteratively, until a satisfactory result is obtained. We implement our approach on widely-used proportion-related infographics, and demonstrate its effectiveness by sample results and expert reviews.
Chunyao Qian, Shizhao Sun, Weiwei Cui 0001, Jian-Guang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.4
2020 You Impress Me: Dialogue Generation via Mutual Persona Perception
abstract
Despite the continuing efforts to improve the engagingness and consistency of chit-chat dialogue systems, the majority of current work simply focus on mimicking human-like responses, leaving understudied the aspects of modeling understanding between interlocutors.The research in cognitive science, instead, suggests that understanding is an essential signal for a high-quality chit-chat conversation.Motivated by this, we propose P 2 BOT, a transmitter-receiver based framework with the aim of explicitly modeling understanding.Specifically, P 2 BOT incorporates mutual persona perception to enhance the quality of personalized dialogue generation.Experiments on a large public dataset, PERSONA-CHAT, demonstrate the effectiveness of our approach, with a considerable boost over the state-of-theart baselines across both automatic metrics and human evaluations.
Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001
ACL4
2020 Benchmarking Meaning Representations in Neural Semantic Parsing
abstract
Meaning representation is an important component of semantic parsing.Although researchers have designed a lot of meaning representations, recent work focuses on only a few of them.Thus, the impact of meaning representation on semantic parsing is less understood.Furthermore, existing work's performance is often not comprehensively evaluated due to the lack of readily-available execution engines.Upon identifying these gaps, we propose UNIMER, a new unified benchmark on meaning representations, by integrating existing semantic parsing datasets, completing the missing logical forms, and implementing the missing execution engines.The resulting unified benchmark contains the complete enumeration of logical forms and execution engines over three datasets × four meaning representations.A thorough experimental study on UNIMER reveals that neural semantic parsing approaches exhibit notably different performance when they are trained to generate different meaning representations.Also, program alias and grammar rules heavily impact the performance of different meaning representations.Our benchmark, execution engines and implementation can be found on: https
Qian Liu 0033, Jian-Guang Lou, Zhenwen Li, Xueqing Liu 0001, Tao Xie 0001, Ting Liu 0002
EMNLP (1)3
2020 "What Do You Mean by That?" A Parser-Independent Interactive Approach for Enhancing Text-to-SQL
abstract
In Natural Language Interfaces to Databases systems, the text-to-SQL technique allows users to query databases by using natural language questions. Though significant progress in this area has been made recently, most parsers may fall short when they are deployed in real systems. One main reason stems from the difficulty of fully understanding the users' natural language questions. In this paper, we include human in the loop and present a novel parser-independent interactive approach (PIIA) that interacts with users using multi-choice questions and can easily work with arbitrary parsers. Experiments were conducted on two cross-domain datasets, the WikiSQL and the more complex Spider, with five state-of-the-art parsers. These demonstrated that PIIA is capable of enhancing the text-to-SQL performance with limited interaction turns by using both simulation and human evaluation.
Bei Chen 0008, Qian Liu 0033, Yan Gao 0002, Jian-Guang Lou, Yan Zhang 0117, Dongmei Zhang 0001
EMNLP (1)5
2020 Incomplete Utterance Rewriting as Semantic Segmentation
abstract
Recent years the task of incomplete utterance rewriting has raised a large attention.Previous works usually shape it as a machine translation task and employ sequence to sequence based architecture with copy mechanism.In this paper, we present a novel and extensive approach, which formulates it as a semantic segmentation task.Instead of generating from scratch, such a formulation introduces edit operations and shapes the problem as prediction of a word-level edit matrix.Benefiting from being able to capture both local and global information, our approach achieves state-ofthe-art performance on several public datasets.Furthermore, our approach is four times faster than the standard approach in inference.
Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001
EMNLP (1)3
2020 How Far are We from Effective Context Modeling? An Exploratory Study on Semantic Parsing in Context
abstract
Recently semantic parsing in context has received a considerable attention, which is challenging since there are complex contextual phenomena. Previous works verified their proposed methods in limited scenarios, which motivates us to conduct an exploratory study on context modeling methods under real-world semantic parsing in context. We present a grammar-based decoding semantic parser and adapt typical context modeling methods on top of it. We evaluate 13 context modeling methods on two large complex cross-domain datasets, and our best model achieves state-of-the-art performances on both datasets with significant improvements. Furthermore, we summarize the most frequent contextual phenomena, with a fine-grained analysis on representative models, which may shed light on potential research directions. Our code is available at https://github.com/microsoft/ContextualSP.
Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001
IJCAI4
2020 RECPARSER: A Recursive Semantic Parsing Framework for Text-to-SQL Task
abstract
Neural semantic parsers usually fail to parse long and complicated utterances into nested SQL queries, due to the large search space. In this paper, we propose a novel recursive semantic parsing framework called RECPARSER to generate the nested SQL query layer-by-layer. It decomposes the complicated nested SQL query generation problem into several progressive non-nested SQL query generation problems. Furthermore, we propose a novel Question Decomposer module to explicitly encourage RECPARSER to focus on different components of an utterance when predicting SQL queries of different layers. Experiments on the Spider dataset show that our approach is more effective compared to the previous works at predicting the nested SQL queries. In addition, we achieve an overall accuracy that is comparable with state-of-the-art approaches.
Yan Gao 0002, Bei Chen 0008, Qian Liu 0033, Jian-Guang Lou, Fei Teng 0001, Dongmei Zhang 0001
IJCAI6
2020 Hierarchical Poset Decoding for Compositional Generalization in Language
abstract
We formalize human language understanding as a structured prediction task where the output is a partially ordered set (poset). Current encoder-decoder architectures do not take the poset structure of semantics into account properly, thus suffering from poor compositional generalization ability. In this paper, we propose a novel hierarchical poset decoding paradigm for compositional generalization in language. Intuitively: (1) the proposed paradigm enforces partial permutation invariance in semantics, thus avoiding overfitting to bias ordering information; (2) the hierarchical mechanism allows to capture high-level structures of posets. We evaluate our proposed decoder on Compositional Freebase Questions (CFQ), a large and realistic natural language question answering dataset that is specifically designed to measure compositional generalization. Results show that it outperforms current decoders.
Yinuo Guo, Zeqi Lin, Jian-Guang Lou, Dongmei Zhang 0001
NeurIPS3
2020 Compositional Generalization by Learning Analytical Expressions
abstract
Compositional generalization is a basic and essential intellective capability of human beings, which allows us to recombine known parts readily. However, existing neural network based models have been proven to be extremely deficient in such a capability. Inspired by work in cognition which argues compositionality can be captured by variable slots with symbolic functions, we present a refreshing view that connects a memory-augmented neural model with analytical expressions, to achieve compositional generalization. Our model consists of two cooperative neural modules, Composer and Solver, fitting well with the cognitive argument while being able to be trained in an end-to-end manner via a hierarchical reinforcement learning algorithm. Experiments on the well-known benchmark SCAN demonstrate that our model seizes a great ability of compositional generalization, solving all challenges addressed by previous works with 100% accuracies.
Qian Liu 0033, Shengnan An, Jian-Guang Lou, Bei Chen 0008, Zeqi Lin, Yan Gao 0002, Nanning Zheng 0001, Dongmei Zhang 0001
NeurIPS3
2020 Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language Statements
abstract
Combining data content with visual embellishments, infographics can effectively deliver messages in an engaging and memorable manner. Various authoring tools have been proposed to facilitate the creation of infographics. However, creating a professional infographic with these authoring tools is still not an easy task, requiring much time and design expertise. Therefore, these tools are generally not attractive to casual users, who are either unwilling to take time to learn the tools or lacking in proper design expertise to create a professional infographic. In this paper, we explore an alternative approach: to automatically generate infographics from natural language statements. We first conducted a preliminary study to explore the design space of infographics. Based on the preliminary study, we built a proof-of-concept system that automatically converts statements about simple proportion-related statistics to a set of infographics with pre-designed styles. Finally, we demonstrated the usability and usefulness of the system through sample results, exhibits, and expert reviews.
Weiwei Cui 0001, Xiaoyu Zhang 0014, Yun Wang 0012, Bei Chen 0008, Lei Fang 0004, Jian-Guang Lou, Dongmei Zhang 0001
IEEE Trans. Vis. Comput. Graph.8
2019 FANDA: A Novel Approach to Perform Follow-Up Query Analysis
abstract
Recent work on Natural Language Interfaces to Databases (NLIDB) has attracted considerable attention. NLIDB allow users to search databases using natural language instead of SQL-like query languages. While saving the users from having to learn query languages, multi-turn interaction with NLIDB usually involves multiple queries where contextual information is vital to understand the users’ query intents. In this paper, we address a typical contextual understanding problem, termed as follow-up query analysis. In spite of its ubiquity, follow-up query analysis has not been well studied due to two primary obstacles: the multifarious nature of follow-up query scenarios and the lack of high-quality datasets. Our work summarizes typical follow-up query scenarios and provides a new FollowUp dataset with 1000 query triples on 120 tables. Moreover, we propose a novel approach FANDA, which takes into account the structures of queries and employs a ranking model with weakly supervised max-margin learning. The experimental results on FollowUp demonstrate the superiority of FANDA over multiple baselines across multiple metrics.
Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Dongmei Zhang 0001
AAAI3
2019 Leveraging Web Semantic Knowledge in Word Representation Learning
Haoyan Liu 0001, Lei Fang 0004, Jian-Guang Lou, Zhoujun Li 0001
AAAI3
2019 Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation
abstract
We present a neural approach called IRNet for complex and cross-domain Text-to-SQL.IR-Net aims to address two challenges: 1) the mismatch between intents expressed in natural language (NL) and the implementation details in SQL; 2) the challenge in predicting columns caused by the large number of outof-domain words.Instead of end-to-end synthesizing a SQL query, IRNet decomposes the synthesis process into three phases.In the first phase, IRNet performs a schema linking over a question and a database schema.Then, IRNet adopts a grammar-based neural model to synthesize a SemQL query which is an intermediate representation that we design to bridge NL and SQL.Finally, IRNet deterministically infers a SQL query from the synthesized SemQL query with domain knowledge.On the challenging Text-to-SQL benchmark Spider, IRNet achieves 46.7% accuracy, obtaining 19.5% absolute improvement over previous state-of-the-art approaches.At the time of writing, IRNet achieves the first position on the Spider leaderboard. * Equal Contributions. Work done during an internship at MSRA.NL: Show the names of students who have a grade higher than 5 and have at least 2 friends. SQL: SELECT T1.name FROM friend AS T1 JOIN highschooler AS T2ON T1.student_id = T2.idWHERE T2
Zecheng Zhan, Yan Gao 0002, Jian-Guang Lou, Ting Liu 0002, Dongmei Zhang 0001
ACL (1)5
2019 Data-Anonymous Encoding for Text-to-SQL Generation
abstract
Zhen Dong, Shizhao Sun, Hongzhi Liu, Jian-Guang Lou, Dongmei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shizhao Sun, Hongzhi Liu 0001, Jian-Guang Lou, Dongmei Zhang 0001
EMNLP/IJCNLP (1)4
2019 A Split-and-Recombine Approach for Follow-up Query Analysis
abstract
Qian Liu, Bei Chen, Haoyan Liu, Jian-Guang Lou, Lei Fang, Bin Zhou, Dongmei Zhang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Qian Liu 0033, Bei Chen 0008, Haoyan Liu 0001, Jian-Guang Lou, Lei Fang 0004, Dongmei Zhang 0001
EMNLP/IJCNLP (1)4
2019 Leveraging Adjective-Noun Phrasing Knowledge for Comparison Relation Prediction in Text-to-SQL
abstract
Haoyan Liu, Lei Fang, Qian Liu, Bei Chen, Jian-Guang Lou, Zhoujun Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Haoyan Liu 0001, Lei Fang 0004, Qian Liu 0033, Bei Chen 0008, Jian-Guang Lou, Zhoujun Li 0001
EMNLP/IJCNLP (1)5
2019 Sara: self-replay augmented record and replay for Android in industrial cases
abstract
Record-and-replay tools are indispensable for quality assurance of mobile applications. Due to its importance, an increasing number of tools are being developed to record and replay user interactions for Android. However, by conducting an empirical study of various existing tools in industrial settings, researchers have revealed a gap between the characteristics requested from industry and the performance of publicly available record-and-replay tools. The study concludes that no existing tools under evaluation are sufficient for industrial applications. In this paper, we present a record-and-replay tool called SARA towards bridging the gap and targeting a wide adoption. Specifically, a dynamic instrumentation technique is used to accommodate rich sources of inputs in the application layer satisfying various constraints requested from industry. A self-replay mechanism is proposed to record more information of user inputs for accurate replaying without degrading user experience. In addition, an adaptive replay method is designed to enable replaying events on different devices with diverse screen sizes and OS versions. Through an evaluation on 53 highly popular industrial Android applications and 265 common usage scenarios, we demonstrate the effectiveness of SARA in recording and replaying rich sources of inputs on the same or different devices.
Shuyue Li, Jian-Guang Lou, Zijiang Yang 0006, Ting Liu 0002
ISSTA3
2019 λOpt: Learn to Regularize Recommender Models in Finer Levels
abstract
Recommendation models mainly deal with categorical variables, such as user/item ID and attributes. Besides the high-cardinality issue, the interactions among such categorical variables are usually long-tailed, with the head made up of highly frequent values and a long tail of rare ones. This phenomenon results in the data sparsity issue, making it essential to regularize the models to ensure generalization. The common practice is to employ grid search to manually tune regularization hyperparameters based on the validation data. However, it requires non-trivial efforts and large computation resources to search the whole candidate space; even so, it may not lead to the optimal choice, for which different parameters should have different regularization strengths. In this paper, we propose a hyperparameter optimization method, lambdaOpt, which automatically and adaptively enforces regularization during training. Specifically, it updates the regularization coefficients based on the performance of validation data. With lambdaOpt, the notorious tuning of regularization hyperparameters can be avoided; more importantly, it allows fine-grained regularization (i.e. each parameter can have an individualized regularization coefficient), leading to better generalized models. We show how to employ lambdaOpt on matrix factorization, a classical model that is representative of a large family of recommender models. Extensive experiments on two public benchmarks demonstrate the superiority of our method in boosting the performance of top-K recommendation.
Bei Chen 0008, Xiangnan He 0001, Chen Gao 0001, Yong Li 0008, Jian-Guang Lou, Yue Wang 0007
KDD6
2019 Robust log-based anomaly detection on unstable log data
abstract
Logs are widely used by large and complex software-intensive systems for troubleshooting. There have been a lot of studies on log-based anomaly detection. To detect the anomalies, the existing methods mainly construct a detection model using log event data extracted from historical logs. However, we find that the existing methods do not work well in practice. These methods have the close-world assumption, which assumes that the log data is stable over time and the set of distinct log events is known. However, our empirical study shows that in practice, log data often contains previously unseen log events or log sequences. The instability of log data comes from two sources: 1) the evolution of logging statements, and 2) the processing noise in log data. In this paper, we propose a new log-based anomaly detection approach, called LogRobust. LogRobust extracts semantic information of log events and represents them as semantic vectors. It then detects anomalies by utilizing an attention-based Bi-LSTM model, which has the ability to capture the contextual information in the log sequences and automatically learn the importance of different log events. In this way, LogRobust is able to identify and handle unstable log events and sequences. We have evaluated LogRobust using logs collected from the Hadoop system and an actual online service system of Microsoft. The experimental results show that the proposed approach can well address the problem of log instability and achieve accurate and robust results on real-world, ever-changing log data.
Xu Zhang 0024, Yong Xu 0010, Qingwei Lin, Bo Qiao 0001, Hongyu Zhang 0002, Yingnong Dang, Chunyu Xie, Xinsheng Yang, Ze Li 0005, Junjie Chen 0003, Xiaoting He 0003, Randolph Yao, Jian-Guang Lou, Murali Chintalapati, Furao Shen, Dongmei Zhang 0001
ESEC/SIGSOFT FSE14
2018 SemRegex: A Semantics-Based Approach for Generating Regular Expressions from Natural Language Specifications
abstract
Recent research proposes syntax-based approaches to address the problem of generating programs from natural language specifications.These approaches typically train a sequence-to-sequence learning model using a syntax-based objective: maximum likelihood estimation (MLE).Such syntax-based approaches do not effectively address the goal of generating semantically correct programs, because these approaches fail to handle Program Aliasing, i.e., semantically equivalent programs may have many syntactically different forms.To address this issue, in this paper, we propose a semantics-based approach named SemRegex.SemRegex provides solutions for a subtask of the program-synthesis problem: generating regular expressions from natural language.Different from the existing syntax-based approaches, SemRegex trains the model by maximizing the expected semantic correctness of the generated regular expressions.The semantic correctness is measured using the DFA-equivalence oracle, random test cases, and distinguishing test cases.The experiments on three public datasets demonstrate the superiority of SemRegex over the existing state-of-the-art approaches.
Zexuan Zhong, Wei Yang 0013, Jian Peng 0001, Tao Xie 0001, Jian-Guang Lou, Ting Liu 0002, Dongmei Zhang 0001
EMNLP6
2018 Learning-to-Ask: Knowledge Acquisition via 20 Questions
abstract
Almost all the knowledge empowered applications rely upon accurate knowledge, which has to be either collected manually with high cost, or extracted automatically with unignorable errors. In this paper, we study 20 Questions, an online interactive game where each question-response pair corresponds to a fact of the target entity, to acquire highly accurate knowledge effectively with nearly zero labor cost. Knowledge acquisition via 20 Questions predominantly presents two challenges to the intelligent agent playing games with human players. The first one is to seek enough information and identify the target entity with as few questions as possible, while the second one is to leverage the remaining questioning opportunities to acquire valuable knowledge effectively, both of which count on good questioning strategies. To address these challenges, we propose the Learning-to-Ask (LA) framework, within which the agent learns smart questioning strategies for information seeking and knowledge acquisition by means of deep reinforcement learning and generalized matrix factorization respectively. In addition, a Bayesian approach to represent knowledge is adopted to ensure robustness to noisy user responses. Simulating experiments on real data show that LA is able to equip the agent with effective questioning strategies, which result in high winning rates and rapid knowledge acquisition. Moreover, the questioning strategies for information seeking and knowledge acquisition boost the performance of each other, allowing the agent to start with a relatively small knowledge set and quickly improve its knowledge base in the absence of constant human supervision.
Bei Chen 0008, Xuguang Duan, Jian-Guang Lou, Yue Wang 0007, Wenwu Zhu 0001
KDD4
2018 BigIN4: Instant, Interactive Insight Identification for Multi-Dimensional Big Data
abstract
The ability to identify insights from multi-dimensional big data is important for business intelligence. To enable interactive identification of insights, a large number of dimension combinations need to be searched and a series of aggregation queries need to be quickly answered. The existing approaches answer interactive queries on big data through data cubes or approximate query processing. However, these approaches can hardly satisfy the performance or accuracy requirements for ad-hoc queries demanded by interactive exploration. In this paper, we present BigIN4, a system for instant, interactive identification of insights from multi-dimensional big data. BigIN4 gives insight suggestions by enumerating subspaces and answers queries by combining data cube and approximate query processing techniques. If a query cannot be answered by the cubes, BigIN4 decomposes it into several low dimensional queries that can be directly answered by the cubes through an online constructed Bayesian Network and gives an approximate answer within a statistical interval. Unlike the related works, BigIN4 does not require any prior knowledge of queries and does not assume a certain data distribution. Our experiments on ten real-world large-scale datasets show that BigIN4 can successfully identify insights from big data. Furthermore, BigIN4 can provide approximate answers to aggregation queries effectively (with less than 10% error on average) and efficiently (50x faster than sampling-based methods).
Qingwei Lin, Weichen Ke, Jian-Guang Lou, Hongyu Zhang 0002, Kaixin Sui, Yong Xu 0010, Bo Qiao 0001, Dongmei Zhang 0001
KDD3
2018 Identifying impactful service system problems via log analysis
abstract
Logs are often used for troubleshooting in large-scale software systems. For a cloud-based online system that provides 24/7 service, a huge number of logs could be generated every day. However, these logs are highly imbalanced in general, because most logs indicate normal system operations, and only a small percentage of logs reveal impactful problems. Problems that lead to the decline of system KPIs (Key Performance Indicators) are impactful and should be fixed by engineers with a high priority. Furthermore, there are various types of system problems, which are hard to be distinguished manually. In this paper, we propose Log3C, a novel clustering-based approach to promptly and precisely identify impactful system problems, by utilizing both log sequences (a sequence of log events) and system KPIs. More specifically, we design a novel cascading clustering algorithm, which can greatly save the clustering time while keeping high accuracy by iteratively sampling, clustering, and matching log sequences. We then identify the impactful problems by correlating the clusters of log sequences with system KPIs. Log3C is evaluated on real-world log data collected from an online service system at Microsoft, and the results confirm its effectiveness and efficiency. Furthermore, our approach has been successfully applied in industrial practice.
Shilin He, Qingwei Lin, Jian-Guang Lou, Hongyu Zhang 0002, Michael R. Lyu, Dongmei Zhang 0001
ESEC/SIGSOFT FSE3
2018 Predicting Node failure in cloud service systems
abstract
In recent years, many traditional software systems have migrated to cloud computing platforms and are provided as online services. The service quality matters because system failures could seriously affect business and user experience. A cloud service system typically contains a large number of computing nodes. In reality, nodes may fail and affect service availability. In this paper, we propose a failure prediction technique, which can predict the failure-proneness of a node in a cloud service system based on historical data, before node failure actually happens. The ability to predict faulty nodes enables the allocation and migration of virtual machines to the healthy nodes, therefore improving service availability. Predicting node failure in cloud service systems is challenging, because a node failure could be caused by a variety of reasons and reflected by many temporal and spatial signals. Furthermore, the failure data is highly imbalanced. To tackle these challenges, we propose MING, a novel technique that combines: 1) a LSTM model to incorporate the temporal data, 2) a Random Forest model to incorporate spatial data; 3) a ranking model that embeds the intermediate results of the two models as feature inputs and ranks the nodes by their failure-proneness, 4) a cost-sensitive function to identify the optimal threshold for selecting the faulty nodes. We evaluate our approach using real-world data collected from a cloud service system. The results confirm the effectiveness of the proposed approach. We have also successfully applied the proposed approach in real industrial practice.
Qingwei Lin, Ken Hsieh, Yingnong Dang, Hongyu Zhang 0002, Kaixin Sui, Yong Xu 0010, Jian-Guang Lou, Chenggang Li, Youjiang Wu, Randolph Yao, Murali Chintalapati, Dongmei Zhang 0001
ESEC/SIGSOFT FSE7
2018 Improving Service Availability of Cloud Systems by Predicting Disk Error
Yong Xu 0010, Kaixin Sui, Randolph Yao, Hongyu Zhang 0002, Qingwei Lin, Yingnong Dang, Peng Li 0062, Keceng Jiang, Wenchi Zhang, Jian-Guang Lou, Murali Chintalapati, Dongmei Zhang 0001
USENIX ATC10
2017 Experience report on applying software analytics in incident management of online service
Jian-Guang Lou, Qingwei Lin, Rui Ding 0001, Qiang Fu 0015, Dongmei Zhang 0001, Tao Xie 0001
Autom. Softw. Eng.1
2016 iDice: problem identification for emerging issues
abstract
One challenge for maintaining a large-scale software system, especially an online service system, is to quickly respond to customer issues. The issue reports typically have many categorical attributes that reflect the characteristics of the issues. For a commercial system, most of the time the volume of reported issues is relatively constant. Sometimes, there are emerging issues that lead to significant volume increase. It is important for support engineers to efficiently and effectively identify and resolve such emerging issues, since they have impacted a large number of customers. Currently, problem identification for an emerging issue is a tedious and error-prone process, because it requires support engineers to manually identify a particular attribute combination that characterizes the emerging issue among a large number of attribute combinations. We call such an attribute combination effective combination, which is important for issue isolation and diagnosis. In this paper, we propose iDice, an approach that can identify the effective combination for an emerging issue with high quality and performance. We evaluate the effectiveness and efficiency of iDice through experiments. We have also successfully applied iDice to several Microsoft online service systems in production. The results confirm that iDice can help identify emerging issues and reduce maintenance effort.
Qingwei Lin, Jian-Guang Lou, Hongyu Zhang 0002, Dongmei Zhang 0001
ICSE2
2016 Roundtable: Research Opportunities and Challenges for Large-Scale Software Systems
Xusheng Xiao, Jian-Guang Lou, Shan Lu 0001, David C. Shepherd, Xin Peng 0001, Qianxiang Wang
J. Comput. Sci. Technol.2
2015 An Empirical Study on Quality Issues of Production Big Data Platform
abstract
Big Data computing platform has evolved to be a multi-tenant service. The service quality matters because system failure or performance slowdown could adversely affect business and user experience. There is few study in literature on service quality issues of production Big Data computing platform. In this paper, we present an empirical study on the service quality issues of Microsoft ProductA, which is a company-wide multi-tenant Big Data computing platform, serving thousands of customers from hundreds of teams. ProductA has a well-defined incident management process, which helps customers report and mitigate service quality issues on 24/7 basis. This paper explores the common symptom, causes and mitigation of service quality issues in Big Data computing. We conduct an empirical study on 210 real service quality issues in ProductA. Our major findings include (1) 21.0% of escalations are caused by hardware faults; (2) 36.2% are caused by system side defects; (3) 37.2% are due to customer side faults. We also studied the general diagnosis process and the commonly adopted mitigation solutions. Our findings can help improve current development and maintenance practice of Big Data computing platform, and motivate tool support.
Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Haoxiang Lin, Tingting Qin
ICSE (2)2
2015 CodeHow: Effective Code Search Based on API Understanding and Extended Boolean Model (E)
abstract
Over the years of software development, a vast amount of source code has been accumulated. Many code search tools were proposed to help programmers reuse previously-written code by performing free-text queries over a large-scale codebase. Our experience shows that the accuracy of these code search tools are often unsatisfactory. One major reason is that existing tools lack of query understanding ability. In this paper, we propose CodeHow, a code search technique that can recognize potential APIs a user query refers to. Having understood the potentially relevant APIs, CodeHow expands the query with the APIs and performs code retrieval by applying the Extended Boolean model, which considers the impact of both text similarity and potential APIs on code search. We deploy the backend of CodeHow as a Microsoft Azure service and implement the front-end as a Visual Studio extension. We evaluate CodeHow on a large-scale codebase consisting of 26K C# projects downloaded from GitHub. The experimental results show that when the top 1 results are inspected, CodeHow achieves a precision score of 0.794 (i.e., 79.4% of the first returned results are relevant code snippets). The results also show that CodeHow outperforms conventional code search tools. Furthermore, we perform a controlled experiment and a survey of Microsoft developers. The results confirm the usefulness and effectiveness of CodeHow in programming practices.
Fei Lv 0002, Hongyu Zhang 0002, Jian-Guang Lou, Shaowei Wang 0002, Dongmei Zhang 0001, Jianjun Zhao 0001
ASE3
2015 Log2: A Cost-Aware Logging Mechanism for Performance Diagnosis
Rui Ding 0001, Hucheng Zhou, Jian-Guang Lou, Hongyu Zhang 0002, Qingwei Lin, Qiang Fu 0015, Dongmei Zhang 0001, Tao Xie 0001
USENIX ATC3
2014 Mining Historical Issue Repositories to Heal Large-Scale Online Service Systems
abstract
Online service systems have been increasingly popular and important nowadays. Reducing the MTTR (Mean Time to Restore) of a service remains one of the most important steps to assure the user-perceived availability of the service. To reduce the MTTR, a common practice is to restore the service by identifying and applying an appropriate healing action. In this paper, we present an automated mining-based approach for suggesting an appropriate healing action for a given new issue. Our approach suggests an appropriate healing action by adapting healing actions from the retrieved similar historical issues. We have applied our approach to a real-world and large-scale product online service. The studies on 243 real issues of the service show that our approach can effectively suggest appropriate healing actions (with 87% accuracy) to reduce the MTTR of the service. In addition, according to issue characteristics, we further study and categorize issues where automatic healing suggestion faces difficulties.
Rui Ding 0001, Qiang Fu 0015, Jian-Guang Lou, Qingwei Lin, Dongmei Zhang 0001, Tao Xie 0001
DSN3
2014 Identifying Recurrent and Unknown Performance Issues
abstract
For a large-scale software system, especially an online service system, when a performance issue occurs, it is desirable to check whether this issue has occurred before. If there are past similar issues, a known remedy could be applied. Otherwise, a new troubleshooting process may have to be initiated. The symptom of a performance issue can be characterized by a set of metrics. Due to the sophisticated nature of software systems, manual diagnosis of performance issues based on metric data is typically expensive and laborious. In this paper, we propose a Hidden Markov Random Field (HMRF) based approach to automatic identification of recurrent and unknown performance issues. We formulate the problem of issue identification as a HMRF-based clustering problem. Our approach incorporates the learning of metric discretization thresholds and the optimization of issue clustering. Based on the learned thresholds and cluster centroids, we can achieve accurate identification of recurrent issues and unknown issues. Experimental evaluations on an open benchmark and a large-scale industrial production system show that our approach is effective and outperforms the related state-of-the-art approaches.
Meng-Hui Lim, Jian-Guang Lou, Hongyu Zhang 0002, Qiang Fu 0015, Andrew Beng Jin Teoh, Qingwei Lin, Rui Ding 0001, Dongmei Zhang 0001
ICDM2
2014 Correlating events with time series for incident diagnosis
abstract
As online services have more and more popular, incident diagnosis has emerged as a critical task in minimizing the service downtime and ensuring high quality of the services provided. For most online services, incident diagnosis is mainly conducted by analyzing a large amount of telemetry data collected from the services at runtime. Time series data and event sequence data are two major types of telemetry data. Techniques of correlation analysis are important tools that are widely used by engineers for data-driven incident diagnosis. Despite their importance, there has been little previous work addressing the correlation between two types of heterogeneous data for incident diagnosis: continuous time series data and temporal event data. In this paper, we propose an approach to evaluate the correlation between time series data and event data. Our approach is capable of discovering three important aspects of event-timeseries correlation in the context of incident diagnosis: existence of correlation, temporal order, and monotonic effect. Our experimental results on simulation data sets and two real data sets demonstrate the effectiveness of the algorithm.
Jian-Guang Lou, Qingwei Lin, Qiang Fu 0015, Rui Ding 0001, Dongmei Zhang 0001, Zhe Wang 0007
KDD2
2014 Querying sequential software engineering data
abstract
We propose a pattern-based approach to effectively and efficiently analyzing sequential software engineering (SE) data. Different from other types of SE data, sequential SE data preserves unique temporal properties, which cannot be easily analyzed without much programming effort. In order to facilitate the analysis of sequential SE data, we design a sequential pattern query language (SPQL), which specifies the temporal properties based on regular expressions, and is enhanced with variables and statements to store and manipulate matching states. We also propose a query engine to effectively process the SPQL queries. We have applied our approach to analyze two types of SE data, namely bug report history and source code change history. We experiment with 181,213 Eclipse bug reports and 323,989 code revisions of Android. SPQL enables us to explore interesting temporal properties underneath these sequential data with a few lines of query code and low matching overhead. The analysis results can help better under- stand a software process and identify process violations.
Chengnian Sun, Jian-Guang Lou, Hongyu Zhang 0002, Dongmei Zhang 0001, Siau-Cheng Khoo
SIGSOFT FSE3
2013 Software analytics for incident management of online services: An experience report
abstract
As online services become more and more popular, incident management has become a critical task that aims to minimize the service downtime and to ensure high quality of the provided services. In practice, incident management is conducted through analyzing a huge amount of monitoring data collected at runtime of a service. Such data-driven incident management faces several significant challenges such as the large data scale, complex problem space, and incomplete knowledge. To address these challenges, we carried out two-year software-analytics research where we designed a set of novel data-driven techniques and developed an industrial system called the Service Analysis Studio (SAS) targeting real scenarios in a large-scale online service of Microsoft. SAS has been deployed to worldwide product datacenters and widely used by on-call engineers for incident management. This paper shares our experience about using software analytics to solve engineers' pain points in incident management, the developed data-analysis techniques, and the lessons learned from the process of research development and technology transfer.
Jian-Guang Lou, Qingwei Lin, Rui Ding 0001, Qiang Fu 0015, Dongmei Zhang 0001, Tao Xie 0001
ASE1
2013 Contextual analysis of program logs for understanding system behaviors
abstract
Understanding the behaviors of a software system is very important for performing daily system maintenance tasks. In practice, one way to gain knowledge about the runtime behavior of a system is to manually analyze system logs collected during the system executions. With the increasing scale and complexity of software systems, it has become challenging for system operators to manually analyze system logs. To address these challenges, in this paper, we propose a new approach for contextual analysis of system logs for understanding a system's behaviors. In particular, we first use execution patterns to represent execution structures reflected by a sequence of system logs, and propose an algorithm to mine execution patterns from the program logs. The mined execution patterns correspond to different execution paths of the system. Based on these execution patterns, our approach further learns essential contextual factors (e.g., the occurrences of specific program logs with specific parameter values) that cause a specific branch or path to be executed by the system. The mining and learning results can help system operators to understand a software system's runtime execution logic and behaviors during various tasks such as system problem diagnosis. We demonstrate the feasibility of our approach upon two real-world software systems (Hadoop and Ethereal).
Qiang Fu 0015, Jian-Guang Lou, Qingwei Lin, Rui Ding 0001, Dongmei Zhang 0001, Tao Xie 0001
MSR2
2012 Healing online service systems via mining historical issue repositories
abstract
Online service systems have been increasingly popular and important nowadays, with an increasing demand on the availability of services provided by these systems, while significant efforts have been made to strive for keeping services up continuously. Therefore, reducing the MTTR (Mean Time to Restore) of a service remains the most important step to assure the user-perceived availability of the service. To reduce the MTTR, a common practice is to restore the service by identifying and applying an appropriate healing action (i.e., a temporary workaround action such as rebooting a SQL machine). However, manually identifying an appropriate healing action for a given new issue (such as service down) is typically time consuming and error prone. To address this challenge, in this paper, we present an automated mining-based approach for suggesting an appropriate healing action for a given new issue. Our approach generates signatures of an issue from its corresponding transaction logs and then retrieves historical issues from a historical issue repository. Finally, our approach suggests an appropriate healing action by adapting healing actions for the retrieved historical issues. We have implemented a healing suggestion system for our approach and applied it to a real-world product online service that serves millions of online customers globally. The studies on 77 incidents (severe issues) over 3 months showed that our approach can effectively provide appropriate healing actions to reduce the MTTR of the service.
Rui Ding 0001, Qiang Fu 0015, Jian-Guang Lou, Qingwei Lin, Dongmei Zhang 0001, Tao Xie 0001
ASE3
2012 Performance Issue Diagnosis for Online Service Systems
abstract
Monitoring and diagnosing performance issues of an online service system are critical to assure satisfactory performance of the system. Given a detected performance issue and collected system metrics for an online service system, engineers usually need to make great efforts to conduct diagnosis by first identifying performance issue beacons, which are metrics that pinpoint to the root causes. In order to reduce the manual efforts, in this paper, we propose a new approach to effectively detecting performance issue beacons to help with performance issue diagnosis. Our approach includes techniques for mining system metric data to address limitations when applying previous classification-based approaches. Our evaluations on both a controlled environment and a real production environment show that our approach can more effectively identify performance issue beacons from system metric data than previous approaches.
Qiang Fu 0015, Jian-Guang Lou, Qingwei Lin, Rui Ding 0001, Dongmei Zhang 0001, Tao Xie 0001
SRDS2
2010 Mining program workflow from interleaved traces
abstract
Successful software maintenance is becoming increasingly critical due to the increasing dependence of our society and economy on software systems. One key problem of software maintenance is the difficulty in understanding the evolving software systems. Program workflows can help system operators and administrators to understand system behaviors and verify system executions so as to greatly facilitate system maintenance. In this paper, we propose an algorithm to automatically discover program workflows from event traces that record system events during system execution. Different from existing workflow mining algorithms, our approach can construct concurrent workflows from traces of interleaved events. Our workflow mining approach is a three-step coarse-to-fine algorithm. At first, we mine temporal dependencies for each pair of events. Then, based on the mined pair-wise tem-poral dependencies, we construct a basic workflow model by a breadth-first path pruning algorithm. After that, we refine the workflow by verifying it with all training event traces. The re-finement algorithm tries to find out a workflow that can interpret all event traces with minimal state transitions and threads. The results of both simulation data and real program data show that our algorithm is highly effective.
Jian-Guang Lou, Qiang Fu 0015, Shengqi Yang, Jiang Li 0008, Bin Wu 0001
KDD1
2010 Mining Invariants from Console Logs for System Problem Detection
Jian-Guang Lou, Qiang Fu 0015, Shengqi Yang, Jiang Li 0008
USENIX ATC1
2009 Execution Anomaly Detection in Distributed Systems through Unstructured Log Analysis
abstract
Detection of execution anomalies is very important for the maintenance, development, and performance refinement of large scale distributed systems. Execution anomalies include both work flow errors and low performance problems. People often use system logs produced by distributed systems for troubleshooting and problem diagnosis. However, manually inspecting system logs to detect anomalies is unfeasible due to the increasing scale and complexity of distributed systems. Therefore, there is a great demand for automatic anomalies detection techniques based on log analysis. In this paper, we propose an unstructured log analysis technique for anomalies detection. In the technique, we propose a novel algorithm to convert free form text messages in log files to log keys without heavily relying on application specific knowledge. The log keys correspond to the log-print statements in the source code which can provide cues of system execution behavior. After converting log messages to log keys, we learn a Finite State Automaton (FSA) from training log sequences to present the normal work flow for each system component. At the same time, a performance measurement model is learned to characterize the normal execution performance based on the log messages' timing information. With these learned models, we can automatically detect anomalies in newly input log files. Experiments on Hadoop and SILK show that the technique can effectively detect running anomalies.
Qiang Fu 0015, Jian-Guang Lou, Yi Wang 0010, Jiang Li 0008
ICDM2
2007 Distributed Density Estimation Using Non-parametric Statistics
abstract
Learning the underlying model from distributed data is often useful for many distributed systems. In this paper, we study the problem of learning a non-parametric model from distributed observations. We propose a gossip-based distributed kernel density estimation algorithm and analyze the convergence and consistency of the estimation process. Furthermore, we extend our algorithm to distributed systems under communication and storage constraints by introducing a fast and efficient data reduction algorithm. Experiments show that our algorithm can estimate underlying density distribution accurately and robustly with only small communication and storage overhead.
Yusuo Hu, Jian-Guang Lou, Jiang Li 0008
ICDCS3
2006 An Effective Epipolar Geometry Assisted Motion Estimation Technique for Multi-View Image and Video Coding
abstract
To efficiently encode data-intensive multi-view imaging content, conventional hybrid predictive coding methodologies choose to address the compression by exploiting temporal and inter-viewpoint redundancy. However, their key yet time-consuming component, motion estimation (ME), is usually not efficient in inter-viewpoint prediction because inter-viewpoint motion is quite different from temporal motion. In essence, inter-viewpoint correlation is subject to epipolar geometry, which provides constraints for multi-view image sequences. A fast inter-viewpoint ME technique is hence proposed in this paper to accelerate the encoding by employing epipolar geometry. Theoretical analysis and experimental results prove that the proposed ME algorithm can greatly reduce search region and effectively track large and irregular motion that is typical for convergent multi-view camera setups. As a result, compared with fast full search at large search size adopted in H.264, our proposed ME algorithm can obtain a similar coding efficiency while achieving a speedup ratio of 2.9.
Jiangbo Lu, Hua Cai, Jian-Guang Lou, Jiang Li 0008
ICIP3
2005 A real-time interactive multi-view video system
abstract
With the rapid development of electronic and computing technology, multi-view video is attracting extensive interest recently due to its greatly enhanced viewing experience. In this paper, we present the system architecture for real-time capturing, processing, and interactive delivery of multi-view video. Unlike previous systems that mainly focus on multi-view video capturing, our system is designed to provide multi-view video service with high degree of interactivity in real time, which is still challenging in the current state of the technology. The proposed architecture tackles many practical problems in system calibration, object tracking, video compression, interactive delivery, etc. With the proposed system, users can interactively select their desired viewing directions and enjoy many exciting visual experiences, such as view switching, frozen moment and view sweeping, in real-time and with great freedom.
Jian-Guang Lou, Hua Cai, Jiang Li 0008
ACM Multimedia1