Zhuang Li 0001

dblp:88/7870-1 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-9808-9992ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Landsat30-AU: A Vision-Language Dataset for Australian Landsat Imagery
abstract
Vision language models (VLMs) that enable natural language interaction with satellite imagery can democratize Earth observation by accelerating expert workflows, making data accessible to non-specialists, and enabling planet-scale automation. However, existing datasets focus mainly on short-term, high-resolution imagery from a limited number of satellites, overlooking low-resolution, multi-satellite, long-term archives, such as Landsat, that are essential for affordable and bias-robust global monitoring. We address this gap with Landsat30-AU, a large-scale vision-language dataset built from 30-meter resolution imagery collected by four Landsat satellites (5, 7, 8, and 9) over Australia, spanning more than 36 years. The dataset includes two components: Landsat30-AU-Cap, containing 196,262 image-caption pairs, and Landsat30-AU-VQA, comprising 17,725 human-verified visual question answering (VQA) samples across eight remote sensing domains. Both datasets are curated through a bootstrapped pipeline that leverages generic VLMs with iterative refinement and human verification to ensure quality. Our evaluation of eight VLMs on our benchmark reveals that off-the-shelf models struggle to understand satellite imagery. The open-source remote-sensing VLM EarthDial achieves only 0.07 SPIDEr in captioning and a VQA accuracy of 0.48, highlighting the limitations of current approaches. Encouragingly, lightweight fine-tuning of Qwen2.5-VL-7B on Landsat30-AU improves captioning performance from 0.11 to 0.31 SPIDEr and boosts VQA accuracy from 0.74 to 0.87.
Zhuang Li 0001, John A. Taylor
AAAI2
2026 ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations
abstract
Online conversations have become more prevalent on public discussion platforms (e.g. Reddit). With growing controversial topics, it is desirable to summarize not only diverse arguments, but also their rationale and justification. Early studies on text summarization focus on capturing general salient information in source documents, overlooking the argumentative nature of online conversations. Recent research on conversation summarization although considers the argumentative relationship among sentences, fail to explicate deeper argument structure within sentences for summarization. In this paper, we propose a novel task of argument-aware quantitative summarization to reveal the claim-reason structure of arguments in conversations, with quantities measuring argument strength. We further propose ARQUSUMM, a novel framework to address the task. To reveal the underlying argument structure within sentences, ARQUSUMM leverages LLM few-shot learning grounded in the argumentation theory to identify propositions within sentences and their claim-reason relationships. For quantitative summarization, ARQUSUMM employs argument structure-aware clustering algorithms to aggregate arguments and quantify their support. Experiments show that ARQUSUMM outperforms existing conversation and quantitative summarization models and generate summaries representing argument structures that are more helpful to users, of high textual quality and quantification accuracy.
An Quang Tang, Xiuzhen Zhang 0001, Minh Ngoc Dinh, Zhuang Li 0001
AAAI4
2026 Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering
abstract
Li Zheng, Xin Zhang, Shuyi He, Fei Li, Chong Teng, Jiang-Ming Yang, Donghong Ji, Zhuang Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shuyi He, Fei Li 0021, Chong Teng, Jiang-Ming Yang, Donghong Ji, Zhuang Li 0001
ACL (1)8
2025 SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
abstract
Zhuang Li, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhuang Li 0001, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari
ACL (1)1
2025 On the Reliability of Large Language Models for Causal Discovery
abstract
This study investigates the efficacy of Large Language Models (LLMs) in causal discovery. Using newly available open-source LLMs, OLMo and BLOOM, which provide access to their pre-training corpora, we investigate how LLMs address causal discovery through three research questions. We examine: (i) the impact of memorization for accurate causal relation prediction, (ii) the influence of incorrect causal relations in pre-training data, and (iii) the contextual nuances that influence LLMs’ understanding of causal relations. Our findings indicate that while LLMs are effective in recognizing causal relations that occur frequently in pre-training data, their ability to generalize to new or rare causal relations is limited. Moreover, the presence of incorrect causal relations significantly undermines the confidence of LLMs in corresponding correct causal relations, and the contextual information critically affects the outcomes of LLMs to discern causal connections between random variables.
Tao Feng 0013, Lizhen Qu, Niket Tandon, Zhuang Li 0001, Xiaoxi Kang, Gholamreza Haffari
ACL (1)4
2025 SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social Media
abstract
Opinion survey research is a crucial method used by social scientists for understanding societal beliefs and behaviors.Traditional methodologies often entail high costs and limited scalability, while current automated methods such as opinion synthesis exhibit severe biases and lack traceability.In this paper, we introduce SUR-VEYPILOT, a novel finite-state orchestrated agentic framework that automates the collection and analysis of human opinions from social media platforms.SURVEYPILOT addresses the limitations of pioneering approaches by (i) providing transparency and traceability in each state of opinion collection and (ii) incorporating several techniques for mitigating biases, notably with a novel genetic algorithm for improving result diversity.Our extensive experiments reveal that SURVEYPILOT achieves a close alignment with authentic survey results across multiple domains, observing average relative improvements of 68.98% and 51.37% when comparing to opinion synthesis and agent-based approaches.Implementation of SURVEYPILOT is available on https: //github.com/thanhpv2102/SurveyPilot
Viet Thanh Pham, Lizhen Qu, Zhuang Li 0001, Suraj Sharma, Gholamreza Haffari
ACL (1)3
2025 LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer Reviews
abstract
Peer review is a cornerstone of quality control in scientific publishing.With the increasing workload, the unintended use of 'quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality.Automated methods to detect such heuristics can help improve the peer-reviewing process.However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools.This work introduces LAZYREVIEW, a dataset of peer-review sentences annotated with finegrained lazy thinking categories.Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zeroshot setting.However, instruction-based finetuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data.Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback.We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community.1 Heuristics Description Example review segmentsThe results are not surprising Many findings seem obvious in retrospect, but this does not mean that the community is already aware of them and can use them as building blocks for future work.
Sukannya Purkayastha, Zhuang Li 0001, Anne Lauscher, Lizhen Qu, Iryna Gurevych
ACL (1)2
2025 QQSUM: A Novel Task and Model of Quantitative Query-Focused Summarization for Review-based Product Question Answering
abstract
Review-based Product Question Answering (PQA) allows e-commerce platforms to automatically address customer queries by leveraging insights from user reviews.However, existing PQA systems generate answers with only a single perspective, failing to capture the diversity of customer opinions.In this paper we introduce a novel task Quantitative Query-Focused Summarization (QQSUM), which aims to summarize diverse customer opinions into representative Key Points (KPs) and quantify their prevalence to effectively answer user queries.While Retrieval-Augmented Generation (RAG) shows promise for PQA, its generated answers still fall short of capturing the full diversity of viewpoints.To tackle this challenge, our model QQSUM-RAG, which extends RAG, employs few-shot learning to jointly train a KP-oriented retriever and a KP summary generator, enabling KP-based summaries that capture diverse and representative opinions.Experimental results demonstrate that QQSUM-RAG achieves superior performance compared to state-of-the-art RAG baselines in both textual quality and quantification accuracy of opinions.Our source code is available at: https://github.com/ antangrocket1312/
An Quang Tang, Xiuzhen Zhang 0001, Minh Ngoc Dinh, Zhuang Li 0001
ACL (1)4
2025 TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis
abstract
Xiaorui Wu, Xiaofeng Mao, Fei Li, Xin Zhang, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xiaorui Wu, Xiaofeng Mao, Fei Li 0021, Xuanhong Li, Chong Teng, Donghong Ji, Zhuang Li 0001
ACL (1)8
2025 DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement
abstract
Vision-Language Models (VLMs) generate discourse-level, multi-sentence visual descriptions, challenging text scene graph parsers built for single-sentence caption-to-graph mapping.Current approaches typically merge sentencelevel parsing outputs for discourse input, often missing phenomena like cross-sentence coreference, resulting in fragmented graphs and degraded downstream VLM task performance.We introduce a new task, Discourse-level text Scene Graph parsing (DiscoSG), and release DiscoSG-DS, a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence captiongraph pairs.Each caption averages 9 sentences, and each graph contains at least 3× more triples than those in existing datasets.Fine-tuning GPT-4o on DiscoSG-DS yields over 40% higher SPICE metric than the best sentence-merging baseline.However, its high inference cost and licensing restrict opensource use.Smaller fine-tuned open-source models (e.g., Flan-T5) perform well on simpler graphs yet degrade on denser, more complex graphs.To bridge this gap, we introduce DiscoSG-Refiner, a lightweight open-source parser that drafts a seed graph and iteratively refines it with a novel learned graph-editing model, achieving 30% higher SPICE than the baseline while delivering 86× faster inference than GPT-4o.It generalises from simple to dense graphs, thereby consistently improving downstream VLM tasks, including discourselevel caption evaluation and hallucination detection, outperforming alternative open-source parsers.
Shaoqing Lin, Chong Teng, Fei Li 0021, Donghong Ji, Lizhen Qu, Zhuang Li 0001
EMNLP6
2025 CultureInstruct: Curating Multi-Cultural Instructions at Scale
abstract
Viet Thanh Pham, Zhuang Li, Lizhen Qu, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Viet Thanh Pham, Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
NAACL (Long Papers)2
2025 EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions
abstract
Large language models (LLMs) frequently refuse to respond to pseudo-malicious instructions: semantically harmless input queries triggering unnecessary LLM refusals due to conservative safety alignment, significantly impairing user experience. Collecting such instructions is crucial for evaluating and mitigating over-refusals, but existing instruction curation methods, like manual creation or instruction rewriting, either lack scalability or fail to produce sufficiently diverse and effective refusal-inducing prompts. To address these limitations, we introduce EVOREFUSE, a prompt optimization approach that generates diverse pseudo-malicious instructions consistently eliciting confident refusals across LLMs. EVOREFUSE employs an evolutionary algorithm exploring the instruction space in more diverse directions than existing methods via mutation strategies and recombination, and iteratively evolves seed instructions to maximize evidence lower bound on LLM refusal probability. Using EVOREFUSE, we create two novel datasets: EVOREFUSE-TEST, a benchmark of 582 pseudo-malicious instructions that outperforms the next-best benchmark with 85.34% higher average refusal triggering rate across 9 LLMs without a safety-prior system prompt, 34.86% greater lexical diversity, and 40.03% improved LLM response confidence scores; and EVOREFUSE-ALIGN, which provides 3,000 pseudo-malicious instructions with responses for supervised and preference-based alignment training. With supervised fine-tuning on EVOREFUSE-ALIGN, LLAMA3.1-8B-INSTRUCT achieves up to 29.85% fewer over-refusals than models trained on the second-best alignment dataset, without compromising safety. Our analysis with EVOREFUSE-TEST reveals models trigger over-refusals by overly focusing on sensitive keywords while ignoring broader context. Our code and datasets are available at https://github.com/FishT0ucher/EVOREFUSE .
Xiaorui Wu, Fei Li 0021, Xiaofeng Mao, Chong Teng, Donghong Ji, Zhuang Li 0001
NeurIPS9
2025 Scalable Frame-Based Construction of Sociocultural Norm Bases for Socially Aware Dialogues
abstract
Sociocultural norms serve as guiding principles for personal conduct in social interactions, emphasizing respect, cooperation, and appropriate behavior, which is able to benefit tasks including conversational information retrieval, contextual information retrieval, and retrieval-enhanced machine learning. We propose a scalable approach for constructing a Sociocultural Norm ( Scn ) Base using large language models (LLMs) for socially aware dialogues. We construct a comprehensive and publicly accessible Chinese Sociocultural NormBase ( ChineseNormBase ). Our approach utilizes socially aware dialogues, enriched with contextual frames, as the primary data source to constrain the generating process and reduce the hallucinations. This enables extracting of high-quality and nuanced natural-language norm statements, leveraging the pragmatic implications of utterances with respect to the situation. As real dialogue annotated with gold frames are not readily available, we propose using synthetic data. Our empirical results show (i) the quality of the Scn s derived from synthetic data is comparable to that from real dialogues annotated with gold frames, and (ii) the quality of the Scn s extracted from real data, annotated with either silver (predicted) or gold frames, surpasses that without the frame annotations. We further show the effectiveness of the extracted Scn s in a Retrieval-Augmented Generation (RAG)-based model to reason about multiple downstream dialogue tasks.
Shilin Qu, Weiqing Wang 0001, Xin Zhou 0023, Haolan Zhan, Zhuang Li 0001, Lizhen Qu, Linhao Luo, Yuan-Fang Li, Gholamreza Haffari
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Compositional Generalization for Multi-Label Text Classification: A Data-Augmentation Approach
abstract
Despite significant advancements in multi-label text classification, the ability of existing models to generalize to novel and seldom-encountered complex concepts, which are compositions of elementary ones, remains underexplored. This research addresses this gap. By creating unique data splits across three benchmarks, we assess the compositional generalization ability of existing multi-label text classification models. Our results show that these models often fail to generalize to compositional concepts encountered infrequently during training, leading to inferior performance on tests with these new combinations. To address this, we introduce a data augmentation method that leverages two innovative text generation models designed to enhance the classification models' capacity for compositional generalization. Our experiments show that this data augmentation approach significantly improves the compositional generalization capabilities of classification models on our benchmarks, with both generation models surpassing other text generation baselines. Our codes available at https://github.com/yychai74/LD-VAE.
Yuyang Chai, Zhuang Li 0001, Fei Li 0021, Donghong Ji, Chong Teng
AAAI2
2024 IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained Models
abstract
Machine learning models have made incredible progress, but they still struggle when applied to examples from unseen domains.This study focuses on a specific problem of domain generalization, where a model is trained on one source domain and tested on multiple target domains that are unseen during training.We propose IMO: Invariant features Masks for Out-of-Distribution text classification, to achieve OOD generalization by learning domain-invariant features.During training, IMO employs a greedy algorithm to learn sparse representations for each layer in a top-down manner.It performs better than the opposite direction and learning of sparse representations for all layers simultaneously.Our comprehensive experiments show that IMO substantially outperforms strong baselines such as prompt-based methods and large language models, in terms of various evaluation metrics and settings.1
Tao Feng 0013, Lizhen Qu, Zhuang Li 0001, Haolan Zhan, Yuncheng Hua, Gholamreza Haffari
ACL (1)3
2024 Detecting AI-Generated Sentences in Human-AI Collaborative Hybrid Texts: Challenges, Strategies, and Insights
Zijie Zeng, Shiqi Liu 0003, Lele Sha, Zhuang Li 0001, Kaixun Yang, Sannyuya Liu, Dragan Gasevic, Guangliang Chen
IJCAI4
2024 Generative Dialogue Sentiment and Act Recognition with Feature Denoising and Set Prediction
Bobo Li 0001, Zhuang Li 0001, Yuyang Chai, Fei Li 0021, Chong Teng, Donghong Ji
NLPCC (5)3
2024 GPTVoiceTasker: Advancing Multi-step Mobile Task Efficiency Through Dynamic Interface Exploration and Learning
abstract
Virtual assistants have the potential to play an important role in helping users achieves different tasks. However, these systems face challenges in their real-world usability, characterized by inefficiency and struggles in grasping user intentions. Leveraging recent advances in Large Language Models (LLMs), we introduce GptVoiceTasker, a virtual assistant poised to enhance user experiences and task efficiency on mobile devices. GptVoiceTasker excels at intelligently deciphering user commands and executing relevant device interactions to streamline task completion. For unprecedented tasks, GptVoiceTasker utilises the contextual information and on-screen content to continuously explore and execute the tasks. In addition, the system continually learns from historical user commands to automate subsequent task invocations, further enhancing execution efficiency. From our experiments, GptVoiceTasker achieved 84.5% accuracy in parsing human commands into executable actions and 85.7% accuracy in automating multi-step tasks. In our user study, GptVoiceTasker boosted task efficiency in real-world scenarios by 34.85%, accompanied by positive participant feedback. We made GptVoiceTasker open-source, inviting further research into LLMs utilization for diverse tasks through prompt engineering and leveraging user usage data to improve efficiency.
Minh Duc Vu, Han Wang 0023, Jieshan Chen, Zhuang Li 0001, Shengdong Zhao 0001, Zhenchang Xing, Chunyang Chen 0001
UIST4
2023 The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active Learning
abstract
Multilingual semantic parsing aims to leverage the knowledge from the high-resource languages to improve low-resource semantic parsing, yet commonly suffers from the data imbalance problem.Prior works propose to utilize the translations by either humans or machines to alleviate such issues.However, human translations are expensive, while machine translations are cheap but prone to error and bias.In this work, we propose an active learning approach that exploits the strengths of both human and machine translations by iteratively adding small batches of human translations into the machine-translated training set.Besides, we propose novel aggregated acquisition criteria that help our active learning method select utterances to be manually translated.Our experiments demonstrate that an ideal utterance selection can significantly reduce the error and bias in the translated data, resulting in higher parser accuracies than the parsers merely trained on the machine-translated data.
Zhuang Li 0001, Lizhen Qu, Phil Cohen 0001, Raj Tumuluri, Gholamreza Haffari
ACL (1)1
2023 On Robustness of Prompt-based Semantic Parsing with Large Pre-trained Language Model: An Empirical Study on Codex
abstract
Terry Yue Zhuo, Zhuang Li, Yujin Huang, Fatemeh Shiri, Weiqing Wang, Gholamreza Haffari, Yuan-Fang Li. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Terry Yue Zhuo, Zhuang Li 0001, Yujin Huang, Fatemeh Shiri, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li
EACL2
2023 Reranking for Natural Language Generation from Logical Forms: A Study based on Large Language Models
abstract
Levon Haroutunian, Zhuang Li, Lucian Galescu, Philip Cohen, Raj Tumuluri, Gholamreza Haffari. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Levon Haroutunian, Zhuang Li 0001, Lucian Galescu, Phil Cohen 0001, Raj Tumuluri, Gholamreza Haffari
IJCNLP (1)2
2023 SocialDial: A Benchmark for Socially-Aware Dialogue Systems
abstract
Content Warning: this paper may contain content that is offensive or upsetting.
Haolan Zhan, Zhuang Li 0001, Yufei Wang 0003, Linhao Luo, Tao Feng 0013, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, Gholamreza Haffari
SIGIR2
2022 Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language Generation
abstract
In this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of taskspecific labeled examples.In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space.Both types of priors satisfy a novel property called ϵ-disentangled.We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work.The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings.Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.The code is available at https://github. com/zhuang-li/VAE-DPrior.
Zhuang Li 0001, Lizhen Qu, Qiongkai Xu, Tongtong Wu, Tianyang Zhan, Gholamreza Haffari
EMNLP1
2022 Paraphrasing Techniques for Maritime QA system
Fatemeh Shiri, Terry Yue Zhuo, Zhuang Li 0001, Shirui Pan, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002
FUSION3
2022 Pretrained Language Model in Continual Learning: A Comparative Study
Tongtong Wu, Massimo Caccia, Zhuang Li 0001, Yuan-Fang Li, Guilin Qi, Gholamreza Haffari
ICLR3
2021 On Robustness of Neural Semantic Parsers
abstract
Semantic parsing maps natural language (NL) utterances into logical forms (LFs), which underpins many advanced NLP problems.Semantic parsers gain performance boosts with deep neural networks, but inherit vulnerabilities against adversarial examples.In this paper, we provide the empirical study on the robustness of semantic parsers in the presence of adversarial attacks.Formally, adversaries of semantic parsing are considered to be the perturbed utterance-LF pairs, whose utterances have exactly the same meanings as the original ones.A scalable methodology is proposed to construct robustness test sets based on existing benchmark corpora.Our results answered five research questions in measuring the sateof-the-art parsers' performance on robustness test sets, and evaluating the effect of data augmentation.
Zhuang Li 0001, Lizhen Qu
EACL2
2021 Few-Shot Semantic Parsing for New Predicates
abstract
In this work, we investigate the problems of semantic parsing in a few-shot learning setting.In this setting, we are provided with k utterance-logical form pairs per new predicate.The state-of-the-art neural semantic parsers achieve less than 25% accuracy on benchmark datasets when k = 1.To tackle this problem, we proposed to i) apply a designated metalearning method to train the model; ii) regularize attention scores with alignment statistics; iii) apply a smoothing technique in pretraining.As a result, our method consistently outperforms all the baselines in both one and two-shot settings.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
EACL1
2021 Total Recall: a Customized Continual Learning Method for Neural Semantic Parsers
abstract
This paper investigates continual learning for semantic parsing.In this setting, a neural semantic parser learns tasks sequentially without accessing full training data from previous tasks.Direct application of the SOTA continual learning algorithms to this problem fails to achieve comparable performance with retraining models with all seen tasks, because they have not considered the special properties of structured outputs, yielded by semantic parsers.Therefore, we propose TO-TAL RECALL, a continual learning method designed for neural semantic parsers from two aspects: i) a sampling method for memory replay that diversifies logical form templates and balances distributions of parse actions in a memory; ii) a two-stage training method that significantly improves generalization capability of the parsers across tasks.We conduct extensive experiments to study the research problems involved in continual semantic parsing, and demonstrate that a neural semantic parser trained with TOTAL RECALL achieves superior performance than the one trained directly with the SOTA continual learning algorithms, and achieve a 3-6 times speedup compared to retraining from scratch.Code and datasets are available at: https://github. com/zhuang-li/cl_nsp.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
EMNLP (1)1
2020 Context Dependent Semantic Parsing: A Survey
abstract
Semantic parsing is the task of translating natural language utterances into machine-readable meaning representations.Currently, most semantic parsing methods are not able to utilize contextual information (e.g.dialogue and comments history), which has a great potential to boost semantic parsing performance.To address this issue, context dependent semantic parsing has recently drawn a lot of attention.In this survey, we investigate progress on the methods for the context dependent semantic parsing, together with the current datasets and tasks.We then point out open problems and challenges for future research in this area.The collected resources for this topic are available at: https://github.
Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari
COLING1