EDBT 2026 Demo / reviewers in the wild / expert
Lizhen Qu
dblp:58/3601
· DBLP profile ↗
67ranked-venue papers
6as first author
45since 2021 · last 2026
0000-0002-7764-431XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 57 · 6 first-author · 37 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LiveCultureBench: a Multi-Agent, Multi-Cultural Benchmark for Large Language Models in Dynamic Social SimulationsabstractLarge language models (LLMs) are increasingly deployed as autonomous agents, yet evaluations focus primarily on task success rather than cultural appropriateness or evaluator reliability.We introduce LIVECULTUREBENCH 1 , a multi-cultural, dynamic benchmark that embeds LLMs as agents in a simulated town and evaluates them on both task completion and adherence to socio-cultural norms.The simulation models a small city as a location graph with synthetic residents having diverse demographic and cultural profiles.Each episode assigns one resident a daily goal while others provide social context.An LLM-based verifier generates structured judgments on norm violations and task progress, which we aggregate into metrics capturing task-norm trade-offs and verifier uncertainty.Using LIVECULTUREBENCH across models and cultural profiles, we study (i) cross-cultural robustness of LLM agents, (ii) how they balance effectiveness against norm sensitivity, and (iii) when LLM-as-a-judge evaluation is reliable for automated benchmarking versus when human oversight is needed. Viet Thanh Pham, Lizhen Qu, Thuy-Trang Vu, Gholamreza Haffari, Dinh Q. Phung |
ACL (1) | 2 |
| 2026 | LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal IssuesabstractFanyu Wang, Xiaoxi Kang, Paul Burgess, Aashish Srivastava, Chetan Arora, Adnan Trakic, Lay-Ki Soon, Md Khalid Hossain, Lizhen Qu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Fanyu Wang, Xiaoxi Kang, Paul Burgess, Aashish Srivastava, Chetan Arora 0002, Adnan Trakic, Lay-Ki Soon, Md Khalid Hossain, Lizhen Qu |
ACL (1) | 9 |
| 2026 | Supporting Multimodal Data Interaction on Refreshable Tactile Displays: An Architecture to Combine Touch and Conversational AIabstractCombining conversational AI with refreshable tactile displays (RTDs) offers significant potential for creating accessible data visualization for people who are blind or have low vision (BLV). To support researchers and developers building accessible data visualizations with RTDs, we present a multimodal data interaction architecture along with an open-source reference implementation. Our system is the first to combine touch input with a conversational agent on an RTD, enabling deictic queries that fuse touch context with spoken language, such as "what is the trend between these points?" The architecture addresses key technical challenges, including touch sensing on RTDs, visual-to-tactile encoding, integrating touch context with conversational AI, and synchronizing multimodal output. Our contributions are twofold: (1) a technical architecture integrating RTD hardware, external touch sensing, and conversational AI to enable multimodal data interaction; and (2) an open-source reference implementation demonstrating its feasibility. This work provides a technical foundation to support future research in multimodal accessible data visualization. Samuel Reinders, Munazza Zaib, Matthew Butler 0002, Bongshin Lee, Ingrid Zukerman, Lizhen Qu, Kim Marriott |
PacificVis | 6 |
| 2025 | SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language ModelsabstractZhuang Li, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhuang Li 0001, Yuncheng Hua, Thuy-Trang Vu, Haolan Zhan, Lizhen Qu, Gholamreza Haffari |
ACL (1) | 5 |
| 2025 | On the Reliability of Large Language Models for Causal DiscoveryabstractThis study investigates the efficacy of Large Language Models (LLMs) in causal discovery. Using newly available open-source LLMs, OLMo and BLOOM, which provide access to their pre-training corpora, we investigate how LLMs address causal discovery through three research questions. We examine: (i) the impact of memorization for accurate causal relation prediction, (ii) the influence of incorrect causal relations in pre-training data, and (iii) the contextual nuances that influence LLMs’ understanding of causal relations. Our findings indicate that while LLMs are effective in recognizing causal relations that occur frequently in pre-training data, their ability to generalize to new or rare causal relations is limited. Moreover, the presence of incorrect causal relations significantly undermines the confidence of LLMs in corresponding correct causal relations, and the contextual information critically affects the outcomes of LLMs to discern causal connections between random variables. Tao Feng 0013, Lizhen Qu, Niket Tandon, Zhuang Li 0001, Xiaoxi Kang, Gholamreza Haffari |
ACL (1) | 2 |
| 2025 | IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular DataabstractCausal discovery is fundamental to scientific research, yet traditional statistical algorithms face significant challenges, including expensive data collection, redundant computation for known relations, and unrealistic assumptions.While recent LLM-based methods excel at identifying commonly known causal relations, they fail to uncover novel relations.We introduce IRIS (Iterative Retrieval and Integrated System for Real-Time Causal Discovery), a novel framework that addresses these limitations.Starting with a set of initial variables, IRIS automatically collects relevant documents, extracts variables, and uncovers causal relations.Our hybrid causal discovery method combines statistical algorithms and LLM-based methods to discover known and novel causal relations.In addition to causal discovery on initial variables, the missing variable proposal component of IRIS identifies and incorporates missing variables to expand the causal graphs.Our approach enables real-time causal discovery from only a set of initial variables without requiring pre-existing datasets. Tao Feng 0013, Lizhen Qu, Niket Tandon, Gholamreza Haffari |
ACL (1) | 2 |
| 2025 | SurveyPilot: an Agentic Framework for Automated Human Opinion Collection from Social MediaabstractOpinion survey research is a crucial method used by social scientists for understanding societal beliefs and behaviors.Traditional methodologies often entail high costs and limited scalability, while current automated methods such as opinion synthesis exhibit severe biases and lack traceability.In this paper, we introduce SUR-VEYPILOT, a novel finite-state orchestrated agentic framework that automates the collection and analysis of human opinions from social media platforms.SURVEYPILOT addresses the limitations of pioneering approaches by (i) providing transparency and traceability in each state of opinion collection and (ii) incorporating several techniques for mitigating biases, notably with a novel genetic algorithm for improving result diversity.Our extensive experiments reveal that SURVEYPILOT achieves a close alignment with authentic survey results across multiple domains, observing average relative improvements of 68.98% and 51.37% when comparing to opinion synthesis and agent-based approaches.Implementation of SURVEYPILOT is available on https: //github.com/thanhpv2102/SurveyPilot Viet Thanh Pham, Lizhen Qu, Zhuang Li 0001, Suraj Sharma, Gholamreza Haffari |
ACL (1) | 2 |
| 2025 | LazyReview: A Dataset for Uncovering Lazy Thinking in NLP Peer ReviewsabstractPeer review is a cornerstone of quality control in scientific publishing.With the increasing workload, the unintended use of 'quick' heuristics, referred to as lazy thinking, has emerged as a recurring issue compromising review quality.Automated methods to detect such heuristics can help improve the peer-reviewing process.However, there is limited NLP research on this issue, and no real-world dataset exists to support the development of detection tools.This work introduces LAZYREVIEW, a dataset of peer-review sentences annotated with finegrained lazy thinking categories.Our analysis reveals that Large Language Models (LLMs) struggle to detect these instances in a zeroshot setting.However, instruction-based finetuning on our dataset significantly boosts performance by 10-20 performance points, highlighting the importance of high-quality training data.Furthermore, a controlled experiment demonstrates that reviews revised with lazy thinking feedback are more comprehensive and actionable than those written without such feedback.We will release our dataset and the enhanced guidelines that can be used to train junior reviewers in the community.1 Heuristics Description Example review segmentsThe results are not surprising Many findings seem obvious in retrospect, but this does not mean that the community is already aware of them and can use them as building blocks for future work. Sukannya Purkayastha, Zhuang Li 0001, Anne Lauscher, Lizhen Qu, Iryna Gurevych |
ACL (1) | 4 |
| 2025 | CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue SystemsabstractAutomatically evaluating the quality of responses in dialogue systems is a challenging yet crucial task. Current metrics often fail to align with human judgments, especially when assessing responses that are grammatically correct. To address this issue, we propose a novel metric, called CausalScore, which assesses the relevance of responses by measuring the causal strength between dialogue histories and responses. The causal strength is estimated by utilizing both unconditional dependence and conditional dependencies from dialogue histories to responses. We compare our metric with the existing competitive metrics in terms of their alignment with human judgements. Our experimental results demonstrate that CausalScore significantly surpasses existing state-of-the-art metrics by aligning better with human judgements. Additionally, we collect a dialogue dataset CGDIALOG+ with human-annotated causal relations and a set of pairwise human judgements to facilitate the development of automatic metrics. Tao Feng 0013, Lizhen Qu, Xiaoxi Kang, Gholamreza Haffari |
COLING | 2 |
| 2025 | DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph RefinementabstractVision-Language Models (VLMs) generate discourse-level, multi-sentence visual descriptions, challenging text scene graph parsers built for single-sentence caption-to-graph mapping.Current approaches typically merge sentencelevel parsing outputs for discourse input, often missing phenomena like cross-sentence coreference, resulting in fragmented graphs and degraded downstream VLM task performance.We introduce a new task, Discourse-level text Scene Graph parsing (DiscoSG), and release DiscoSG-DS, a dataset of 400 expert-annotated and 8,430 synthesised multi-sentence captiongraph pairs.Each caption averages 9 sentences, and each graph contains at least 3× more triples than those in existing datasets.Fine-tuning GPT-4o on DiscoSG-DS yields over 40% higher SPICE metric than the best sentence-merging baseline.However, its high inference cost and licensing restrict opensource use.Smaller fine-tuned open-source models (e.g., Flan-T5) perform well on simpler graphs yet degrade on denser, more complex graphs.To bridge this gap, we introduce DiscoSG-Refiner, a lightweight open-source parser that drafts a seed graph and iteratively refines it with a novel learned graph-editing model, achieving 30% higher SPICE than the baseline while delivering 86× faster inference than GPT-4o.It generalises from simple to dense graphs, thereby consistently improving downstream VLM tasks, including discourselevel caption evaluation and hallucination detection, outperforming alternative open-source parsers. Shaoqing Lin, Chong Teng, Fei Li 0021, Donghong Ji, Lizhen Qu, Zhuang Li 0001 |
EMNLP | 5 |
| 2025 | Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language ModelsabstractLarge Audio Language Models (LALMs) have extended the capabilities of Large Language Models (LLMs) by enabling audio-based human interactions.However, recent research has revealed that LALMs remain vulnerable to harmful queries due to insufficient safetyalignment.Despite advances in defence measures for text and vision LLMs, effective safetyalignment strategies and audio-safety dataset specifically targeting LALMs are notably absent.Meanwhile defence measures based on Supervised Fine-tuning (SFT) struggle to address safety improvement while avoiding overrejection issues, significantly compromising helpfulness.In this work, we propose an unsupervised safety-fine-tuning strategy as remedy that reshapes model's representation space to enhance existing LALMs safety-alignment while balancing the risk of over-rejection.Our experiments, conducted across three generations of Qwen LALMs, demonstrate that our approach significantly improves LALMs safety under three modality input conditions (audiotext, text-only, and audio-only) while increasing over-rejection rate by only 0.88% on average. 1 Warning: this paper contains harmful examples. Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari |
EMNLP | 2 |
| 2025 | The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite GraphabstractThe performance of large language models (LLMs) is strongly influenced by the quality and diversity of data used during supervised fine-tuning (SFT). However, current data selection methods often prioritize one aspect over the other, resulting in suboptimal training outcomes. To address this, we formulate data selection as a set cover problem and present GraphFilter, a novel approach that balances both quality and diversity in data selection. GraphFilter models the dataset as a bipartite graph connecting sentences to their constituent n-grams, then employs a priority function that combines quality and diversity metrics multiplicatively. GraphFilter iteratively selects sentences with the highest priority, removes covered n-grams from the bipartite graph, and recomputes priorities to reflect the changing data landscape. We validate GraphFilter using three model backbones across six widely-used benchmarks, demonstrating that it outperforms nine existing baselines in both model performance and computational efficiency. Further analysis shows that our design choices lead to more effective subset selection, underscores the value of instruction diversity, and provides insights into how quality and diversity interact with different subset sizes. Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Gholamreza Haffari |
ICML | 3 |
| 2025 | CultureInstruct: Curating Multi-Cultural Instructions at ScaleabstractViet Thanh Pham, Zhuang Li, Lizhen Qu, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Viet Thanh Pham, Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari |
NAACL (Long Papers) | 3 |
| 2025 | ACCESS : A Benchmark for Abstract Causal Event Discovery and ReasoningabstractVy Vo, Lizhen Qu, Tao Feng, Yuncheng Hua, Xiaoxi Kang, Songhai Fan, Tim Dwyer, Lay-Ki Soon, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Vy Vo, Lizhen Qu, Tao Feng 0013, Yuncheng Hua, Xiaoxi Kang, Songhai Fan, Tim Dwyer, Lay-Ki Soon, Gholamreza Haffari |
NAACL (Long Papers) | 2 |
| 2025 | Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal ModelsabstractHao Yang, Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Lizhen Qu, Ehsan Shareghi, Gholamreza Haffari |
NAACL (Long Papers) | 2 |
| 2025 | Unbiased Sliced Wasserstein Kernels for High-Quality Audio CaptioningabstractAudio captioning systems face a fundamental challenge: teacher-forcing training creates exposure bias that leads to caption degeneration during inference. While contrastive methods have been proposed as solutions, they typically fail to capture the crucial temporal relationships between acoustic and linguistic modalities. We address this limitation by introducing the unbiased sliced Wasserstein RBF (USW-RBF) kernel with rotary positional embedding, specifically designed to preserve temporal information across modalities. Our approach offers a practical advantage: the kernel enables efficient stochastic gradient optimization, making it computationally feasible for real-world applications. Building on this foundation, we develop a complete audio captioning framework that integrates stochastic decoding to further mitigate caption degeneration. Extensive experiments on AudioCaps and Clotho datasets demonstrate that our method significantly improves caption quality, lexical diversity, and text-to-audio retrieval accuracy. Furthermore, we demonstrate the generalizability of our USW-RBF kernel by applying it to audio reasoning tasks, where it enhances the reasoning capabilities of large audio language models on the CompA-R in terms of correctness and quality. Our kernel also improves the reasoning accuracy of the MMAU-test-mini benchmarks by $4\%$. These results establish our approach as a powerful and generalizable solution for cross-modal alignment challenges in audio-language tasks. Manh Luong, Dinh Q. Phung, Gholamreza Haffari, Lizhen Qu |
NeurIPS | 5 |
| 2025 | From continuous pre-training to alignment: A comprehensive toolkit for large language models in federated learning
Zhuo Zhang 0007, Lizhen Qu, Xun Zhou 0001, Wendy Hui Wang, Zenglin Xu |
Neurocomputing | 4 |
| 2025 | Scalable Frame-Based Construction of Sociocultural Norm Bases for Socially Aware DialoguesabstractSociocultural norms serve as guiding principles for personal conduct in social interactions, emphasizing respect, cooperation, and appropriate behavior, which is able to benefit tasks including conversational information retrieval, contextual information retrieval, and retrieval-enhanced machine learning. We propose a scalable approach for constructing a Sociocultural Norm ( Scn ) Base using large language models (LLMs) for socially aware dialogues. We construct a comprehensive and publicly accessible Chinese Sociocultural NormBase ( ChineseNormBase ). Our approach utilizes socially aware dialogues, enriched with contextual frames, as the primary data source to constrain the generating process and reduce the hallucinations. This enables extracting of high-quality and nuanced natural-language norm statements, leveraging the pragmatic implications of utterances with respect to the situation. As real dialogue annotated with gold frames are not readily available, we propose using synthetic data. Our empirical results show (i) the quality of the Scn s derived from synthetic data is comparable to that from real dialogues annotated with gold frames, and (ii) the quality of the Scn s extracted from real data, annotated with either silver (predicted) or gold frames, surpasses that without the frame annotations. We further show the effectiveness of the extracted Scn s in a Retrieval-Augmented Generation (RAG)-based model to reason about multiple downstream dialogue tasks. Shilin Qu, Weiqing Wang 0001, Xin Zhou 0023, Haolan Zhan, Zhuang Li 0001, Lizhen Qu, Linhao Luo, Yuan-Fang Li, Gholamreza Haffari |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | When Refreshable Tactile Displays Meet Conversational Agents: Investigating Accessible Data Presentation and Analysis with Touch and SpeechabstractDespite the recent surge of research efforts to make data visualizations accessible to people who are blind or have low vision (BLV), how to support BLV people's data analysis remains an important and challenging question. As refreshable tactile displays (RTDs) become cheaper and conversational agents continue to improve, their combination provides a promising approach to support BLV people's interactive data exploration and analysis. To understand how BLV people would use and react to a system combining an RTD with a conversational agent, we conducted a Wizard-of-Oz study with 11 BLV participants, where they interacted with line charts, bar charts, and isarithmic maps. Our analysis of participants' interactions led to the identification of nine distinct patterns. We also learned that the choice of modalities depended on the type of task and prior experience with tactile graphics, and that participants strongly preferred the combination of RTD and speech to a single modality. In addition, participants with more tactile experience described how tactile images facilitated a deeper engagement with the data and supported independent interpretation. Our findings will inform the design of interfaces for such interactive mixed-modality systems. Samuel Reinders, Matthew Butler 0002, Ingrid Zukerman, Bongshin Lee, Lizhen Qu, Kim Marriott |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | IMO: Greedy Layer-Wise Sparse Representation Learning for Out-of-Distribution Text Classification with Pre-trained ModelsabstractMachine learning models have made incredible progress, but they still struggle when applied to examples from unseen domains.This study focuses on a specific problem of domain generalization, where a model is trained on one source domain and tested on multiple target domains that are unseen during training.We propose IMO: Invariant features Masks for Out-of-Distribution text classification, to achieve OOD generalization by learning domain-invariant features.During training, IMO employs a greedy algorithm to learn sparse representations for each layer in a top-down manner.It performs better than the opposite direction and learning of sparse representations for all layers simultaneously.Our comprehensive experiments show that IMO substantially outperforms strong baselines such as prompt-based methods and large language models, in terms of various evaluation metrics and settings.1 Tao Feng 0013, Lizhen Qu, Zhuang Li 0001, Haolan Zhan, Yuncheng Hua, Gholamreza Haffari |
ACL (1) | 2 |
| 2024 | Revisiting Data Reconstruction Attacks on Real-world Dataset for Federated Natural Language UnderstandingabstractWith the growing privacy concerns surrounding natural language understanding (NLU) applications, the need to train high-quality models while safeguarding data privacy has reached unprecedented importance. Federated learning (FL) offers a promising approach to collaborative model training by exchanging model gradients. However, many studies show that eavesdroppers in FL could develop sophisticated data reconstruction attack (DRA) to accurately reconstruct clients’ data from the shared gradients. Regrettably, current DRA methods in federated NLU have been mostly conducted on public datasets, lacking a comprehensive evaluation of real-world privacy datasets. To address this limitation, this paper presents a pioneering study that reexamines the performance of these DRA methods as well as corresponding defense methods. Specifically, we introduce a novel real-world privacy dataset called FedAttack which leads to a significant discovery: existing DRA methods usually fail to accurately recover the original text of real-world privacy data. In detail, the tokens within a recovery sentence are disordered and intertwined with tokens from other sentences in the same training batch. Moreover, our experiments demonstrate that the performance of DRA is also influenced by different languages and domains. By discovering these findings, our work lays a solid foundation for further research into the development of more practical DRA methods and corresponding defenses. Zhuo Zhang 0007, Xiangjing Hu, Wendy Hui Wang, Yue Yu 0001, Qifan Wang 0001, Lizhen Qu, Zenglin Xu |
LREC/COLING | 9 |
| 2024 | Generative Region-Language Pretraining for Open-Ended Object DetectionabstractIn recent research, significant attention has been devoted to the open-vocabulary object detection task, aiming to generalize beyond the limited number of classes labeled during training and detect objects described by arbitrary category names at inference. Compared with conventional object detection, open vocabulary object detection largely extends the object detection categories. However, it relies on calculating the similarity between image regions and a set of arbitrary category names with a pretrained vision-and-language model. This implies that, despite its open-set nature, the task still needs the predefined object categories during the inference stage. This raises the question: What if we do not have exact knowledge of object categories during inference? In this paper, we call such a new setting as generative open-ended object detection, which is a more general and practical problem. To address it, we formulate object detection as a generative problem and propose a simple framework named GenerateU, which can detect dense objects and generate their names in a free-form way. Particularly, we employ Deformable DETR as a region proposal generator with a language model translating visual regions to object names. To assess the free-form object detection task, we introduce an evaluation method designed to quantitatively measure the performance of generative out-comes. Extensive experiments demonstrate strong zero-shot detection performance of our GenerateU. For example, on the LVIS dataset, our GenerateU achieves comparable results to the open-vocabulary object detection method GLIP, even though the category names are not seen by GenerateU during inference. Code is available at: https://github.com/FoundationVision/GenerateU. Chuang Lin 0003, Yi Jiang 0009, Lizhen Qu, Zehuan Yuan, Jianfei Cai 0001 |
CVPR | 3 |
| 2024 | Importance-Aware Data Augmentation for Document-Level Neural Machine TranslationabstractMinghao Wu, Yufei Wang, George Foster, Lizhen Qu, Gholamreza Haffari. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Minghao Wu, Yufei Wang 0003, George F. Foster, Lizhen Qu, Gholamreza Haffari |
EACL (1) | 4 |
| 2024 | Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language ModelsabstractLarge language models (LLMs) are typically fine-tuned on diverse and extensive datasets sourced from various origins to develop a comprehensive range of skills, such as writing, reasoning, chatting, coding, and more.Each skill has unique characteristics, and these datasets are often heterogeneous and imbalanced, making the fine-tuning process highly challenging.Balancing the development of each skill while ensuring the model maintains its overall performance requires sophisticated techniques and careful dataset curation.In this work, we propose a general, model-agnostic, reinforcement learning framework, MIXTURE-OF-SKILLS (MOS), that learns to optimize data usage automatically during the fine-tuning process.This framework ensures the optimal comprehensive skill development of LLMs by dynamically adjusting the focus on different datasets based on their current learning state.To validate the effectiveness of MOS, we conduct extensive experiments using three diverse LLM backbones on two widely used benchmarks and demonstrate that MOS substantially enhances model performance.Building on the success of MOS, we propose MOSPEC, an adaptation for task-specific fine-tuning, which harnesses the utilities of various datasets for a specific purpose.Our work underlines the significance of dataset rebalancing and present MOS as a powerful, general solution for optimizing data usage in the fine-tuning of LLMs for various purposes. Minghao Wu, Thuy-Trang Vu, Lizhen Qu, Reza Haf |
EMNLP | 3 |
| 2024 | Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and InsightsabstractLarge Multimodal Models (LMMs) have achieved great success recently, demonstrating a strong capability to understand multimodal information and to interact with human users.Despite the progress made, the challenge of detecting high-risk interactions in multimodal settings, and in particular in speech modality, remains largely unexplored.Conventional research on risk for speech modality primarily emphasises the content (e.g., what is captured as transcription).However, in speechbased interactions, paralinguistic cues in audio can significantly alter the intended meaning behind utterances.In this work, we propose a speech-specific risk taxonomy, covering 8 risk categories under hostility (malicious sarcasm and threats), malicious imitation (age, gender, ethnicity), and stereotypical biases (age, gender, ethnicity).Based on the taxonomy, we create a small-scale dataset for evaluating current LMMs capability in detecting these categories of risk.We observe even the latest models remain ineffective to detect various paralinguistic-specific risks in speech (e.g., Gemini 1.5 Pro is performing only slightly above random baseline).1 Warning: this paper contains biased and offensive examples. A Experimental ResultsWe provide complete experimental results including accuracy and macro-averaged F1 score as metrics in Table 6. B Examples for Sub-categoriesWe provide examples from our text sets for each sub-category in Table 7. C Description of Speech Generation from AudioboxWe provide the examples for speech generation from Audiobox in Table 8. D Prompting StrategiesWe provide a complete list covering prompting strategies used in our evaluation experiments and analysis in Table 9 and Table 10, respectively. E Computational Hardware and APIWe conduct all our evaluation experiments and analysis on 4×A100 GPUs.No fine-tuning was done and the experiments only involved inference.For Gemini 1.5 Pro we used gemini-1.5-proAPI, and for GPT-4 we used gpt-4-turbo API.Temperature was set to 0 and sampling at decoding was switched off. Lizhen Qu, Ehsan Shareghi, Reza Haf |
EMNLP | 2 |
| 2024 | Revisiting Deep Audio-Text Retrieval Through the Lens of TransportationabstractThe Learning-to-match (LTM) framework proves to be an effective inverse optimal transport approach for learning the underlying ground metric between two sources of data, facilitating subsequent matching. However, the conventional LTM framework faces scalability challenges, necessitating the use of the entire dataset each time the parameters of the ground metric are updated. In adapting LTM to the deep learning context, we introduce the mini-batch Learning-to-match (m-LTM) framework for audio-text retrieval problems. This framework leverages mini-batch subsampling and Mahalanobis-enhanced family of ground metrics. Moreover, to cope with misaligned training data in practice, we propose a variant using partial optimal transport to mitigate the harm of misaligned data pairs in training data. We conduct extensive experiments on audio-text matching problems using three datasets: AudioCaps, Clotho, and ESC-50. Results demonstrate that our proposed method is capable of learning rich and expressive joint embedding space, which achieves SOTA performance. Beyond this, the proposed m-LTM framework is able to close the modality gap across audio and text embedding, which surpasses both triplet and contrastive loss in the zero-shot sound event detection task on the ESC-50 dataset. Notably, our strategy of employing partial optimal transport with m-LTM demonstrates greater noise tolerance than contrastive loss, especially under varying noise ratios in training data on the AudioCaps dataset. Our code is available at https://github.com/v-manhlt3/m-LTM-Audio-Text-Retrieval Manh Luong, Nhat Ho, Gholamreza Haffari, Dinh Q. Phung, Lizhen Qu |
ICLR | 6 |
| 2024 | Learning in Order! A Sequential Strategy to Learn Invariant Features for Multimodal Sentiment Analysis
Xianbing Zhao, Lizhen Qu, Tao Feng 0013, Jianfei Cai 0001, Buzhou Tang |
ACM Multimedia | 2 |
| 2023 | The Best of Both Worlds: Combining Human and Machine Translations for Multilingual Semantic Parsing with Active LearningabstractMultilingual semantic parsing aims to leverage the knowledge from the high-resource languages to improve low-resource semantic parsing, yet commonly suffers from the data imbalance problem.Prior works propose to utilize the translations by either humans or machines to alleviate such issues.However, human translations are expensive, while machine translations are cheap but prone to error and bias.In this work, we propose an active learning approach that exploits the strengths of both human and machine translations by iteratively adding small batches of human translations into the machine-translated training set.Besides, we propose novel aggregated acquisition criteria that help our active learning method select utterances to be manually translated.Our experiments demonstrate that an ideal utterance selection can significantly reduce the error and bias in the translated data, resulting in higher parser accuracies than the parsers merely trained on the machine-translated data. Zhuang Li 0001, Lizhen Qu, Phil Cohen 0001, Raj Tumuluri, Gholamreza Haffari |
ACL (1) | 2 |
| 2023 | FEDLEGAL: The First Real-World Federated Learning Benchmark for Legal NLPabstractZhuo Zhang, Xiangjing Hu, Jingyuan Zhang, Yating Zhang, Hui Wang, Lizhen Qu, Zenglin Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhuo Zhang 0007, Xiangjing Hu, Wendy Hui Wang, Lizhen Qu, Zenglin Xu |
ACL (1) | 6 |
| 2023 | Document Flattening: Beyond Concatenating Context for Document-Level Neural Machine TranslationabstractExisting work in document-level neural machine translation commonly concatenates several consecutive sentences as a pseudodocument, and then learns inter-sentential dependencies.This strategy limits the model's ability to leverage information from distant context.We overcome this limitation with a novel Document Flattening (DOCFLAT) technique that integrates FLAT-BATCH ATTEN-TION (FBA) and NEURAL CONTEXT GATE (NCG) into Transformer model to utilize information beyond the pseudo-document boundaries.FBA allows the model to attend to all the positions in the batch and learns the relationships between positions explicitly and NCG identifies the useful information from the distant context.We conduct comprehensive experiments and analyses on three benchmark datasets for English-German translation, and validate the effectiveness of two variants of DOCFLAT.Empirical results show that our approach outperforms strong baselines with statistical significance on BLEU, COMET and accuracy on the contrastive test set.The analyses highlight that DOCFLAT is highly effective in capturing the long-range information. Minghao Wu, George F. Foster, Lizhen Qu, Gholamreza Haffari |
EACL | 3 |
| 2023 | Language Independent Neuro-Symbolic Semantic Parsing for Form Understanding
Bhanu Prakash Voutharoja, Lizhen Qu, Fatemeh Shiri |
ICDAR (2) | 2 |
| 2023 | Learning Object-Language Alignments for Open-Vocabulary Object Detection
Chuang Lin 0003, Peize Sun, Yi Jiang 0009, Ping Luo 0002, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan, Jianfei Cai 0001 |
ICLR | 5 |
| 2023 | SocialDial: A Benchmark for Socially-Aware Dialogue SystemsabstractContent Warning: this paper may contain content that is offensive or upsetting. Haolan Zhan, Zhuang Li 0001, Yufei Wang 0003, Linhao Luo, Tao Feng 0013, Xiaoxi Kang, Yuncheng Hua, Lizhen Qu, Lay-Ki Soon, Suraj Sharma, Ingrid Zukerman, Zhaleh Semnani-Azad, Gholamreza Haffari |
SIGIR | 8 |
| 2023 | Less is More: Mitigate Spurious Correlations for Open-Domain Dialogue Response Generation Models by Causal DiscoveryabstractAbstract In this paper, we conduct the first study on spurious correlations for open-domain response generation models based on a corpus CGDialog curated by ourselves. The current models indeed suffer from spurious correlations and have a tendency to generate irrelevant and generic responses. Inspired by causal discovery algorithms, we propose a novel model-agnostic method for training and inference using a conditional independence classifier. The classifier is trained by a constrained self-training method, coined ConSTrain, to overcome data sparsity. The experimental results based on both human and automatic evaluation show that our method significantly outperforms the competitive baselines in terms of relevance, informativeness, and fluency. Tao Feng 0013, Lizhen Qu, Gholamreza Haffari |
Trans. Assoc. Comput. Linguistics | 2 |
| 2022 | Student Surpasses Teacher: Imitation Attack for Black-Box NLP APIsabstractMachine-learning-as-a-service (MLaaS) has attracted millions of users to their splendid large-scale models. Although published as black-box APIs, the valuable models behind these services are still vulnerable to imitation attacks. Recently, a series of works have demonstrated that attackers manage to steal or extract the victim models. Nonetheless, none of the previous stolen models can outperform the original black-box APIs. In this work, we conduct unsupervised domain adaptation and multi-victim ensemble to showing that attackers could potentially surpass victims, which is beyond previous understanding of model extraction. Extensive experiments on both benchmark datasets and real-world APIs validate that the imitators can succeed in outperforming the original black-box models on transferred domains. We consider our work as a milestone in the research of imitation attack, especially on NLP APIs, as the superior performance could influence the defense or even publishing strategy of API providers. Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, Gholamreza Haffari |
COLING | 4 |
| 2022 | Multimodal Transformer with Variable-Length Memory for Vision-and-Language Navigation
Chuang Lin 0003, Yi Jiang 0009, Jianfei Cai 0001, Lizhen Qu, Gholamreza Haffari, Zehuan Yuan |
ECCV (36) | 4 |
| 2022 | Variational Autoencoder with Disentanglement Priors for Low-Resource Task-Specific Natural Language GenerationabstractIn this paper, we propose a variational autoencoder with disentanglement priors, VAE-DPRIOR, for task-specific natural language generation with none or a handful of taskspecific labeled examples.In order to tackle compositional generalization across tasks, our model performs disentangled representation learning by introducing a conditional prior for the latent content space and another conditional prior for the latent label space.Both types of priors satisfy a novel property called ϵ-disentangled.We show both empirically and theoretically that the novel priors can disentangle representations even without specific regularizations as in the prior work.The content prior enables directly sampling diverse content representations from the content space learned from the seen tasks, and fuse them with the representations of novel tasks for generating semantically diverse texts in the low-resource settings.Our extensive experiments demonstrate the superior performance of our model over competitive baselines in terms of i) data augmentation in continuous zero/few-shot learning, and ii) text style transfer in the few-shot setting.The code is available at https://github. com/zhuang-li/VAE-DPrior. Zhuang Li 0001, Lizhen Qu, Qiongkai Xu, Tongtong Wu, Tianyang Zhan, Gholamreza Haffari |
EMNLP | 2 |
| 2022 | Federated Model Decomposition with Private Vocabulary for Text ClassificationabstractWith the necessity of privacy protection, it becomes increasingly vital to train deep neural models in a federated learning manner for natural language processing (NLP) tasks.However, recent studies show eavesdroppers (i.e., dishonest servers) can still reconstruct the private input in federated learning (FL).Such a data reconstruction attack relies on the mappings between vocabulary and associated word embedding in NLP tasks, which are unfortunately less studied in current FL methods.In this paper, we propose a fedrated model decomposition method that protects the privacy of vocabularies, shorted as FEDEVOCAB.In FEDEVOCAB, each participant keeps the local embedding layer in the local device and detaches the local embedding parameters from federated aggregation.However, it is challenging to train an accurate NLP model when the private mappings are unknown and vary across participants in a cross-device FL setting.To address this problem, we further propose an adaptive updating technique to improve the performance of local models.Experimental results show that FEDEVOCAB maintains competitive performance and provides better privacy-preserving capacity compared to status quo methods. * Co-corresponding authorHi, recently I feel very thirsty and urinate more.And I get hungry easily and eat more than usual.But being 165cm tall, my weight even dropped to 40kg.It has continued for half a month.Hi! Does your family have a history of diabetes?If so, there is a high probability of diabetes. Zhuo Zhang 0007, Xiangjing Hu, Lizhen Qu, Qifan Wang 0001, Zenglin Xu |
EMNLP | 3 |
| 2022 | CD-VulD: Cross-Domain Vulnerability Discovery Based on Deep Domain AdaptationabstractA major cause of security incidents such as cyber attacks is rooted in software vulnerabilities. These vulnerabilities should ideally be found and fixed before the code gets deployed. Machine learning-based approaches achieve state-of-the-art performance in capturing vulnerabilities. These methods are predominantly supervised. Their prediction models are trained on a set of ground truth data where the training data and test data are assumed to be drawn from the same probability distribution. However, in practice, the test data often differs from the training data in terms of distribution because they are from different projects or they differ in the types of vulnerability. In this article, we present a new system forCrossDomain SoftwareVulnerabilityDiscovery (CD-VulD) using deep learning (DL) and domain adaptation (DA). We employ DL because it has the capacity of automatically constructing high-level abstract feature representations of programs, which are likely of more cross-domain useful than the handcrafted features driven by domain knowledge. The divergence between distributions is reduced by learning cross-domain representations. First, given software program representations, CD-VulD converts them into token sequences and learns the token embeddings for generalization across tokens. Next, CD-VulD employs a deep feature model to build abstract high-level presentations based on those sequences. Then, the metric transfer learning framework (MTLF) technique is employed to learn cross-domain representations by minimizing the distribution divergence between the source domain and the target domain. Finally, the cross-domain representations are used to build a classifier for vulnerability detection. Experimental results show that CD-VulD outperforms the state-of-the-art vulnerability detection approaches by a wide margin. We make the new datasets publicly available so that our work is replicable and can be further improved. Shigang Liu, Guanjun Lin, Lizhen Qu, Jun Zhang 0010, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2021 | Causal Relationships Between Emotions and Dialog ActsabstractEmotions and Dialog Acts (DAs) are phenomena in interpersonal communication that are informative of a person’s underlying thoughts and feelings. Thus, affective computing and dialog system researchers are interested in understanding and detecting emotions and DAs in dialogs. Previous studies on emotion recognition and dialog act classification observed that utilizing one feature for classification of the other can improve model performance, which benefits the advancement of affective dialog systems. However, theoretical explanation of such gain remains unclear. In linguistic research, the relationship between an emotion and a DA is often investigated qualitatively. In this work, we conducted both qualitative and quantitative analyses examining the relationships between emotions and DAs. Through statistical analyses, causal discovery methods, and a crowd-sourcing study, using two datasets of emotional dialog we identified emotion-DA causal pairs that provide empirical evidences to existing linguistic theories, while supporting addressing these two features together in affective dialog systems. Moreover, our work offers an effective methodology for revealing linguistic discoveries with a data-driven approach. Shuyi Cao, Lizhen Qu, Leimin Tian |
ACII | 2 |
| 2021 | On Robustness of Neural Semantic ParsersabstractSemantic parsing maps natural language (NL) utterances into logical forms (LFs), which underpins many advanced NLP problems.Semantic parsers gain performance boosts with deep neural networks, but inherit vulnerabilities against adversarial examples.In this paper, we provide the empirical study on the robustness of semantic parsers in the presence of adversarial attacks.Formally, adversaries of semantic parsing are considered to be the perturbed utterance-LF pairs, whose utterances have exactly the same meanings as the original ones.A scalable methodology is proposed to construct robustness test sets based on existing benchmark corpora.Our results answered five research questions in measuring the sateof-the-art parsers' performance on robustness test sets, and evaluating the effect of data augmentation. Zhuang Li 0001, Lizhen Qu |
EACL | 3 |
| 2021 | Few-Shot Semantic Parsing for New PredicatesabstractIn this work, we investigate the problems of semantic parsing in a few-shot learning setting.In this setting, we are provided with k utterance-logical form pairs per new predicate.The state-of-the-art neural semantic parsers achieve less than 25% accuracy on benchmark datasets when k = 1.To tackle this problem, we proposed to i) apply a designated metalearning method to train the model; ii) regularize attention scores with alignment statistics; iii) apply a smoothing technique in pretraining.As a result, our method consistently outperforms all the baselines in both one and two-shot settings. Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari |
EACL | 2 |
| 2021 | Total Recall: a Customized Continual Learning Method for Neural Semantic ParsersabstractThis paper investigates continual learning for semantic parsing.In this setting, a neural semantic parser learns tasks sequentially without accessing full training data from previous tasks.Direct application of the SOTA continual learning algorithms to this problem fails to achieve comparable performance with retraining models with all seen tasks, because they have not considered the special properties of structured outputs, yielded by semantic parsers.Therefore, we propose TO-TAL RECALL, a continual learning method designed for neural semantic parsers from two aspects: i) a sampling method for memory replay that diversifies logical form templates and balances distributions of parse actions in a memory; ii) a two-stage training method that significantly improves generalization capability of the parsers across tasks.We conduct extensive experiments to study the research problems involved in continual semantic parsing, and demonstrate that a neural semantic parser trained with TOTAL RECALL achieves superior performance than the one trained directly with the SOTA continual learning algorithms, and achieve a 3-6 times speedup compared to retraining from scratch.Code and datasets are available at: https://github. com/zhuang-li/cl_nsp. Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari |
EMNLP (1) | 2 |
| 2021 | Privacy Monitoring Service for ConversationsabstractLeakage of personal information in conversations raises serious privacy concerns. Malicious people or bots could pry into sensitive personal information of vulnerable people, such as juveniles, through conversations with them or their digital personal assistants. To address the problem, we present a privacy-leakage warning system that monitors conversations in social media and intercepts the outgoing text messages from a user or a digital assistant, if they impose potential privacy leakage risks. Such messages are redirected to authorized users for approval, before they are sent out. We demonstrate how our system is deployed and used on a social media conversation platform, e.g., Facebook Messenger. Qiongkai Xu, Lizhen Qu |
WSDM | 3 |
| 2021 | Easy-to-Deploy API Extraction by Multi-Level Feature Embedding and Transfer LearningabstractApplication Programming Interfaces (APIs) have been widely discussed on social-technical platforms (e.g., Stack Overflow). Extracting API mentions from such informal software texts is the prerequisite for API-centric search and summarization of programming knowledge. Machine learning based API extraction has demonstrated superior performance than rule-based methods in informal software texts that lack consistent writing forms and annotations. However, machine learning based methods have a significant overhead in preparing training data and effective features. In this paper, we propose a multi-layer neural network based architecture for API extraction. Our architecture automatically learns character-, word- and sentence-level features from the input texts, thus removing the need for manual feature engineering and the dependence on advanced features (e.g., API gazetteers) beyond the input texts. We also propose to adopt transfer learning to adapt a source-library-trained model to a target-library, thus reducing the overhead of manual training-data labeling when the software text of multiple programming languages and libraries need to be processed. We conduct extensive experiments with six libraries of four programming languages which support diverse functionalities and have different API-naming and API-mention characteristics. Our experiments investigate the performance of our neural architecture for API extraction in informal software texts, the importance of different features, the effectiveness of transfer learning. Our results confirm not only the superior performance of our neural architecture than existing machine learning based methods for API extraction in informal software texts, but also the easy-to-deploy characteristic of our neural architecture. Suyu Ma, Zhenchang Xing, Chunyang Chen 0001, Lizhen Qu, Guoqiang Li 0001 |
IEEE Trans. Software Eng. | 5 |
| 2020 | Adhering, Steering, and Queering: Treatment of Gender in Natural Language GenerationabstractNatural Language Generation (NLG) supports the creation of personalized, contextualized, and targeted content. However, the algorithms underpinning NLG have come under scrutiny for reinforcing gender, racial, and other problematic biases. Recent research in NLG seeks to remove these biases through principles of fairness and privacy. Drawing on gender and queer theories from sociology and Science and Technology studies, we consider how NLG can contribute towards the advancement of gender equity in society. We propose a conceptual framework and technical parameters for aligning NLG with feminist HCI qualities. We present three approaches: (1) adhering to current approaches of removing sensitive gender attributes, (2) steering gender differences away from the norm, and (3) queering gender by troubling stereotypes. We discuss the advantages and limitations of these approaches across three hypothetical scenarios; newspaper headlines, job advertisements, and chatbots. We conclude by discussing considerations for implementing this framework and related ethical and equity agendas. Yolande A. A. Strengers, Lizhen Qu, Qiongkai Xu, Jarrod Knibbe |
CHI | 2 |
| 2020 | Context Dependent Semantic Parsing: A SurveyabstractSemantic parsing is the task of translating natural language utterances into machine-readable meaning representations.Currently, most semantic parsing methods are not able to utilize contextual information (e.g.dialogue and comments history), which has a great potential to boost semantic parsing performance.To address this issue, context dependent semantic parsing has recently drawn a lot of attention.In this survey, we investigate progress on the methods for the context dependent semantic parsing, together with the current datasets and tasks.We then point out open problems and challenges for future research in this area.The collected resources for this topic are available at: https://github. Zhuang Li 0001, Lizhen Qu, Gholamreza Haffari |
COLING | 2 |
| 2020 | CosMo: Conditional Seq2Seq-based Mixture Model for Zero-Shot Commonsense Question AnsweringabstractCommonsense reasoning refers to the ability of evaluating a social situation and acting accordingly.Identification of the implicit causes and effects of a social context is the driving capability which can enable machines to perform commonsense reasoning.The dynamic world of social interactions requires context-dependent on-demand systems to infer such underlying information.However, current approaches in this realm lack the ability to perform commonsense reasoning upon facing an unseen situation, mostly due to incapability of identifying a diverse range of implicit social relations.Hence they fail to estimate the correct reasoning path.In this paper, we present Conditional SEQ2SEQ-based Mixture model (COSMO), which provides us with the capabilities of dynamic and diverse content generation.We use COSMO to generate context-dependent clauses, which form a dynamic Knowledge Graph (KG) on-the-fly for commonsense reasoning.To show the adaptability of our model to context-dependant knowledge generation, we address the task of zero-shot commonsense question answering.The empirical results indicate an improvement of up to +5.2% over the state-of-the-art models. Farhad Moghimifar, Lizhen Qu, Terry Yue Zhuo, Mahsa Baktash, Gholamreza Haffari |
COLING | 2 |
| 2020 | Personal Information Leakage Detection in ConversationsabstractThe global market size of conversational assistants (chatbots) is expected to grow to USD 9.4 billion by 2024, according to Marketsand-Markets.Despite the wide use of chatbots, leakage of personal information through chatbots poses serious privacy concerns for their users.In this work, we propose to protect personal information by warning users of detected suspicious sentences generated by conversational assistants.The detection task is formulated as an alignment optimization problem and a new dataset PERSONA-LEAKAGE is collected for evaluation.In this paper, we propose two novel constrained alignment models, which consistently outperform baseline methods on PERSONA-LEAKAGE 1 .Moreover, we conduct analysis on the behavior of recently proposed personalized chit-chat dialogue systems.The empirical results show that those systems suffer more from personal information disclosure than the widely used Seq2Seq model and the language model.In those cases, a significant number of information leaking utterances can be detected by our models with high precision. Qiongkai Xu, Lizhen Qu, Gholamreza Haffari |
EMNLP (1) | 2 |
| 2019 | Machine Reading Comprehension: Matching and OrdersabstractIn this paper, we study the machine reading comprehension of temporal order in text. Given a document of instruction sequences, a model aims to find out the most coherent sequences of activities matching the document among all answer candidates. To tackle the task, we proposeOrdMatch model, which is able to match each activity in a sequence to the corresponding instruction in the document and regularizes the partial order of activities to match the order of instructions. We evaluate the task using the RecipeQA dataset, which includes step-by-step instructions of cooking recipes. Our model outperforms the state-of-the-art models with a wide margin. The experimental results demonstrate the effectiveness of our novel ordering regularizer. Our code will be made available at \hrefhttps://github.com/Aolius/OrdMatch https://github.com/Aolius/OrdMatch. Ao Liu 0008, Lizhen Qu, Chenbin Zhang, Zenglin Xu |
CIKM | 2 |
| 2019 | Maximal Divergence Sequential Autoencoder for Binary Software Vulnerability Detection
Tue Le, Tuan Nguyen 0004, Trung Le 0001, Dinh Q. Phung, Paul Montague, Olivier Y. de Vel, Lizhen Qu |
ICLR (Poster) | 7 |
| 2019 | Deep Domain Adaptation for Vulnerable Code Function IdentificationabstractDue to the ubiquity of computer software, software vulnerability detection (SVD) has become crucial in the software industry and in the field of computer security. Two significant issues in SVD arise when using machine learning, namely: i) how to learn automatic features that can help improve the predictive performance of vulnerability detection and ii) how to overcome the scarcity of labeled vulnerabilities in projects that require the laborious labeling of code by software security experts. In this paper, we address these two crucial concerns by proposing a novel architecture which leverages deep domain adaptation with automatic feature learning for software vulnerability identification. Based on this architecture, we keep the principles and reapply the state-of-the-art deep domain adaptation methods to indicate that deep domain adaptation for SVD is plausible and promising. Moreover, we further propose a novel method named Semi-supervised Code Domain Adaptation Network (SCDAN) that can efficiently utilize and exploit information carried in unlabeled target data by considering them as the unlabeled portion in a semi-supervised learning context. The proposed SCDAN method enforces the clustering assumption, which is a key principle in semi-supervised learning. The experimental results using six real-world software project datasets show that our SCDAN method and the baselines using our architecture have better predictive performance by a wide margin compared with the Deep Code Network (VulDeePecker) method without domain adaptation. Also, the proposed SCDAN significantly outperforms the DIRT-T which to the best of our knowledge is currently the-state-of-the-art method in deep domain adaptation and other baselines. Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, Lizhen Qu, Dinh Q. Phung |
IJCNN | 7 |
| 2019 | Privacy-Aware Text RewritingabstractBiased decisions made by automatic systems have led to growing concerns in research communities. Recent work from the NLP community focuses on building systems that make fair decisions based on text. Instead of relying on unknown decision systems or human decision-makers, we argue that a better way to protect data providers is to remove the trails of sensitive information before publishing the data. In light of this, we propose a new privacy-aware text rewriting task and explore two privacy-aware back-translation methods for the task, based on adversarial training and approximate fairness risk. Our extensive experiments on three real-world datasets with varying demo-graphical attributes show that our methods are effective in obfuscating sensitive attributes. We have also observed that the fairness risk method retains better semantics and fluency, while the adversarial training method tends to leak less sensitive information. Qiongkai Xu, Lizhen Qu, Ran Cui |
INLG | 2 |
| 2017 | Attentive Graph-based Recursive Neural Network for Collective Vertex ClassificationabstractVertex classification is a critical task in graph analysis, where both contents and linkage of vertices are incorporated during classification. Recently, researchers proposed using deep neural network to build an end-to-end framework, which can capture both local content and structure information. These approaches were proved effective in incorporating semantic meanings of neighbouring vertices, while the usefulness of this information was not properly considered. In this paper, we propose an Attentive Graph-based Recursive Neural Network (AGRNN), which exerts attention on neural network to make our model focus on vertices with more relevant semantic information. We evaluated our approach on three real-world datasets and also datasets with synthetic noise. Our experimental results show that AGRNN achieves the state-of-the-art performance, in terms of effectiveness and robustness. We have also illustrated some attention weight samples to demonstrate the rationality of our model. Qiongkai Xu, Qing Wang 0002, Lizhen Qu |
CIKM | 4 |
| 2017 | Making Deep Neural Networks Robust to Label Noise: A Loss Correction ApproachabstractWe present a theoretically grounded approach to train deep neural networks, including recurrent networks, subject to class-dependent label noise. We propose two procedures for loss correction that are agnostic to both application domain and network architecture. They simply amount to at most a matrix inversion and multiplication, provided that we know the probability of each class being corrupted into another. We further show how one can estimate these probabilities, adapting a recent technique for noise estimation to the multi-class setting, and thus providing an end-to-end framework. Extensive experiments on MNIST, IMDB, CIFAR-10, CIFAR-100 and a large scale dataset of clothing images employing a diversity of architectures - stacking dense, convolutional, pooling, dropout, batch normalization, word embedding, LSTM and residual layers - demonstrate the noise robustness of our proposals. Incidentally, we also prove that, when ReLU is the only non-linearity, the loss curvature is immune to class-dependent label noise. Giorgio Patrini, Alessandro Rozza, Aditya Krishna Menon, Richard Nock, Lizhen Qu |
CVPR | 5 |
| 2017 | Automatic Generation of Grounded Visual QuestionsabstractIn this paper, we propose the first model to be able to generate visually grounded questions with diverse types for a single image. Visual question generation is an emerging topic which aims to ask questions in natural language based on visual input. To the best of our knowledge, it lacks automatic methods to generate meaningful questions with various types for the same visual input. To circumvent the problem, we propose a model that automatically generates visually grounded questions with varying types. Our model takes as input both images and the captions generated by a dense caption model, samples the most probable question types, and generates the questions in sequel. The experimental results on two real world datasets show that our model outperforms the strongest baseline in terms of both correctness and diversity with a wide margin. Lizhen Qu, Shaodi You, Zhenglu Yang, Jiawan Zhang |
IJCAI | 2 |
| 2017 | f-GANs in an Information Geometric NutshellabstractNowozin \textit{et al} showed last year how to extend the GAN \textit{principle} to all $f$-divergences. The approach is elegant but falls short of a full description of the supervised game, and says little about the key player, the generator: for example, what does the generator actually converge to if solving the GAN game means convergence in some space of parameters? How does that provide hints on the generator's design and compare to the flourishing but almost exclusively experimental literature on the subject? In this paper, we unveil a broad class of distributions for which such convergence happens --- namely, deformed exponential families, a wide superset of exponential families ---. We show that current deep architectures are able to factorize a very large number of such densities using an especially compact design, hence displaying the power of deep architectures and their concinnity in the $f$-GAN game. This result holds given a sufficient condition on \textit{activation functions} --- which turns out to be satisfied by popular choices. The key to our results is a variational generalization of an old theorem that relates the KL divergence between regular exponential families and divergences between their natural parameters. We complete this picture with additional results and experimental insights on how these results may be used to ground further improvements of GAN architectures, via (i) a principled design of the activation functions in the generator and (ii) an explicit integration of proper composite losses' link function in the discriminator. Richard Nock, Zac Cranko, Aditya Krishna Menon, Lizhen Qu, Robert C. Williamson |
NIPS | 4 |
| 2016 | Neighborhood Mixture Model for Knowledge Base CompletionabstractKnowledge bases are useful resources for many natural language processing tasks, however, they are far from complete.In this paper, we define a novel entity representation as a mixture of its neighborhood in the knowledge base and apply this technique on TransE-a well-known embedding model for knowledge base completion.Experimental results show that the neighborhood information significantly helps to improve the results of the TransE, leading to better performance than obtained by other state-of-the-art embedding models on three benchmark datasets for triple classification, entity prediction and relation prediction tasks. Dat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson 0001 |
CoNLL | 3 |
| 2016 | Named Entity Recognition for Novel Types by Transfer LearningabstractIn named entity recognition, we often don't have a large in-domain training corpus or a knowledge base with adequate coverage to train a model directly.In this paper, we propose a method where, given training data in a related domain with similar (but not identical) named entity (NE) types and a small amount of in-domain training data, we use transfer learning to learn a domain-specific NE model.That is, the novelty in the task setup is that we assume not just domain mismatch, but also label mismatch. Lizhen Qu, Gabriela Ferraro, Liyuan Zhou, Weiwei Hou, Timothy Baldwin |
EMNLP | 1 |
| 2016 | STransE: a novel embedding model of entities and relationships in knowledge basesabstractDat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Dat Quoc Nguyen, Kairit Sirts, Lizhen Qu, Mark Johnson 0001 |
HLT-NAACL | 3 |
| 2015 | Big Data Small Data, In Domain Out-of Domain, Known Word Unknown Word: The Impact of Word Representations on Sequence Labelling TasksabstractWord embeddings -distributed word representations that can be learned from unlabelled data -have been shown to have high utility in many natural language processing applications.In this paper, we perform an extrinsic evaluation of four popular word embedding methods in the context of four sequence labelling tasks: part-of-speech tagging, syntactic chunking, named entity recognition, and multiword expression identification.A particular focus of the paper is analysing the effects of task-based updating of word representations.We show that when using word embeddings as features, as few as several hundred training instances are sufficient to achieve competitive results, and that word embeddings lead to improvements over out-of-vocabulary words and also out of domain.Perhaps more surprisingly, our results indicate there is little difference between the different word embedding methods, and that simple Brown clusters are often competitive with word embeddings across all tasks we consider. Lizhen Qu, Gabriela Ferraro, Liyuan Zhou, Weiwei Hou, Nathan Schneider 0001, Timothy Baldwin |
CoNLL | 1 |
| 2014 | Senti-LSSVM: Sentiment-Oriented Multi-Relation Extraction with Latent Structural SVMabstractExtracting instances of sentiment-oriented relations from user-generated web documents is important for online marketing analysis. Unlike previous work, we formulate this extraction task as a structured prediction problem and design the corresponding inference as an integer linear program. Our latent structural SVM based model can learn from training corpora that do not contain explicit annotations of sentiment-bearing expressions, and it can simultaneously recognize instances of both binary (polarity) and ternary (comparative) relations with regard to entity mentions of interest. The empirical evaluation shows that our approach significantly outperforms state-of-the-art systems across domains (cameras and movies) and across genres (reviews and forum posts). The gold standard corpus that we built will also be a valuable resource for the community. Lizhen Qu, Yi Zhang 0003, Rui Wang 0005, Lili Jiang 0002, Rainer Gemulla, Gerhard Weikum |
Trans. Assoc. Comput. Linguistics | 1 |
| 2012 | A Weakly Supervised Model for Sentence-Level Semantic Orientation Analysis with Multiple Experts
Lizhen Qu, Rainer Gemulla, Gerhard Weikum |
EMNLP-CoNLL | 1 |
| 2011 | Harvesting facts from textual web sources by constrained label propagationabstractThere have been major advances on automatically constructing large knowledge bases by extracting relational facts from Web and text sources. However, the world is dynamic: periodic events like sports competitions need to be interpreted with their respective timepoints, and facts such as coaching a sports team, holding political or business positions, and even marriages do not hold forever and should be augmented by their respective timespans. This paper addresses the problem of automatically harvesting temporal facts with such extended time-awareness. We employ pattern-based gathering techniques for fact candidates and construct a weighted pattern-candidate graph. Our key contribution is a system called PRAVDA based on a new kind of label propagation algorithm with a judiciously designed loss function, which iteratively processes the graph to label good temporal facts for a given set of target relations. Our experiments with online news and Wikipedia articles demonstrate the accuracy of this method. Yafang Wang, Bin Yang 0002, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
CIKM | 3 |
| 2010 | The Bag-of-Opinions Method for Review Rating Prediction from Sparse Text Patterns
Lizhen Qu, Georgiana Ifrim, Gerhard Weikum |
COLING | 1 |
| 2010 | Timely YAGO: harvesting, querying, and visualizing temporal knowledge from WikipediaabstractRecent progress in information extraction has shown how to automatically build large ontologies from high-quality sources like Wikipedia. But knowledge evolves over time; facts have associated validity intervals. Therefore, ontologies should include time as a first-class dimension. In this paper, we introduce Timely YAGO, which extends our previously built knowledge base YAGO with temporal aspects. This prototype system extracts temporal facts from Wikipedia infoboxes, categories, and lists in articles, and integrates these into the Timely YAGO knowledge base. We also support querying temporal facts, by temporal predicates in a SPARQL-style language. Visualization of query results is provided in order to better understand of the dynamic nature of knowledge. Yafang Wang, Mingjie Zhu, Lizhen Qu, Marc Spaniol, Gerhard Weikum |
EDBT | 3 |
| 2008 | Using tag semantic network for keyphrase extraction in blogsabstractFolksonomies provide a comfortable way to search and browse the blogosphere. As the tags in the blogosphere are sparse, ambiguous and too general, this paper proposes both a supervised and an unsupervised approach that extract tags from posts using a tag semantic network. We evaluate the two methods on a blog dataset and observe an improvement in F1-measure from 0.23 to 0.50 when compared to the baseline system. Lizhen Qu, Christof Müller, Iryna Gurevych |
CIKM | 1 |