Xin Jiang 0002

dblp:42/4142-2 · DBLP profile ↗
← Back
93ranked-venue papers
1as first author
74since 2021 · last 2026
0000-0002-9117-8247ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 82 · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 15 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Computer networks · 4 · 4 since 2021
YearPublicationVenuePosition
2026 ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learning
abstract
Tool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning.
Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001
AAAI9
2026 Analyzing how pre-trained language models capture factual knowledge using attribution methods
Shaobo Li 0004, Chengjie Sun, Bingquan Liu, Lifeng Shang, Zhenhua Dong, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001
Knowl. Based Syst.8
2026 An Efficient Edge-Cloud Collaboration System With Foundational Models for Open-Set IoT Applications
abstract
Artificial intelligence (AI) models have been widely deployed on edge devices, enabling various IoT applications. However, lightweight on-device AI models on resource-limited edge devices hinder their adaptability to dynamic environments and tasks. Despite the superior generalization capabilities of recently developed Foundation Models (FMs), utilizing their extensive knowledge on the resource-constrained edge platforms remains unexplored. In this work, we introduce DeepEdgeFM, an edge-cloud collaborative system with FMs that enables open-set learning, simultaneously achieving generalizability and efficiency for IoT applications. DeepEdgeFM employs a spatiotemporalaware semantic customization approach that leverages spatial, temporal, and domain-specific knowledge from FMs to continuously customize edge models using unlabeled sensor data in emerging IoT environments. Meanwhile, DeepEdgeFM utilizes a dynamic model switching strategy to selectively query the knowledge of FMs based on sensor-data uncertainty and real-time network fluctuations. We implement DeepEdgeFM on five FMs and multi-modal large language models (MLLMs), covering four types of sensor data modalities. We evaluate DeepEdgeFM on two edge platforms, five public datasets, and two self-collected datasets covering both indoor and outdoor real-world environments. The results show that DeepEdgeFM outperforms state-ofthe- art baselines, achieving up to an 18.6% accuracy gain and a 38.6
Bufang Yang, Wenrui Lu, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002
IEEE Trans. Mob. Comput.9
2025 Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-Alignment
abstract
Zhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok
ACL (1)9
2025 Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
abstract
Kaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002
ACL (1)7
2025 Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
abstract
Qiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Qiyuan Zhang 0001, Yufei Wang 0005, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ACL (1)7
2025 Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP8
2025 Bridging and Modeling Correlations in Pairwise Data for Direct Preference Optimization
abstract
Direct preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants.
Bo Huang 0017, Yufei Wang 0005, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Wei Wang 0011
ICLR7
2025 Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced Exploration
abstract
Large language models (LLMs) have significantly benefited from training on diverse, high-quality task-specific data, leading to impressive performance across a range of downstream applications. Current methods often rely on human-annotated data or predefined task templates to direct powerful LLMs in synthesizing task-relevant data for effective model training. However, this dependence on manually designed components may constrain the scope of generated data, potentially overlooking critical edge cases or novel scenarios that could challenge the model. In this paper, we present a novel approach, ReverseGen, designed to automatically generate effective training samples that expose the weaknesses of LLMs. Specifically, we introduce a dedicated proposer trained to produce queries that lead target models to generate unsatisfactory responses. These failure-inducing queries are then used to construct training data, helping to address the models' shortcomings and improve overall performance. Our approach is flexible and can be applied to models of various scales (3B, 7B, and 8B). We evaluate ReverseGen on three key applications—safety, honesty, and math—demonstrating that our generated data is both highly effective and diverse. Models fine-tuned with ReverseGen-generated data consistently outperform those trained on human-annotated or general model-generated data, offering a new perspective on data synthesis for task-specific LLM enhancement.
Qintong Li, Jiahui Gao 0002, Renjie Pi, Xueliang Zhao, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR7
2025 ToolACE: Winning the Points of LLM Function Calling
abstract
Function calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE.
Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen
ICLR23
2025 Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
abstract
Autoregressive language models, despite their impressive capabilities, struggle with complex reasoning and long-term planning tasks. We introduce discrete diffusion models as a novel solution to these challenges. Through the lens of subgoal imbalance, we demonstrate how diffusion models effectively learn difficult subgoals that elude autoregressive approaches. We propose Multi-Granularity Diffusion Modeling (MGDM), which prioritizes subgoals based on difficulty during learning. On complex tasks like Countdown, Sudoku, and Boolean Satisfiability Problems, MGDM significantly outperforms autoregressive models without using search techniques. For instance, MGDM achieves 91.5\% and 100\% accuracy on Countdown and Sudoku, respectively, compared to 45.8\% and 20.7\% for autoregressive models. Our work highlights the potential of diffusion-based approaches in advancing AI capabilities for sophisticated language understanding and problem-solving tasks. All associated codes are available at \href{https://github.com/HKUNLP/diffusion-vs-ar}{https://github.com/HKUNLP/diffusion-vs-ar}.
Jiacheng Ye, Jiahui Gao 0002, Shansan Gong, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR5
2025 Implicit Search via Discrete Diffusion: A Study on Chess
abstract
In the post-AlphaGo era, there has been a renewed interest in search techniques such as Monte Carlo Tree Search (MCTS), particularly in their application to Large Language Models (LLMs). This renewed attention is driven by the recognition that current next-token prediction models often lack the ability for long-term planning. Is it possible to instill search-like abilities within the models to enhance their planning abilities without relying on explicit search? We propose DiffuSearch , a model that does \textit{implicit search} by looking into the future world via discrete diffusion modeling. We instantiate DiffuSearch on a classical board game, Chess, where explicit search is known to be essential. Through extensive controlled experiments, we show DiffuSearch outperforms both the searchless and explicit search-enhanced policies. Specifically, DiffuSearch outperforms the one-step policy by 19.2\% and the MCTS-enhanced policy by 14\% on action accuracy. Furthermore, DiffuSearch demonstrates a notable 30\% enhancement in puzzle-solving abilities compared to explicit search-based policies, along with a significant 540 Elo increase in game-playing strength assessment. These results indicate that implicit search via discrete diffusion is a viable alternative to explicit search over a one-step policy. All codes are publicly available at \href{https://github.com/HKUNLP/DiffuSearch}{https://github.com/HKUNLP/DiffuSearch}.
Jiacheng Ye, Jiahui Gao 0002, Zhiyong Wu 0003, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong
ICLR5
2025 RevisEval: Improving LLM-as-a-Judge via Response-Adapted References
abstract
With significant efforts in recent studies, LLM-as-a-Judge has become a cost-effective alternative to human evaluation for assessing text generation quality in a wide range of tasks. However, there still remains a reliability gap between LLM-as-a-Judge and human evaluation. One important reason is the lack of guided oracles in the evaluation process. Motivated by the role of reference pervasively used in classic text evaluation, we introduce RevisEval, a novel text generation evaluation paradigm via the response-adapted references. RevisEval is driven by the key observation that an ideal reference should maintain the necessary relevance to the response to be evaluated. Specifically, RevisEval leverages the text revision capabilities of large language models (LLMs) to adaptively revise the response, then treat the revised text as the reference (response-adapted reference) for the subsequent evaluation. Extensive experiments demonstrate that RevisEval outperforms traditional reference-free and reference-based evaluation paradigms that use LLM-as-a-Judge across NLG tasks and open-ended instruction-following tasks. More importantly, our response-adapted references can further boost the classical text metrics, e.g., BLEU and BERTScore, compared to traditional references and even rival the LLM-as-a-Judge. A detailed analysis is also conducted to confirm RevisEval's effectiveness in bias reduction, the impact of inference cost, and reference relevance.
Qiyuan Zhang 0001, Yufei Wang 0005, Tiezheng Yu, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001
ICLR8
2025 FlatQuant: Flatness Matters for LLM Quantization
abstract
Recently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with equally spaced quantization points. Prior research explores various pre-quantization transformations to suppress outliers, such as per-channel scaling and Hadamard transformation. However, we observe that these transformed weights and activations can still exhibit steep and dispersed distributions. In this paper, we propose FlatQuant (Fast and Learnable Affine Transformation), a new post-training quantization approach that enhances the flatness of weights and activations. Our approach identifies optimal affine transformations for each linear layer, calibrated in hours via a lightweight objective. To reduce runtime overhead of affine transformation, we apply Kronecker product with two lightweight matrices, and fuse all operations in FlatQuant into a single kernel. Extensive experiments demonstrate that FlatQuant establishes a new state-of-the-art benchmark for quantization. For example, it achieves less than 1\% accuracy drop for W4A4 quantization on the LLaMA-3-70B model, surpassing SpinQuant by 7.5\%. Additionally, it provides up to 2.3x prefill speedup and 1.7x decoding speedup compared to the FP16 model. Code is available at: https://github.com/ruikangliu/FlatQuant.
Ruikang Liu, Haoli Bai, Yuening Li, Xianzhi Yu, Lu Hou 0002, Chun Yuan 0003, Xin Jiang 0002, Wulong Liu
ICML11
2025 ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis
abstract
Zezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL (Long Papers)7
2025 TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMs
abstract
An increasing number of environments, such as smart homes and factories, are being equipped with multiple sensor systems to enable diverse intelligent applications. However, most existing sensor coordination systems require manually predefined rules, limiting their ability to handle flexible and complex tasks. While recent approaches leverage large language models (LLMs) to interact with external APIs, they struggle to fully understand the capabilities and data dependencies of practical sensor systems. This paper introduces TaskSense, a novel system that coordinates multiple sensor systems in response to users' complex queries. TaskSense introduces a sensor language that automatically translates the capabilities and data dependencies of sensor systems into vocabularies and grammar rules that can be understood by LLMs. It then interprets user intentions into executable task plans for sensor systems using this sensor language in combination with LLMs. Meanwhile, TaskSense checks the solvability of user queries and verifies the correctness of task plan dependencies. To further enhance robustness, TaskSense incorporates a dynamic plan execution mechanism that adjusts plans based on real-time feedback from sensor data availability, data quality and execution results. TaskSense is deployed on real-world smart home systems, utilizing six popular LLMs. The system is evaluated across 4 scenarios involving 9 types of sensor systems, over 60 APIs, 170 tasks and 5 types of data modalities. Results show that TaskSense achieves up to 2× higher planning accuracy and a 75% increase in answer accuracy using the similar amount of tokens compared with baseline approaches.
Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002
SenSys8
2025 Enhancing inter-sentence coherence of extractive summarization with multitask learning
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
J. Intell. Inf. Syst.4
2024 Preparing Lessons for Progressive Training on Language Models
abstract
The rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
AAAI8
2024 Unsupervised Extractive Summarization with Learnable Length Control Strategies
abstract
Unsupervised extractive summarization is an important technique in information extraction and retrieval. Compared with supervised method, it does not require high-quality human-labelled summaries for training and thus can be easily applied for documents with different types, domains or languages. Most of existing unsupervised methods including TextRank and PACSUM rely on graph-based ranking on sentence centrality. However, this scorer can not be directly applied in end-to-end training, and the positional-related prior assumption is often needed for achieving good summaries. In addition, less attention is paid to length-controllable extractor, where users can decide to summarize texts under particular length constraint. This paper introduces an unsupervised extractive summarization model based on a siamese network, for which we develop a trainable bidirectional prediction objective between the selected summary and the original document. Different from the centrality-based ranking methods, our extractive scorer can be trained in an end-to-end manner, with no other requirement of positional assumption. In addition, we introduce a differentiable length control module by approximating 0-1 knapsack solver for end-to-end length-controllable extracting. Experiments show that our unsupervised method largely outperforms the centrality-based baseline using a same sentence encoder. In terms of length control ability, via our trainable knapsack module, the performance consistently outperforms the strong baseline without utilizing end-to-end training. Human evaluation further evidences that our method performs the best among baselines in terms of relevance and consistency.
Renlong Jie, Xiaojun Meng, Xin Jiang 0002, Qun Liu 0001
AAAI3
2024 FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
abstract
Yuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011
ACL (1)8
2024 Learning to Edit: Aligning LLMs with Knowledge Editing
abstract
Yuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011
ACL (1)8
2024 MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
abstract
Wai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
EMNLP7
2024 Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
abstract
The rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility.
Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR13
2024 Retrieval-based Disentangled Representation Learning with Natural Language Supervision
abstract
Disentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all its variations within a finite set of factors. However, it is worth noting that most real-world data have linguistic equivalents, typically in the form of textual descriptions. These linguistic counterparts can represent the data and effortlessly decomposed into distinct tokens. In light of this, we present Vocabulary Disentangled Retrieval (VDR), a retrieval-based framework that harnesses natural language as proxies of the underlying data variation to drive disentangled representation learning. Our approach employ a bi-encoder model to represent both data and natural language in a vocabulary space, enabling the model to distinguish dimensions that capture intrinsic characteristics within data through its natural language counterpart, thus facilitating disentanglement. We extensively assess the performance of VDR across 15 retrieval benchmark datasets, covering text-to-text and cross-modal retrieval scenarios, as well as human evaluation. Our experimental results compellingly demonstrate the superiority of VDR over previous bi-encoder retrievers with comparable model size and training costs, achieving an impressive 8.7% improvement in NDCG@10 on the BEIR benchmark, a 5.3\% increase on MS COCO, and a 6.0% increase on Flickr30k in terms of mean recall in the zero-shot setting. Moreover, The results from human evaluation indicate that interpretability of our method is on par with SOTA captioning models.
Jiawei Zhou 0003, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002
ICLR4
2024 Visually Guided Generative Text-Layout Pre-training for Document Intelligence
abstract
Zhiming Mao, Haoli Bai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhiming Mao, Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong
NAACL-HLT5
2024 Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language Models
abstract
Recently, diffusion models have garnered significant interest in the field of text processing due to their many potential advantages compared to conventional autoregressive models. In this work, we propose Diffusion-of-Thought (DoT), a novel approach that integrates diffusion models with Chain-of-Thought, a well-established technique for improving the reasoning ability of autoregressive language models. In contrast to autoregressive language models that make decisions in a left-to-right, token-by-token manner, DoT allows reasoning steps to diffuse over time through a diffusion language model and offers greater flexibility in trading-off computation for reasoning performance. Our experimental results demonstrate the effectiveness of DoT in multi-digit multiplication, boolean logic, and grade school math problems. In addition to that, DoT showcases promising self-correction abilities and benefits from existing reasoning-enhancing techniques like self-consistency decoding. Our findings contribute to the understanding and development of reasoning with diffusion language models.
Jiacheng Ye, Shansan Gong, Jiahui Gao 0002, Xin Jiang 0002, Zhenguo Li, Wei Bi, Lingpeng Kong
NeurIPS8
2024 Poster Abstract: Tasking Heterogeneous Sensor Systems with LLMs
abstract
Despite the extensive use of sensors enabling intelligent applications, the complementary potential of co-existing sensor systems is often not fully utilized, limiting more advanced applications. This paper introduces a novel solution using Large Language Models (LLMs) to coordinate sensor systems for handling complex user queries. It defines a sensor language for sensor systems, including vocabulary set and grammar rules, analogous to natural language components, enabling LLMs to translate user intentions into sensor coordination plans. Preliminary results show that our approach significantly outperforms the existing solution at plan generation, execution and response generation stages.
Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002
SenSys10
2023 KPT: Keyword-Guided Pre-training for Grounded Dialog Generation
abstract
Incorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.r.t. different types of knowledge. In this work, we propose KPT (Keyword-guided Pre-Training), a novel self-supervised pre-training method for grounded dialog generation without relying on extra knowledge annotation. Specifically, we use a pre-trained language model to extract the most uncertain tokens in the dialog as keywords. With these keywords, we construct two kinds of knowledge and pre-train a knowledge-grounded response generation model, aiming at handling two different scenarios: (1) the knowledge should be faithfully grounded; (2) it can be selectively used. For the former, the grounding knowledge consists of keywords extracted from the response. For the latter, the grounding knowledge is additionally augmented with keywords extracted from other utterances in the same dialog. Since the knowledge is extracted from the dialog itself, KPT can be easily performed on a large volume and variety of dialogue data. We considered three data sources (open-domain, task-oriented, conversational QA) with a total of 2.5M dialogues. We conduct extensive experiments on various few-shot knowledge-grounded generation tasks, including grounding on dialog acts, knowledge graphs, persona descriptions, and Wikipedia passages. Our comprehensive experiments and analyses demonstrate that KPT consistently outperforms state-of-the-art methods on these tasks with diverse grounding knowledge.
Qi Zhu 0007, Fei Mi, Zheng Zhang 0020, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang
AAAI6
2023 Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document Understanding
abstract
Haoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001
ACL (1)12
2023 mCLIP: Multilingual CLIP via Cross-lingual Transfer
abstract
Guanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001
ACL (1)6
2023 One Cannot Stand for Everyone! Leveraging Multiple User Simulators to train Task-oriented Dialogue Systems
abstract
Yajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yajiao Liu, Xin Jiang 0002, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu 0001, Benyou Wang
ACL (1)2
2023 CAME: Confidence-guided Adaptive Memory Efficient Optimization
abstract
Adaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models.Nevertheless, the need for adaptivity requires maintaining second-moment estimates of the per-parameter gradients, which entails a high cost of extra memory overheads.To solve this problem, several memory-efficient optimizers (e.g., Adafactor) have been proposed to obtain a drastic reduction in auxiliary memory usage, but with a performance penalty.In this paper, we first study a confidence-guided strategy to reduce the instability of existing memory efficient optimizers.Based on this strategy, we propose CAME to simultaneously achieve two goals: fast convergence as in traditional adaptive methods, and low memory usage as in memory-efficient methods.Extensive experiments demonstrate the training stability and superior performance of CAME across various NLP tasks such as BERT and GPT-2 training.Notably, for BERT pre-training on the large batch size of 32,768, our proposed optimizer attains faster convergence and higher accuracy compared with the Adam optimizer.The implementation of CAME is publicly available 1 .
Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang 0002, Yang You 0001
ACL (1)5
2023 Lexicon-injected Semantic Parsing for Task-Oriented Dialog
abstract
Recently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents and slots as labels of nested tree nodes, has been proposed for parsing user utterances. Previous TOP parsing methods are limited on leveraging lexicon resources, which are often used to guide the real dialog system. To mitigate this issue, we first propose a novel span-splitting representation for span-based parser that outperforms existing methods. Then we present a novel lexicon-injected semantic parser, which collects slot labels of tree representation as a lexicon, and injects lexical features to the span representation of parser. An additional slot disambiguation technique is involved to remove inappropriate span match occurrences from the lexicon. Experiments show that our best parser produces a new state-of-the-art result (87.62%) on the TOP dataset, and also confirm the effectiveness of our proposed lexicon-injected parser and slot disambiguation model.
Xiaojun Meng, Wenlin Dai, Yasheng Wang, Baojun Wang, Zhiyong Wu 0003, Xin Jiang 0002, Qun Liu 0001
ICASSP6
2023 History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future Prompt
abstract
Dialogue history and response in open-domain dialogue are loosely coupled. Generating informative responses solely based on the original dialogue history is not easy, as dialogue history may not contain enough information or it may contain irrelevant noises. Intuitively, if a generation model can foresee possible dialogue future, or obtain real useful histories, it could generate more informative responses. In this paper, we propose a novel lightweight dialogue generation framework named few-shot history-future prompt that utilizes useful histories and simulated futures to help generate informative responses, without the need for fine-tuning or adding extra parameters. To obtain useful histories, we retrieve and combine relevant utterances from noisy multi- turn histories. Then we adopt a retrieval-generation hybrid approach to obtain diversified simulated futures. Such that our model could learn to condition on history combinations and simulated futures via few-shot learning. Experiments over publicly available datasets demonstrate that our method can help models generate better responses.
Yasheng Wang, Fei Mi, Pingyi Zhou, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001
ICASSP7
2023 A Study on Transformer Configuration and Training Objective
abstract
Transformer-based models have delivered impressive results on many tasks, particularly vision and language tasks. In many model training situations, conventional configurations are often adopted. For example, we usually set the base model with hidden size (i.e. model width) to be 768 and the number of transformer layers (i.e. model depth) to be 12. In this paper, we revisit these conventional configurations by studying the the relationship between transformer configuration and training objective. We show that the optimal transformer configuration is closely related to the training objective. Specifically, compared with the simple classification objective, the masked autoencoder is effective in alleviating the over-smoothing issue in deep transformer training. Based on this finding, we propose “Bamboo”, a notion of using deeper and narrower transformer configurations, for masked autoencoder training. On ImageNet, with such a simple change in configuration, the re-designed Base-level transformer achieves 84.2% top-1 accuracy and outperforms SoTA models like MAE by $0.9%$. On language tasks, re-designed model outperforms BERT with the default setting by 1.1 points on average, on GLUE benchmark with 8 datasets.
Fuzhao Xue, Jianghai Chen, Aixin Sun, Xiaozhe Ren, Zangwei Zheng, Xiao-Xin He, Yongming Chen, Xin Jiang 0002, Yang You 0001
ICML8
2023 Learning Summary-Worthy Visual Representation for Abstractive Summarization in Video
abstract
Multimodal abstractive summarization for videos (MAS) requires generating a concise textual summary to describe the highlights of a video according to multimodal resources, in our case, the video content and its transcript. Inspired by the success of the large-scale generative pre-trained language model (GPLM) in generating high-quality textual content (e.g., summary), recent MAS methods have proposed to adapt the GPLM to this task by equipping it with the visual information, which is often obtained through a general-purpose visual feature extractor. However, the generally extracted visual features may overlook some summary-worthy visual information, which impedes model performance. In this work, we propose a novel approach to learning the summary-worthy visual representation that facilitates abstractive summarization. Our method exploits the summary-worthy information from both the cross-modal transcript data and the knowledge that distills from the pseudo summary. Extensive experiments on three public multimodal datasets show that our method outperforms all competing baselines. Furthermore, with the advantages of summary-worthy visual information, our model can have a significant improvement on small datasets or even datasets with limited training data.
Zenan Xu, Xiaojun Meng, Yasheng Wang, Qinliang Su, Zexuan Qiu, Xin Jiang 0002, Qun Liu 0001
IJCAI6
2023 Reusing Pretrained Models by Multi-linear Operators for Efficient Training
abstract
Training large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively.
Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
NeurIPS6
2023 Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline
abstract
Large language models (LLMs) have revolutionized the field of AI, demonstrating unprecedented capacity across various tasks. However, the inference process for LLMs comes with significant computational costs. In this paper, we propose an efficient LLM inference pipeline that harnesses the power of LLMs. Our approach begins by tapping into the potential of LLMs to accurately perceive and predict the response length with minimal overhead. By leveraging this information, we introduce an efficient sequence scheduling technique that groups queries with similar response lengths into micro-batches. We evaluate our approach on real-world instruction datasets using the LLaMA-based model, and our results demonstrate an impressive 86% improvement in inference throughput without compromising effectiveness. Notably, our method is orthogonal to other inference acceleration techniques, making it a valuable addition to many existing toolkits (e.g., FlashAttention, Quantization) for LLM inference.
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Xin Jiang 0002, Yang You 0001
NeurIPS5
2023 EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge
abstract
Deep Learning (DL) models have been widely deployed on IoT devices with the help of advancements in DL algorithms and chips. However, the limited resources of edge devices make these on-device DL models hard to be generalizable to diverse environments and tasks. Although the recently emerged foundation models (FMs) show impressive generalization power, how to effectively leverage the rich knowledge of FMs on resource-limited edge devices is still not explored. In this paper, we propose EdgeFM, a novel edge-cloud cooperative system with open-set recognition capability. EdgeFM selectively uploads unlabeled data to query the FM on the cloud and customizes the specific knowledge and architectures for edge models. Meanwhile, EdgeFM conducts dynamic model switching at run-time taking into account both data uncertainty and dynamic network variations, which ensures the accuracy always close to the original FM. We implement EdgeFM using two FMs on two edge platforms. We evaluate EdgeFM on three public datasets and two self-collected datasets. Results show that EdgeFM can reduce the end-to-end latency up to 3.2x and achieve 34.3% accuracy increase compared with the baseline.
Bufang Yang, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002
SenSys8
2022 AutoBERT-Zero: Evolving BERT Backbone from Scratch
abstract
Transformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely stacking the manually designed global self-attention layers, introducing inductive bias and thus leads to sub-optimal. In this work, we make the first attempt to automatically discover novel pre-trained language model (PLM) backbone on a flexible search space containing the most fundamental operations from scratch. Specifically, we propose a well-designed search space which (i) contains primitive math operations in the intra-layer level to explore novel attention structures, and (ii) leverages convolution blocks to be the supplementary for attentions in the inter-layer level to better learn local dependency. To enhance the efficiency for finding promising architectures, we propose an Operation-Priority Neural Architecture Search (OP-NAS) algorithm, which optimizes both the search algorithm and evaluation of candidate models. Specifically, we propose Operation-Priority (OP) evolution strategy to facilitate model search via balancing exploration and exploitation. Furthermore, we design a Bi-branch Weight-Sharing (BIWS) training strategy for fast model evaluation. Extensive experiments show that the searched architecture (named AutoBERT-Zero) significantly outperforms BERT and its variants of different model capacities in various downstream tasks, proving the architecture's transfer and scaling abilities. Remarkably, AutoBERT-Zero-base outperforms RoBERTa-base (using much more data) and BERT-large (with much larger model size) by 2.4 and 1.4 higher score on GLUE test set.
Jiahui Gao 0002, Hang Xu 0004, Xiaozhe Ren, Philip L. H. Yu, Xiaodan Liang, Xin Jiang 0002, Zhenguo Li
AAAI7
2022 UniMS: A Unified Framework for Multimodal Summarization with Knowledge Distillation
abstract
With the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task.
Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang
AAAI4
2022 bert2BERT: Towards Reusable Pretrained Language Models
abstract
Cheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001
ACL (1)4
2022 Compression of Generative Pre-trained Language Models via Quantization
abstract
Chaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, Ngai Wong. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Chaofan Tao, Lu Hou 0002, Wei Zhang 0196, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Ping Luo 0002, Ngai Wong 0001
ACL (1)5
2022 ClusterFormer: Neural Clustering Attention for Efficient and Effective Transformer
abstract
Ningning Wang, Guobing Gan, Peng Zhang, Shuai Zhang, Junqiu Wei, Qun Liu, Xin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Guobing Gan, Peng Zhang 0002, Victor Junqiu Wei, Qun Liu 0001, Xin Jiang 0002
ACL (1)7
2022 Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question Answering
abstract
Jiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Jiawei Zhou 0003, Lifeng Shang, Ke Zhan, Enrui Hu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002
ACL (1)11
2022 Pan More Gold from the Sand: Refining Open-domain Dialogue Training with Noisy Self-Retrieval Generation
abstract
Real human conversation data are complicated, heterogeneous, and noisy, from which building open-domain dialogue systems remains a challenging task. In fact, such dialogue data still contains a wealth of information and knowledge, however, they are not fully explored. In this paper, we show existing open-domain dialogue generation methods that memorize context-response paired data with autoregressive or encode-decode language models underutilize the training data. Different from current approaches, using external knowledge, we explore a retrieval-generation training framework that can take advantage of the heterogeneous and noisy training data by considering them as “evidence”. In particular, we use BERTScore for retrieval, which gives better qualities of the evidence and generation. Experiments over publicly available datasets demonstrate that our method can help models generate better responses, even such training data are usually impressed as low-quality data. Such performance gain is comparable with those improved by enlarging the training set, even better. We also found that the model performance has a positive correlation with the relevance of the retrieved evidence. Moreover, our method performed well on zero-shot experiments, which indicates that our method can be more robust to real-world data.
Yasheng Wang, Fei Mi, Pingyi Zhou, Xin Wang 0114, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001
COLING8
2022 UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual Dialog
abstract
Visual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1.
Zhenshan Tan, Qingrong Cheng, Xin Jiang 0002, Qun Liu 0001, Yudong Zhu, Xiaodong Gu 0001
CVPR4
2022 LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal Modeling
abstract
Recent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks.However, the pretraining process is computationally expensive due to the requirement of millions of videotext pairs and the redundant data structure of each video.To mitigate these problems, we propose LiteVL, which adapts a pre-trained image-language model BLIP into a video-text model directly on downstream tasks, without heavy pre-training.To enhance the temporal modeling lacking in the image-language model, we propose to add temporal attention modules in the image encoder of BLIP with dynamic temporal scaling.Besides the model-wise adaptation, we also propose a non-parametric pooling mechanism to adaptively reweight the fine-grained video embedding conditioned on the text.Experimental results on text-video retrieval and video question answering show that the proposed LiteVL even outperforms previous video-language pre-trained models by a clear margin, though without any videolanguage pre-training.
Chaofan Tao, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
EMNLP5
2022 Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing
abstract
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022.
Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais
EMNLP12
2022 Pre-training Language Models with Deterministic Factual Knowledge
abstract
Previous works show that Pre-trained Language Models (PLMs) can capture factual knowledge.However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitive to the changes of prompts when extracting factual knowledge.To mitigate this issue, we propose to let PLMs learn the deterministic relationship between the remaining context and the masked content.The deterministic relationship ensures that the masked factual content can be deterministically inferable based on the existing clues in the context.That would provide more stable patterns for PLMs to capture factual knowledge than randomly masking.Two pre-training tasks are further introduced to motivate PLMs to rely on the deterministic relationship when filling masks.Specifically, we use an external Knowledge Base (KB) to identify deterministic relationships and continuously pre-train PLMs with the proposed methods.The factual knowledge probing experiments indicate that the continuously pre-trained PLMs achieve better robustness in factual knowledge capturing.Further experiments on question-answering datasets show that trying to learn a deterministic relationship with the proposed methods can also help other knowledge-intensive tasks.
Shaobo Li 0004, Lifeng Shang, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001
EMNLP7
2022 Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine Translation
abstract
Transformer has been demonstrated effective in Neural Machine Translation (NMT).However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback.To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, which retains full rank and meanwhile reduces operations and parameters.A Transformer using HTT, named Hypoformer, consistently and notably outperforms the recent lightweight SOTA methods on three standard translation tasks under different parameter and speed scales.In extreme low resource scenarios, Hypoformer has a 7.1 point absolute improvement in BLEU and 1.27× speedup than the vanilla Transformer on the IWSLT'14 De-En task.
Sunzhu Li, Peng Zhang 0002, Guobing Gan, Xiuqing Lv, Benyou Wang, Victor Junqiu Wei, Xin Jiang 0002
EMNLP7
2022 G-MAP: General Memory-Augmented Pre-trained Language Model for Domain Tasks
abstract
Recently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks.
Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001
EMNLP7
2022 SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang 0002, Qun Liu 0001
ICLR4
2022 Exploring extreme parameter compression for pre-trained language models
Benyou Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001
ICLR4
2022 FILIP: Fine-grained Interactive Language-Image Pre-Training
Lewei Yao, Runhui Huang, Lu Hou 0002, Guansong Lu, Minzhe Niu, Hang Xu 0004, Xiaodan Liang, Zhenguo Li, Xin Jiang 0002, Chunjing Xu
ICLR9
2022 Boosting Graph Structure Learning with Dummy Nodes
abstract
With the development of graph kernels and graph representation learning, many superior methods have been proposed to handle scalability and oversmoothing issues on graph structure learning. However, most of those strategies are designed based on practical experience rather than theoretical analysis. In this paper, we use a particular dummy node connecting to all existing vertices without affecting original vertex and edge properties. We further prove that such the dummy node can help build an efficient monomorphic edge-to-vertex transform and an epimorphic inverse to recover the original graph back. It also indicates that adding dummy nodes can preserve local and global structures for better graph representation learning. We extend graph kernels and graph neural networks with dummy nodes and conduct experiments on graph classification and subgraph isomorphism matching tasks. Empirical results demonstrate that taking graphs with dummy nodes as input significantly boosts graph structure learning, and using their edge-to-vertex graphs can also achieve similar results. We also discuss the gain of expressive power from the dummy in neural networks.
Xin Liu 0039, Cheng Jiayang, Yangqiu Song, Xin Jiang 0002
ICML4
2022 CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and Diagnosis
abstract
Mispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762.
Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
INTERSPEECH9
2022 Towards Efficient Post-training Quantization of Pre-trained Language Models
abstract
Network quantization has gained increasing attention with the rapid growth of large pre-trained language models~(PLMs). However, most existing quantization methods for PLMs follow quantization-aware training~(QAT) that requires end-to-end training with full access to the entire dataset. Therefore, they suffer from slow training, large memory overhead, and data accessibility issues. In this paper, we study post-training quantization~(PTQ) of PLMs, and propose module-wise quantization error minimization~(MREM), an efficient solution to mitigate these issues. By partitioning the PLM into multiple modules, we minimize the reconstruction error incurred by quantization for each module. In addition, we design a new model parallel training strategy such that each module can be trained locally on separate computing devices without waiting for preceding modules, which brings nearly the theoretical training speed-up (e.g., $4\times$ on $4$ GPUs). Experiments on GLUE and SQuAD benchmarks show that our proposed PTQ solution not only performs close to QAT, but also enjoys significant reductions in training time, memory overhead, and data consumption.
Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Irwin King, Michael R. Lyu
NeurIPS4
2022 Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark
abstract
Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_\text{ViT-L}$ achieves an average accuracy of 73.03%. For the image-text retrieval task, it achieves a mean recall of 71.6% on AIC-ICC which is 12.9% higher than WenLan 2.0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e.g., Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to https://wukong-dataset.github.io/wukong-dataset/.
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 0002, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang 0196, Xin Jiang 0002, Chunjing Xu, Hang Xu 0004
NeurIPS10
2021 HopRetriever: Retrieve Hops over Wikipedia to Answer Complex Questions
abstract
Collecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required to be gathered together to support the answer extraction. In this paper, we propose a new retrieval target, hop, to collect the hidden reasoning evidence from Wikipedia for complex question answering. Specifically, the hop in this paper is defined as the combination of a hyperlink and the corresponding outbound link document. The hyperlink is encoded as the mention embedding which models the structured knowledge of how the outbound link entity is mentioned in the textual context, and the corresponding outbound link document is encoded as the document embedding representing the unstructured knowledge within it. Accordingly, we build HopRetriever which retrieves hops over Wikipedia to answer complex questions. Experiments on the HotpotQA dataset demonstrate that HopRetriever outperforms previously published evidence retrieval methods by large margins. Moreover, our approach also yields quantifiable interpretations of the evidence collection process.
Shaobo Li 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Chengjie Sun, Zhenzhou Ji, Bingquan Liu
AAAI4
2021 Continuous Self-Attention Models with Neural ODE Networks
abstract
Stacked self-attention models receive widespread attention, due to its ability of capturing global dependency among words. However, the stacking of many layers and components generates huge parameters, leading to low parameter efficiency. In response to this issue, we propose a lightweight architecture named Continuous Self-Attention models with neural ODE networks (CSAODE). In CSAODE, continuous dynamical models (i.e., neural ODEs) are coupled with our proposed self-attention block to form a self-attention ODE solver. This solver continuously calculates and optimizes the hidden states via only one layer of parameters to improve the parameter efficiency. In addition, we design a novel accelerated continuous dynamical model to reduce computing costs, and integrate it in CSAODE. Moreover, since the original self-attention ignores local information, CSAODE makes use of N-gram convolution to encode local representations, and a fusion layer with only two trainable scalars are designed for generating sentence vectors. We perform a series of experiments on text classification, neural language inference (NLI) and text matching tasks. With fewer parameters, CSAODE outperforms state-of-the-art models on text classification tasks (e.g., 1.3% accuracy improved on SUBJ task), and has competitive performances for NLI and text matching tasks as well.
Peng Zhang 0002, Baiwen Kong, Victor Junqiu Wei, Xin Jiang 0002
AAAI5
2021 BinaryBERT: Pushing the Limit of BERT Quantization
abstract
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Haoli Bai, Wei Zhang 0196, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Michael R. Lyu, Irwin King
ACL/IJCNLP (1)6
2021 GhostBERT: Generate More Features with Cheap Operations for BERT
abstract
Zhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zhiqi Huang 0001, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ACL/IJCNLP (1)4
2021 Exploring Discourse Structures for Argument Impact Classification
abstract
Xin Liu, Jiefu Ou, Yangqiu Song, Xin Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xin Liu 0039, Jiefu Ou, Yangqiu Song, Xin Jiang 0002
ACL/IJCNLP (1)4
2021 AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models
abstract
Yichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ACL/IJCNLP (1)4
2021 EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional Fusion
abstract
This paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/.
Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee
ASRU4
2021 Improving Unsupervised Question Answering via Summarization-Informed Question Generation
abstract
Question Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach.
Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)5
2021 DyLex: Incorporating Dynamic Lexicons into BERT for Sequence Labeling
abstract
Baojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)9
2021 Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
ICANN (3)5
2021 On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 0002, Hao Yang 0006, Qun Liu 0001, Jakob Grue Simonsen
ICLR4
2021 Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma
ICLR4
2021 Improved OOD Generalization via Adversarial Training and Pretraing
abstract
Recently, learning a model that generalizes well on out-of-distribution (OOD) data has attracted great attention in the machine learning community. In this paper, after defining OOD generalization by Wasserstein distance, we theoretically justify that a model robust to input perturbation also generalizes well on OOD data. Inspired by previous findings that adversarial training helps improve robustness, we show that models trained by adversarial training have converged excess risk on OOD data. Besides, in the paradigm of pre-training then fine-tuning, we theoretically justify that the input perturbation robust model in the pre-training stage provides an initialization that generalizes well on downstream OOD data. Finally, various experiments conducted on image classification and natural language understanding tasks verify our theoretical findings.
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma
ICML5
2021 Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001
Neurocomputing5
2021 EyelashNet: a dataset and a baseline method for eyelash matting
abstract
Eyelashes play a crucial part in the human facial structure and largely affect the facial attractiveness in modern cosmetic design. However, the appearance and structure of eyelashes can easily induce severe artifacts in high-fidelity multi-view 3D face reconstruction. Unfortunately it is highly challenging to remove eyelashes from portrait images using both traditional and learning-based matting methods due to the delicate nature of eyelashes and the lack of eyelash matting dataset. To this end, we present EyelashNet, the first eyelash matting dataset which contains 5,400 high-quality eyelash matting data captured from real world and 5,272 virtual eyelash matting data created by rendering avatars. Our work consists of a capture stage and an inference stage to automatically capture and annotate eyelashes instead of tedious manual efforts. The capture is based on a specifically-designed fluorescent labeling system. By coloring the eyelashes with a safe and invisible fluorescent substance, our system takes paired photos with colored and normal eyelashes by turning the equipped ultraviolet (UVA) flash on and off. We further correct the alignment between each pair of photos and use a novel alpha matte inference network to extract the eyelash alpha matte. As there is no prior eyelash dataset, we propose a progressive training strategy that progressively fuses captured eyelash data with virtual eyelash data to learn the latent semantics of real eyelashes. As a result, our method can accurately extract eyelash alpha mattes from fuzzy and self-shadow regions such as pupils, which is almost impossible by manual annotations. To validate the advantage of EyelashNet, we present a baseline method based on deep learning that achieves state-of-the-art eyelash matting performance with RGB portrait images as input. We also demonstrate that our work can largely benefit important real applications including high-fidelity personalized avatar and cosmetic design.
Qinjie Xiao, Luyuan Wang, Xiaogang Jin 0001, Xin Jiang 0002, Tianjia Shao, Kun Zhou 0001
ACM Trans. Graph.7
2020 Dialog State Tracking with Reinforced Data Augmentation
abstract
Neural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data.
Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
AAAI3
2020 Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order
abstract
Masked language model and autoregressive language model are two types of language models.While pretrained masked language models such as BERT (Devlin et al., 2019) overwhelm the line of natural language understanding (NLU) tasks, autoregressive language models such as GPT (Radford et al., 2018) are especially capable in natural language generation (NLG).In this paper, we propose a probabilistic masking scheme for the masked language model, which we call probabilistically masked language model (PMLM).We implement a specific PMLM with a uniform prior distribution on the masking ratio named u-PMLM.We prove that u-PMLM is equivalent to an autoregressive permutated language model.One main advantage of the model is that it supports text generation in arbitrary order with surprisingly good quality, which could potentially enable new applications over traditional unidirectional generation.Besides, the pretrained u-PMLM also outperforms BERT on a set of downstream NLU tasks.
Xin Jiang 0002, Qun Liu 0001
ACL2
2020 Accurate Word Alignment Induction from Neural Machine Translation
abstract
Despite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism.In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET.The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work.SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change.SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments.Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8AER points. 1
Yun Chen 0007, Yang Liu 0005, Guanhua Chen 0001, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)4
2020 TernaryBERT: Distillation-aware Ultra-low Bit BERT
abstract
Transformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller.
Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)6
2020 Progressive Memory Banks for Incremental Domain Adaptation
Nabiha Asghar, Lili Mou, Kira A. Selby, Kevin D. Pantasdo, Pascal Poupart, Xin Jiang 0002
ICLR6
2020 On the Importance of Word and Sentence Representation Learning in Implicit Discourse Relation Classification
abstract
Implicit discourse relation classification is one of the most difficult parts in shallow discourse parsing as the relation prediction without explicit connectives requires the language understanding at both the text span level and the sentence level. Previous studies mainly focus on the interactions between two arguments. We argue that a powerful contextualized representation module, a bilateral multi-perspective matching module, and a global information fusion module are all important to implicit discourse analysis. We propose a novel model to combine these modules together. Extensive experiments show that our proposed model outperforms BERT and other state-of-the-art systems on the PDTB dataset by around 8% and CoNLL 2016 datasets around 16%. We also analyze the effectiveness of different modules in the implicit discourse relation classification task and demonstrate how different levels of representation learning can affect the results.
Xin Liu 0039, Jiefu Ou, Yangqiu Song, Xin Jiang 0002
IJCAI4
2020 An Investigation of Few-Shot Learning in Spoken Term Classification
abstract
202402 bcch
Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qing Li 0001
INTERSPEECH5
2020 Neural Subgraph Isomorphism Counting
abstract
In this paper, we study a new graph learning problem: learning to count subgraph isomorphisms. Different from other traditional graph learning problems such as node classification and link prediction, subgraph isomorphism counting is NP-complete and requires more global inference to oversee the whole graph. To make it scalable for large-scale graphs and patterns, we propose a learning framework that augments different representation learning architectures and iteratively attends pattern and target data graphs to memorize intermediate states of subgraph isomorphism searching for global counting. We develop both small graphs (<= 1,024 subgraph isomorphisms in each) and large graphs (<= 4,096 subgraph isomorphisms in each) sets to evaluate different representation and interaction modules. A mutagenic compound dataset, MUTAG, is also used to evaluate neural models and demonstrate the success of transfer learning. While the learning based approach is inexact, we are able to generalize to count large patterns and data graphs in linear time compared to the exponential time of the original NP-complete problem. Experimental results show that learning based subgraph isomorphism counting can speed up the traditional algorithm, VF2, 10-1,000 times with acceptable errors. Domain adaptation based on fine-tuning also shows the usefulness of our approach in real-world applications.
Xin Liu 0039, Haojie Pan, Mutian He 0001, Yangqiu Song, Xin Jiang 0002, Lifeng Shang
KDD5
2020 DynaBERT: Dynamic BERT with Adaptive Width and Depth
abstract
The pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compress the large BERT model to a fixed smaller size, and can not fully satisfy the requirements of different edge devices with various hardware performances. In this paper, we propose a novel dynamic BERT model (abbreviated as DynaBERT), which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. Comprehensive experiments under various efficiency constraints demonstrate that our proposed dynamic BERT (or RoBERTa) at its largest size has comparable performance as BERT-base (or RoBERTa-base), while at smaller widths and depths consistently outperforms existing BERT compression methods. Code is available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/DynaBERT.
Lu Hou 0002, Zhiqi Huang 0001, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001
NeurIPS4
2020 Unsupervised Text Generation by Learning from Search
abstract
In this work, we propose TGLS, a novel framework for unsupervised Text Generation by Learning from Search. We start by applying a strong search algorithm (in particular, simulated annealing) towards a heuristically defined objective that (roughly) estimates the quality of sentences. Then, a conditional generative model learns from the search results, and meanwhile smooth out the noise of search. The alternation between search and learning can be repeated for performance bootstrapping. We demonstrate the effectiveness of TGLS on two real-world natural language generation tasks, unsupervised paraphrasing and text formalization. Our model significantly outperforms unsupervised baseline methods in both tasks. Especially, it achieves comparable performance to strong supervised methods for paraphrase generation.
Jingjing Li 0007, Zichao Li 0001, Lili Mou, Xin Jiang 0002, Michael R. Lyu, Irwin King
NeurIPS4
2019 Decomposable Neural Paraphrase Generation
abstract
Paraphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level.This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence at different levels of granularity in a disentangled way.Specifically, the model is composed of multiple encoders and decoders with different structures, each of which corresponds to a specific granularity.The empirical study shows that the decomposition mechanism of DNPG makes paraphrase generation more interpretable and controllable.Based on DNPG, we further develop an unsupervised domain adaptation method for paraphrase generation.Experimental results show that the proposed model achieves competitive in-domain performance compared to the state-of-the-art neural models, and significantly better performance when adapting to a new domain.What is the population of New York?How many people is there in NYC?Who wrote the Winnie the Pooh books?Who is the author of winnie the pooh?What is the best phone to buy below 15k?Which are best mobile phones to buy under 15000?How can I be a good geologist?What should I do to be a great geologist?How do I reword a sentence to avoid plagiarism?How can I paraphrase my essay and avoid plagiarism?
Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001
ACL (1)2
2019 ERNIE: Enhanced Language Representation with Informative Entities
abstract
Neural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding.We argue that informative entities in KGs can enhance language representation with external knowledge.In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously.The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks.The source code and experiment details of this paper can be obtained from https:// github.com/thunlp/ERNIE.
Zhengyan Zhang, Xu Han 0007, Zhiyuan Liu 0001, Xin Jiang 0002, Maosong Sun 0001, Qun Liu 0001
ACL (1)4
2019 Exploring Diverse Expressions for Paraphrase Generation
abstract
Lihua Qian, Lin Qiu, Weinan Zhang, Xin Jiang, Yong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Lihua Qian, Weinan Zhang 0001, Xin Jiang 0002, Yong Yu 0001
EMNLP/IJCNLP (1)4
2019 Triple-to-Text: Converting RDF Triples into High-Quality Natural Languages via Optimizing an Inverse KL Divergence
abstract
Knowledge base is one of the main forms to represent information in a structured way. A knowledge base typically consists of Resource Description Frameworks (RDF) triples which describe the entities and their relations. Generating natural language description of the knowledge base is an important task in NLP, which has been formulated as a conditional language generation task and tackled using the sequence-to-sequence framework. Current works mostly train the language models by maximum likelihood estimation, which tends to generate lousy sentences. In this paper, we argue that such a problem of maximum likelihood estimation is intrinsic, which is generally irrevocable via changing network structures. Accordingly, we propose a novel Triple-to-Text (T2T) framework, which approximately optimizes the inverse Kullback-Leibler (KL) divergence between the distributions of the real and generated sentences. Due to the nature that inverse KL imposes large penalty on fake-looking samples, the proposed method can significantly reduce the probability of generating low-quality sentences. Our experiments on three real-world datasets demonstrate that T2T can generate higher-quality sentences and outperform baseline models in several evaluation metrics.
Yaoming Zhu, Juncheng Wan, Zhiming Zhou 0001, Weinan Zhang 0001, Xin Jiang 0002, Yong Yu 0001
SIGIR7
2018 Affective Neural Response Generation
Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang 0002, Lili Mou
ECIR4
2018 Paraphrase Generation with Deep Reinforcement Learning
abstract
Automatic generation of paraphrases from a given sentence is an important yet challenging task in natural language processing (NLP).In this paper, we present a deep reinforcement learning approach to paraphrase generation.Specifically, we propose a new framework for the task, which consists of a generator and an evaluator, both of which are learned from data.The generator, built as a sequenceto-sequence learning model, can produce paraphrases given a sentence.The evaluator, constructed as a deep matching model, can judge whether two sentences are paraphrases of each other.The generator is first trained by deep learning and then further fine-tuned by reinforcement learning in which the reward is given by the evaluator.For the learning of the evaluator, we propose two methods based on supervised learning and inverse reinforcement learning respectively, depending on the type of available training data.Experimental results on two datasets demonstrate the proposed models (the generators) can produce more accurate paraphrases and outperform the stateof-the-art methods in paraphrase generation in both automatic evaluation and human evaluation.
Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Hang Li 0001
EMNLP2
2016 Neural Generative Question Answering
Xin Jiang 0002, Zhengdong Lu, Lifeng Shang, Hang Li 0001
IJCAI2
2014 Ranking Optimization with Constraints
abstract
This paper addresses the problem of post-processing of ranking in search, referred to as post ranking. Although important, no research seems to have been conducted on the problem, particularly with a principled approach, and in practice ad-hoc ways of performing the task are being adopted. This paper formalizes the problem as constrained optimization in which the constraints represent the post-processing rules and the objective function represents the trade-off between adherence to the original ranking and satisfaction of the rules. The optimization amounts to refining the original ranking result based on the rules. We further propose a specific probabilistic implementation of the general formalization on the basis of the Bradley-Terry model, which is theoretically sound, effective, and efficient. Our experimental results, using benchmark datasets and enterprise search dataset, show that the proposed method works much better than several baseline methods of utilizing rules.
Fangzhao Wu, Jun Xu 0001, Hang Li 0001, Xin Jiang 0002
CIKM4
2009 A ranking approach to keyphrase extraction
abstract
This paper addresses the issue of automatically extracting keyphrases from a document. Previously, this problem was formalized as classification and learning methods for classification were utilized. This paper points out that it is more essential to cast the problem as ranking and employ a learning to rank method to perform the task. Specifically, it employs Ranking SVM, a state-of-art method of learning to rank, in keyphrase extraction. Experimental results on three datasets show that Ranking SVM significantly outperforms the baseline methods of SVM and Naive Bayes, indicating that it is better to exploit learning to rank techniques in keyphrase extraction.
Xin Jiang 0002, Yunhua Hu, Hang Li 0001
SIGIR1