EDBT 2026 Demo / reviewers in the wild / expert
Xin Jiang 0002
dblp:42/4142-2
· DBLP profile ↗
93ranked-venue papers
1as first author
74since 2021 · last 2026
0000-0002-9117-8247ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 82 · 66 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 15 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Computer networks · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool learningabstractTool learning, which allows Large Language Models (LLMs) to leverage external tools for solving complex user tasks, has emerged as a promising avenue for extending model capabilities. However, existing approaches primarily focus on data synthesis for fine-tuning LLMs to invoke tools effectively, largely ignoring how to fully stimulate the potential of the model. In this paper, we propose ToolACE-R, a novel framework that includes both model-aware iterative training and adaptive refinement for tool learning. ToolACE-R features a model-aware iterative training procedure that progressively adjust training samples based on the model’s evolving capabilities to maximize its potential. Additionally, it incorporates self-refinement training corpus which emphasizes LLM's ability to iteratively refine their tool calls, optimizing performance without requiring external feedback. Furthermore, we introduce adaptive self-refinement for efficient test-time scaling, where the trained model can autonomously determine when to stop the process based on iterative self-refinement. We conduct extensive experiments across several benchmark datasets, showing that ToolACE-R achieves competitive performance compared to advanced LLMs. The performance can be further improved efficiently through adaptive self-refinement. These results highlight the effectiveness and generalizability of ToolACE-R, offering a promising direction for more efficient and scalable tool learning. Xingshan Zeng, Weiwen Liu, Xu Huang 0008, Zezhong Wang 0004, Lingzhi Wang 0001, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Qun Liu 0001 |
AAAI | 9 |
| 2026 | Analyzing how pre-trained language models capture factual knowledge using attribution methods
Shaobo Li 0004, Chengjie Sun, Bingquan Liu, Lifeng Shang, Zhenhua Dong, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001 |
Knowl. Based Syst. | 8 |
| 2026 | An Efficient Edge-Cloud Collaboration System With Foundational Models for Open-Set IoT ApplicationsabstractArtificial intelligence (AI) models have been widely deployed on edge devices, enabling various IoT applications. However, lightweight on-device AI models on resource-limited edge devices hinder their adaptability to dynamic environments and tasks. Despite the superior generalization capabilities of recently developed Foundation Models (FMs), utilizing their extensive knowledge on the resource-constrained edge platforms remains unexplored. In this work, we introduce DeepEdgeFM, an edge-cloud collaborative system with FMs that enables open-set learning, simultaneously achieving generalizability and efficiency for IoT applications. DeepEdgeFM employs a spatiotemporalaware semantic customization approach that leverages spatial, temporal, and domain-specific knowledge from FMs to continuously customize edge models using unlabeled sensor data in emerging IoT environments. Meanwhile, DeepEdgeFM utilizes a dynamic model switching strategy to selectively query the knowledge of FMs based on sensor-data uncertainty and real-time network fluctuations. We implement DeepEdgeFM on five FMs and multi-modal large language models (MLLMs), covering four types of sensor data modalities. We evaluate DeepEdgeFM on two edge platforms, five public datasets, and two self-collected datasets covering both indoor and outdoor real-world environments. The results show that DeepEdgeFM outperforms state-ofthe- art baselines, achieving up to an 18.6% accuracy gain and a 38.6 Bufang Yang, Wenrui Lu, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
IEEE Trans. Mob. Comput. | 9 |
| 2025 | Mixture of insighTful Experts (MoTE): The Synergy of Reasoning Chains and Expert Mixtures in Self-AlignmentabstractZhili Liu, Yunhao Gou, Kai Chen, Lanqing Hong, Jiahui Gao, Fei Mi, Yu Zhang, Zhenguo Li, Xin Jiang, Qun Liu, James Kwok. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhili Liu, Yunhao Gou, Kai Chen 0023, Lanqing Hong, Jiahui Gao 0002, Fei Mi, Yu Zhang 0006, Zhenguo Li, Xin Jiang 0002, Qun Liu 0001, James T. Kwok |
ACL (1) | 9 |
| 2025 | Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editingabstractKaishuai Xu, Tiezheng Yu, Wenjun Hou, Yi Cheng, Chak Tou Leong, Liangyou Li, Xin Jiang, Lifeng Shang, Qun Liu, Wenjie Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kaishuai Xu, Tiezheng Yu, Chak Tou Leong, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001, Wenjie Li 0002 |
ACL (1) | 7 |
| 2025 | Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-JudgeabstractQiyuan Zhang, Yufei Wang, Yuxin Jiang, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Qiyuan Zhang 0001, Yufei Wang 0005, Liangyou Li, Chuhan Wu, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001 |
ACL (1) | 7 |
| 2025 | Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' ReasoningabstractZezhong Wang, Xingshan Zeng, Weiwen Liu, Yufei Wang, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Yufei Wang 0005, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
EMNLP | 8 |
| 2025 | Bridging and Modeling Correlations in Pairwise Data for Direct Preference OptimizationabstractDirect preference optimization (DPO), a widely adopted offline preference optimization algorithm, aims to align large language models (LLMs) with human-desired behaviors using pairwise preference data. However, the generation of the winning response and the losing response within pairwise data are typically isolated, leading to weak correlations between them as well as suboptimal alignment performance. To address this issue, we propose an effective framework for Bridging and Modeling Correlations in pairwise data, named BMC. Firstly, we increase the consistency and informativeness of the pairwise preference signals through targeted modifications, synthesizing a pseudo-winning response by improving the losing response with the winning response as a reference. Secondly, we identify that DPO alone is insufficient to model these correlations and capture nuanced variations. Therefore, we propose learning token-level correlations by dynamically leveraging the policy model's confidence during training. Comprehensive experiments on QA, math, and instruction-following tasks demonstrate the effectiveness of our approach, significantly surpassing competitive baselines, including DPO. Additionally, our in-depth quantitative analysis reveals the reasons behind our method's superior performance over DPO and showcases its versatility to other DPO variants. Bo Huang 0017, Yufei Wang 0005, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Wei Wang 0011 |
ICLR | 7 |
| 2025 | Forewarned is Forearmed: Harnessing LLMs for Data Synthesis via Failure-induced ExplorationabstractLarge language models (LLMs) have significantly benefited from training on diverse, high-quality task-specific data, leading to impressive performance across a range of downstream applications. Current methods often rely on human-annotated data or predefined task templates to direct powerful LLMs in synthesizing task-relevant data for effective model training. However, this dependence on manually designed components may constrain the scope of generated data, potentially overlooking critical edge cases or novel scenarios that could challenge the model. In this paper, we present a novel approach, ReverseGen, designed to automatically generate effective training samples that expose the weaknesses of LLMs. Specifically, we introduce a dedicated proposer trained to produce queries that lead target models to generate unsatisfactory responses. These failure-inducing queries are then used to construct training data, helping to address the models' shortcomings and improve overall performance. Our approach is flexible and can be applied to models of various scales (3B, 7B, and 8B). We evaluate ReverseGen on three key applications—safety, honesty, and math—demonstrating that our generated data is both highly effective and diverse. Models fine-tuned with ReverseGen-generated data consistently outperform those trained on human-annotated or general model-generated data, offering a new perspective on data synthesis for task-specific LLM enhancement. Qintong Li, Jiahui Gao 0002, Renjie Pi, Xueliang Zhao, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong |
ICLR | 7 |
| 2025 | ToolACE: Winning the Points of LLM Function CallingabstractFunction calling significantly extends the application boundary of large language models (LLMs), where high-quality and diverse training data is critical for unlocking this capability. However, collecting and annotating real function-calling data is challenging, while synthetic data from existing pipelines often lack coverage and accuracy. In this paper, we present ToolACE, an automatic agentic pipeline designed to generate accurate, complex, and diverse tool-learning data, specifically tailored to the capabilities of LLMs. ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs. Dialogs are further generated through the interplay among multiple agents, under the guidance of a complexity evaluator. To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks. We demonstrate that models trained on our synthesized data---even with only 8B parameters---achieve state-of-the-art performance, comparable to the latest GPT-4 models. Our model and a subset of the data are publicly available at https://huggingface.co/Team-ACE. Weiwen Liu, Xu Huang 0008, Xingshan Zeng, Xinlong Hao, Dexun Li, Shuai Wang 0020, Weinan Gan, Zhengying Liu, Yuanqing Yu, Zezhong Wang 0004, Yuxian Wang, Wu Ning, Yutai Hou, Bin Wang 0004, Chuhan Wu, Yong Liu 0020, Yasheng Wang, Duyu Tang, Dandan Tu, Lifeng Shang, Xin Jiang 0002, Ruiming Tang, Defu Lian, Qun Liu 0001, Enhong Chen |
ICLR | 23 |
| 2025 | Beyond Autoregression: Discrete Diffusion for Complex Reasoning and PlanningabstractAutoregressive language models, despite their impressive capabilities, struggle with complex reasoning and long-term planning tasks. We introduce discrete diffusion models as a novel solution to these challenges. Through the lens of subgoal imbalance, we demonstrate how diffusion models effectively learn difficult subgoals that elude autoregressive approaches. We propose Multi-Granularity Diffusion Modeling (MGDM), which prioritizes subgoals based on difficulty during learning. On complex tasks like Countdown, Sudoku, and Boolean Satisfiability Problems, MGDM significantly outperforms autoregressive models without using search techniques. For instance, MGDM achieves 91.5\% and 100\% accuracy on Countdown and Sudoku, respectively, compared to 45.8\% and 20.7\% for autoregressive models. Our work highlights the potential of diffusion-based approaches in advancing AI capabilities for sophisticated language understanding and problem-solving tasks. All associated codes are available at \href{https://github.com/HKUNLP/diffusion-vs-ar}{https://github.com/HKUNLP/diffusion-vs-ar}. Jiacheng Ye, Jiahui Gao 0002, Shansan Gong, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong |
ICLR | 5 |
| 2025 | Implicit Search via Discrete Diffusion: A Study on ChessabstractIn the post-AlphaGo era, there has been a renewed interest in search techniques such as Monte Carlo Tree Search (MCTS), particularly in their application to Large Language Models (LLMs).
This renewed attention is driven by the recognition that current next-token prediction models often lack the ability for long-term planning. Is it possible to instill search-like abilities within the models to enhance their planning abilities without relying on explicit search? We propose DiffuSearch , a model that does \textit{implicit search} by looking into the future world via discrete diffusion modeling. We instantiate DiffuSearch on a classical board game, Chess, where explicit search is known to be essential. Through extensive controlled experiments, we show DiffuSearch outperforms both the searchless and explicit search-enhanced policies. Specifically, DiffuSearch outperforms the one-step policy by 19.2\% and the MCTS-enhanced policy by 14\% on action accuracy. Furthermore, DiffuSearch demonstrates a notable 30\% enhancement in puzzle-solving abilities compared to explicit search-based policies, along with a significant 540 Elo increase in game-playing strength assessment. These results indicate that implicit search via discrete diffusion is a viable alternative to explicit search over a one-step policy. All codes are publicly available at \href{https://github.com/HKUNLP/DiffuSearch}{https://github.com/HKUNLP/DiffuSearch}. Jiacheng Ye, Jiahui Gao 0002, Zhiyong Wu 0003, Xin Jiang 0002, Zhenguo Li, Lingpeng Kong |
ICLR | 5 |
| 2025 | RevisEval: Improving LLM-as-a-Judge via Response-Adapted ReferencesabstractWith significant efforts in recent studies, LLM-as-a-Judge has become a cost-effective alternative to human evaluation for assessing text generation quality in a wide range of tasks. However, there still remains a reliability gap between LLM-as-a-Judge and human evaluation. One important reason is the lack of guided oracles in the evaluation process. Motivated by the role of reference pervasively used in classic text evaluation, we introduce RevisEval, a novel text generation evaluation paradigm via the response-adapted references. RevisEval is driven by the key observation that an ideal reference should maintain the necessary relevance to the response to be evaluated. Specifically, RevisEval leverages the text revision capabilities of large language models (LLMs) to adaptively revise the response, then treat the revised text as the reference (response-adapted reference) for the subsequent evaluation. Extensive experiments demonstrate that RevisEval outperforms traditional reference-free and reference-based evaluation paradigms that use LLM-as-a-Judge across NLG tasks and open-ended instruction-following tasks. More importantly, our response-adapted references can further boost the classical text metrics, e.g., BLEU and BERTScore, compared to traditional references and even rival the LLM-as-a-Judge. A detailed analysis is also conducted to confirm RevisEval's effectiveness in bias reduction, the impact of inference cost, and reference relevance. Qiyuan Zhang 0001, Yufei Wang 0005, Tiezheng Yu, Chuhan Wu, Liangyou Li, Yasheng Wang, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Fuyuan Lyu, Chen Ma 0001 |
ICLR | 8 |
| 2025 | FlatQuant: Flatness Matters for LLM QuantizationabstractRecently, quantization has been widely used for the compression and acceleration of large language models (LLMs). Due to the outliers in LLMs, it is crucial to flatten weights and activations to minimize quantization error with equally spaced quantization points. Prior research explores various pre-quantization transformations to suppress outliers, such as per-channel scaling and Hadamard transformation. However, we observe that these transformed weights and activations can still exhibit steep and dispersed distributions. In this paper, we propose FlatQuant (Fast and Learnable Affine Transformation), a new post-training quantization approach that enhances the flatness of weights and activations. Our approach identifies optimal affine transformations for each linear layer, calibrated in hours via a lightweight objective. To reduce runtime overhead of affine transformation, we apply Kronecker product with two lightweight matrices, and fuse all operations in FlatQuant into a single kernel. Extensive experiments demonstrate that FlatQuant establishes a new state-of-the-art benchmark for quantization. For example, it achieves less than 1\% accuracy drop for W4A4 quantization on the LLaMA-3-70B model, surpassing SpinQuant by 7.5\%. Additionally, it provides up to 2.3x prefill speedup and 1.7x decoding speedup compared to the FP16 model. Code is available at: https://github.com/ruikangliu/FlatQuant. Ruikang Liu, Haoli Bai, Yuening Li, Xianzhi Yu, Lu Hou 0002, Chun Yuan 0003, Xin Jiang 0002, Wulong Liu |
ICML | 11 |
| 2025 | ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue SynthesisabstractZezhong Wang, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Zezhong Wang 0004, Xingshan Zeng, Weiwen Liu, Liangyou Li, Yasheng Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
NAACL (Long Papers) | 7 |
| 2025 | TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMsabstractAn increasing number of environments, such as smart homes and factories, are being equipped with multiple sensor systems to enable diverse intelligent applications. However, most existing sensor coordination systems require manually predefined rules, limiting their ability to handle flexible and complex tasks. While recent approaches leverage large language models (LLMs) to interact with external APIs, they struggle to fully understand the capabilities and data dependencies of practical sensor systems. This paper introduces TaskSense, a novel system that coordinates multiple sensor systems in response to users' complex queries. TaskSense introduces a sensor language that automatically translates the capabilities and data dependencies of sensor systems into vocabularies and grammar rules that can be understood by LLMs. It then interprets user intentions into executable task plans for sensor systems using this sensor language in combination with LLMs. Meanwhile, TaskSense checks the solvability of user queries and verifies the correctness of task plan dependencies. To further enhance robustness, TaskSense incorporates a dynamic plan execution mechanism that adjusts plans based on real-time feedback from sensor data availability, data quality and execution results. TaskSense is deployed on real-world smart home systems, utilizing six popular LLMs. The system is evaluated across 4 scenarios involving 9 types of sensor systems, over 60 APIs, 170 tasks and 5 types of data modalities. Results show that TaskSense achieves up to 2× higher planning accuracy and a 75% increase in answer accuracy using the similar amount of tokens compared with baseline approaches. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 8 |
| 2025 | Enhancing inter-sentence coherence of extractive summarization with multitask learning
Renlong Jie, Xiaojun Meng, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
J. Intell. Inf. Syst. | 4 |
| 2024 | Preparing Lessons for Progressive Training on Language ModelsabstractThe rapid progress of Transformers in artificial intelligence has come at the cost of increased resource consumption and greenhouse gas emissions due to growing model sizes. Prior work suggests using pretrained small models to improve training efficiency, but this approach may not be suitable for new model structures. On the other hand, training from scratch can be slow, and progressively stacking layers often fails to achieve significant acceleration. To address these challenges, we propose a novel method called Apollo, which prepares lessons for expanding operations by learning high-layer functionality during training of low layers. Our approach involves low-value-prioritized sampling (LVPS) to train different depths and weight sharing to facilitate efficient expansion. We also introduce an interpolation method for stable model depth extension. Experiments demonstrate that Apollo achieves state-of-the-art acceleration ratios, even rivaling methods using pretrained models, making it a universal and efficient solution for training deep models while reducing time, financial, and environmental costs. Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Jiaxin Shi, Zenglin Xu, Ming Zhang 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 8 |
| 2024 | Unsupervised Extractive Summarization with Learnable Length Control StrategiesabstractUnsupervised extractive summarization is an important technique in information extraction and retrieval. Compared with supervised method, it does not require high-quality human-labelled summaries for training and thus can be easily applied for documents with different types, domains or languages. Most of existing unsupervised methods including TextRank and PACSUM rely on graph-based ranking on sentence centrality. However, this scorer can not be directly applied in end-to-end training, and the positional-related prior assumption is often needed for achieving good summaries. In addition, less attention is paid to length-controllable extractor, where users can decide to summarize texts under particular length constraint. This paper introduces an unsupervised extractive summarization model based on a siamese network, for which we develop a trainable bidirectional prediction objective between the selected summary and the original document. Different from the centrality-based ranking methods, our extractive scorer can be trained in an end-to-end manner, with no other requirement of positional assumption. In addition, we introduce a differentiable length control module by approximating 0-1 knapsack solver for end-to-end length-controllable extracting. Experiments show that our unsupervised method largely outperforms the centrality-based baseline using a same sentence encoder. In terms of length control ability, via our trainable knapsack module, the performance consistently outperforms the strong baseline without utilizing end-to-end training. Human evaluation further evidences that our method performs the best among baselines in terms of relevance and consistency. Renlong Jie, Xiaojun Meng, Xin Jiang 0002, Qun Liu 0001 |
AAAI | 3 |
| 2024 | FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language ModelsabstractYuxin Jiang, Yufei Wang, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufei Wang 0005, Xingshan Zeng, Wanjun Zhong, Liangyou Li, Fei Mi, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Wei Wang 0011 |
ACL (1) | 8 |
| 2024 | Learning to Edit: Aligning LLMs with Knowledge EditingabstractYuxin Jiang, Yufei Wang, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao, Liangyou Li, Xin Jiang, Lifeng Shang, Ruiming Tang, Qun Liu, Wei Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufei Wang 0005, Chuhan Wu, Wanjun Zhong, Xingshan Zeng, Jiahui Gao 0002, Liangyou Li, Xin Jiang 0002, Lifeng Shang, Ruiming Tang, Qun Liu 0001, Wei Wang 0011 |
ACL (1) | 8 |
| 2024 | MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language ModelsabstractWai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang, Liangyou Li, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wai-Chung Kwan, Xingshan Zeng, Yufei Wang 0005, Liangyou Li, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
EMNLP | 7 |
| 2024 | Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisabstractThe rapid development of large language models (LLMs) has not only provided numerous opportunities but also presented significant challenges. This becomes particularly evident when LLMs inadvertently generate harmful or toxic content, either unintentionally or because of intentional inducement. Existing alignment methods usually direct LLMs toward the favorable outcomes by utilizing human-annotated, flawless instruction-response pairs. Conversely, this study proposes a novel alignment technique based on mistake analysis, which deliberately exposes LLMs to erroneous content to learn the reasons for mistakes and how to avoid them. In this case, mistakes are repurposed into valuable data for alignment, effectively helping to avoid the production of erroneous responses. Without external models or human annotations, our method leverages a model's intrinsic ability to discern undesirable mistakes and improves the safety of its generated responses. Experimental results reveal that our method outperforms existing alignment approaches in enhancing model safety while maintaining the overall utility. Kai Chen 0023, Chunwei Wang, Jianhua Han, Lanqing Hong, Fei Mi, Hang Xu 0004, Zhengying Liu, Wenyong Huang, Zhenguo Li, Dit-Yan Yeung, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 13 |
| 2024 | Retrieval-based Disentangled Representation Learning with Natural Language SupervisionabstractDisentangled representation learning remains challenging as the underlying factors of variation in the data do not naturally exist. The inherent complexity of real-world data makes it unfeasible to exhaustively enumerate and encapsulate all its variations within a finite set of factors. However, it is worth noting that most real-world data have linguistic equivalents, typically in the form of textual descriptions. These linguistic counterparts can represent the data and effortlessly decomposed into distinct tokens. In light of this, we present Vocabulary Disentangled Retrieval (VDR), a retrieval-based framework that harnesses natural language as proxies of the underlying data variation to drive disentangled representation learning. Our approach employ a bi-encoder model to represent both data and natural language in a vocabulary space, enabling the model to distinguish dimensions that capture intrinsic characteristics within data through its natural language counterpart, thus facilitating disentanglement. We extensively assess the performance of VDR across 15 retrieval benchmark datasets, covering text-to-text and cross-modal retrieval scenarios, as well as human evaluation. Our experimental results compellingly demonstrate the superiority of VDR over previous bi-encoder retrievers with comparable model size and training costs, achieving an impressive 8.7% improvement in NDCG@10 on the BEIR benchmark, a 5.3\% increase on MS COCO, and a 6.0% increase on Flickr30k in terms of mean recall in the zero-shot setting. Moreover, The results from human evaluation indicate that interpretability of our method is on par with SOTA captioning models. Jiawei Zhou 0003, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002 |
ICLR | 4 |
| 2024 | Visually Guided Generative Text-Layout Pre-training for Document IntelligenceabstractZhiming Mao, Haoli Bai, Lu Hou, Lifeng Shang, Xin Jiang, Qun Liu, Kam-Fai Wong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Zhiming Mao, Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Kam-Fai Wong |
NAACL-HLT | 5 |
| 2024 | Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language ModelsabstractRecently, diffusion models have garnered significant interest in the field of text processing due to their many potential advantages compared to conventional autoregressive models.
In this work, we propose Diffusion-of-Thought (DoT), a novel approach that integrates diffusion models with Chain-of-Thought, a well-established technique for improving the reasoning ability of autoregressive language models. In contrast to autoregressive language models that make decisions in a left-to-right, token-by-token manner, DoT allows reasoning steps to diffuse over time through a diffusion language model and offers greater flexibility in trading-off computation for reasoning performance. Our experimental results demonstrate the effectiveness of DoT in multi-digit multiplication, boolean logic, and grade school math problems. In addition to that, DoT showcases promising self-correction abilities and benefits from existing reasoning-enhancing techniques like self-consistency decoding. Our findings contribute to the understanding and development of reasoning with diffusion language models. Jiacheng Ye, Shansan Gong, Jiahui Gao 0002, Xin Jiang 0002, Zhenguo Li, Wei Bi, Lingpeng Kong |
NeurIPS | 8 |
| 2024 | Poster Abstract: Tasking Heterogeneous Sensor Systems with LLMsabstractDespite the extensive use of sensors enabling intelligent applications, the complementary potential of co-existing sensor systems is often not fully utilized, limiting more advanced applications. This paper introduces a novel solution using Large Language Models (LLMs) to coordinate sensor systems for handling complex user queries. It defines a sensor language for sensor systems, including vocabulary set and grammar rules, analogous to natural language components, enabling LLMs to translate user intentions into sensor coordination plans. Preliminary results show that our approach significantly outperforms the existing solution at plan generation, execution and response generation stages. Kaiwei Liu 0001, Bufang Yang, Lilin Xu, Yunqi Guo, Neiwen Ling, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002, Zhenyu Yan 0002 |
SenSys | 10 |
| 2023 | KPT: Keyword-Guided Pre-training for Grounded Dialog GenerationabstractIncorporating external knowledge into the response generation process is essential to building more helpful and reliable dialog agents. However, collecting knowledge-grounded conversations is often costly, calling for a better pre-trained model for grounded dialog generation that generalizes well w.r.t. different types of knowledge. In this work, we propose KPT (Keyword-guided Pre-Training), a novel self-supervised pre-training method for grounded dialog generation without relying on extra knowledge annotation. Specifically, we use a pre-trained language model to extract the most uncertain tokens in the dialog as keywords. With these keywords, we construct two kinds of knowledge and pre-train a knowledge-grounded response generation model, aiming at handling two different scenarios: (1) the knowledge should be faithfully grounded; (2) it can be selectively used. For the former, the grounding knowledge consists of keywords extracted from the response. For the latter, the grounding knowledge is additionally augmented with keywords extracted from other utterances in the same dialog. Since the knowledge is extracted from the dialog itself, KPT can be easily performed on a large volume and variety of dialogue data. We considered three data sources (open-domain, task-oriented, conversational QA) with a total of 2.5M dialogues. We conduct extensive experiments on various few-shot knowledge-grounded generation tasks, including grounding on dialog acts, knowledge graphs, persona descriptions, and Wikipedia passages. Our comprehensive experiments and analyses demonstrate that KPT consistently outperforms state-of-the-art methods on these tasks with diverse grounding knowledge. Qi Zhu 0007, Fei Mi, Zheng Zhang 0020, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Xiaoyan Zhu 0001, Minlie Huang |
AAAI | 6 |
| 2023 | Wukong-Reader: Multi-modal Pre-training for Fine-grained Visual Document UnderstandingabstractHaoli Bai, Zhiguang Liu, Xiaojun Meng, Li Wentao, Shuang Liu, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang, Lu Hou, Jiansheng Wei, Xin Jiang, Qun Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Haoli Bai, Xiaojun Meng, Yifeng Luo, Nian Xie, Rongfu Zheng, Liangwei Wang 0004, Lu Hou 0002, Jiansheng Wei, Xin Jiang 0002, Qun Liu 0001 |
ACL (1) | 12 |
| 2023 | mCLIP: Multilingual CLIP via Cross-lingual TransferabstractGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai, Lifeng Shang, Xin Jiang, Qun Liu, Jia Pan, Wenping Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Guanhua Chen 0001, Lu Hou 0002, Yun Chen 0007, Wenliang Dai, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Jia Pan 0001, Wenping Wang 0001 |
ACL (1) | 6 |
| 2023 | One Cannot Stand for Everyone! Leveraging Multiple User Simulators to train Task-oriented Dialogue SystemsabstractYajiao Liu, Xin Jiang, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu, Xiang Wan, Benyou Wang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yajiao Liu, Xin Jiang 0002, Yichun Yin, Yasheng Wang, Fei Mi, Qun Liu 0001, Benyou Wang |
ACL (1) | 2 |
| 2023 | CAME: Confidence-guided Adaptive Memory Efficient OptimizationabstractAdaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models.Nevertheless, the need for adaptivity requires maintaining second-moment estimates of the per-parameter gradients, which entails a high cost of extra memory overheads.To solve this problem, several memory-efficient optimizers (e.g., Adafactor) have been proposed to obtain a drastic reduction in auxiliary memory usage, but with a performance penalty.In this paper, we first study a confidence-guided strategy to reduce the instability of existing memory efficient optimizers.Based on this strategy, we propose CAME to simultaneously achieve two goals: fast convergence as in traditional adaptive methods, and low memory usage as in memory-efficient methods.Extensive experiments demonstrate the training stability and superior performance of CAME across various NLP tasks such as BERT and GPT-2 training.Notably, for BERT pre-training on the large batch size of 32,768, our proposed optimizer attains faster convergence and higher accuracy compared with the Adam optimizer.The implementation of CAME is publicly available 1 . Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang 0002, Yang You 0001 |
ACL (1) | 5 |
| 2023 | Lexicon-injected Semantic Parsing for Task-Oriented DialogabstractRecently, semantic parsing using hierarchical representations for dialog systems has captured substantial attention. Task-Oriented Parse (TOP), a tree representation with intents and slots as labels of nested tree nodes, has been proposed for parsing user utterances. Previous TOP parsing methods are limited on leveraging lexicon resources, which are often used to guide the real dialog system. To mitigate this issue, we first propose a novel span-splitting representation for span-based parser that outperforms existing methods. Then we present a novel lexicon-injected semantic parser, which collects slot labels of tree representation as a lexicon, and injects lexical features to the span representation of parser. An additional slot disambiguation technique is involved to remove inappropriate span match occurrences from the lexicon. Experiments show that our best parser produces a new state-of-the-art result (87.62%) on the TOP dataset, and also confirm the effectiveness of our proposed lexicon-injected parser and slot disambiguation model. Xiaojun Meng, Wenlin Dai, Yasheng Wang, Baojun Wang, Zhiyong Wu 0003, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 6 |
| 2023 | History, Present and Future: Enhancing Dialogue Generation with Few-Shot History-Future PromptabstractDialogue history and response in open-domain dialogue are loosely coupled. Generating informative responses solely based on the original dialogue history is not easy, as dialogue history may not contain enough information or it may contain irrelevant noises. Intuitively, if a generation model can foresee possible dialogue future, or obtain real useful histories, it could generate more informative responses. In this paper, we propose a novel lightweight dialogue generation framework named few-shot history-future prompt that utilizes useful histories and simulated futures to help generate informative responses, without the need for fine-tuning or adding extra parameters. To obtain useful histories, we retrieve and combine relevant utterances from noisy multi- turn histories. Then we adopt a retrieval-generation hybrid approach to obtain diversified simulated futures. Such that our model could learn to condition on history combinations and simulated futures via few-shot learning. Experiments over publicly available datasets demonstrate that our method can help models generate better responses. Yasheng Wang, Fei Mi, Pingyi Zhou, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
ICASSP | 7 |
| 2023 | A Study on Transformer Configuration and Training ObjectiveabstractTransformer-based models have delivered impressive results on many tasks, particularly vision and language tasks. In many model training situations, conventional configurations are often adopted. For example, we usually set the base model with hidden size (i.e. model width) to be 768 and the number of transformer layers (i.e. model depth) to be 12. In this paper, we revisit these conventional configurations by studying the the relationship between transformer configuration and training objective. We show that the optimal transformer configuration is closely related to the training objective. Specifically, compared with the simple classification objective, the masked autoencoder is effective in alleviating the over-smoothing issue in deep transformer training. Based on this finding, we propose “Bamboo”, a notion of using deeper and narrower transformer configurations, for masked autoencoder training. On ImageNet, with such a simple change in configuration, the re-designed Base-level transformer achieves 84.2% top-1 accuracy and outperforms SoTA models like MAE by $0.9%$. On language tasks, re-designed model outperforms BERT with the default setting by 1.1 points on average, on GLUE benchmark with 8 datasets. Fuzhao Xue, Jianghai Chen, Aixin Sun, Xiaozhe Ren, Zangwei Zheng, Xiao-Xin He, Yongming Chen, Xin Jiang 0002, Yang You 0001 |
ICML | 8 |
| 2023 | Learning Summary-Worthy Visual Representation for Abstractive Summarization in VideoabstractMultimodal abstractive summarization for videos (MAS) requires generating a concise textual summary to describe the highlights of a video according to multimodal resources, in our case, the video content and its transcript. Inspired by the success of the large-scale generative pre-trained language model (GPLM) in generating high-quality textual content (e.g., summary), recent MAS methods have proposed to adapt the GPLM to this task by equipping it with the visual information, which is often obtained through a general-purpose visual feature extractor. However, the generally extracted visual features may overlook some summary-worthy visual information, which impedes model performance. In this work, we propose a novel approach to learning the summary-worthy visual representation that facilitates abstractive summarization. Our method exploits the summary-worthy information from both the cross-modal transcript data and the knowledge that distills from the pseudo summary. Extensive experiments on three public multimodal datasets show that our method outperforms all competing baselines. Furthermore, with the advantages of summary-worthy visual information, our model can have a significant improvement on small datasets or even datasets with limited training data. Zenan Xu, Xiaojun Meng, Yasheng Wang, Qinliang Su, Zexuan Qiu, Xin Jiang 0002, Qun Liu 0001 |
IJCAI | 6 |
| 2023 | Reusing Pretrained Models by Multi-linear Operators for Efficient TrainingabstractTraining large models from scratch usually costs a substantial amount of resources. Towards this problem, recent studies such as bert2BERT and LiGO have reused small pretrained models to initialize a large model (termed the ``target model''), leading to a considerable acceleration in training. Despite the successes of these previous studies, they grew pretrained models by mapping partial weights only, ignoring potential correlations across the entire model. As we show in this paper, there are inter- and intra-interactions among the weights of both the pretrained and the target models. As a result, the partial mapping may not capture the complete information and lead to inadequate growth. In this paper, we propose a method that linearly correlates each weight of the target model to all the weights of the pretrained model to further enhance acceleration ability. We utilize multi-linear operators to reduce computational and spacial complexity, enabling acceptable resource requirements. Experiments demonstrate that our method can save 76\% computational costs on DeiT-base transferred from DeiT-small, which outperforms bert2BERT by +12\% and LiGO by +21\%, respectively. Yu Pan 0005, Ye Yuan 0016, Yichun Yin, Zenglin Xu, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
NeurIPS | 6 |
| 2023 | Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineabstractLarge language models (LLMs) have revolutionized the field of AI, demonstrating unprecedented capacity across various tasks. However, the inference process for LLMs comes with significant computational costs. In this paper, we propose an efficient LLM inference pipeline that harnesses the power of LLMs. Our approach begins by tapping into the potential of LLMs to accurately perceive and predict the response length with minimal overhead. By leveraging this information, we introduce an efficient sequence scheduling technique that groups queries with similar response lengths into micro-batches. We evaluate our approach on real-world instruction datasets using the LLaMA-based model, and our results demonstrate an impressive 86% improvement in inference throughput without compromising effectiveness. Notably, our method is orthogonal to other inference acceleration techniques, making it a valuable addition to many existing toolkits (e.g., FlashAttention, Quantization) for LLM inference. Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Xin Jiang 0002, Yang You 0001 |
NeurIPS | 5 |
| 2023 | EdgeFM: Leveraging Foundation Model for Open-set Learning on the EdgeabstractDeep Learning (DL) models have been widely deployed on IoT devices with the help of advancements in DL algorithms and chips. However, the limited resources of edge devices make these on-device DL models hard to be generalizable to diverse environments and tasks. Although the recently emerged foundation models (FMs) show impressive generalization power, how to effectively leverage the rich knowledge of FMs on resource-limited edge devices is still not explored. In this paper, we propose EdgeFM, a novel edge-cloud cooperative system with open-set recognition capability. EdgeFM selectively uploads unlabeled data to query the FM on the cloud and customizes the specific knowledge and architectures for edge models. Meanwhile, EdgeFM conducts dynamic model switching at run-time taking into account both data uncertainty and dynamic network variations, which ensures the accuracy always close to the original FM. We implement EdgeFM using two FMs on two edge platforms. We evaluate EdgeFM on three public datasets and two self-collected datasets. Results show that EdgeFM can reduce the end-to-end latency up to 3.2x and achieve 34.3% accuracy increase compared with the baseline. Bufang Yang, Lixing He, Neiwen Ling, Zhenyu Yan 0002, Guoliang Xing, Xian Shuai, Xiaozhe Ren, Xin Jiang 0002 |
SenSys | 8 |
| 2022 | AutoBERT-Zero: Evolving BERT Backbone from ScratchabstractTransformer-based pre-trained language models like BERT and its variants have recently achieved promising performance in various natural language processing (NLP) tasks. However, the conventional paradigm constructs the backbone by purely stacking the manually designed global self-attention layers, introducing inductive bias and thus leads to sub-optimal. In this work, we make the first attempt to automatically discover novel pre-trained language model (PLM) backbone on a flexible search space containing the most fundamental operations from scratch. Specifically, we propose a well-designed search space which (i) contains primitive math operations in the intra-layer level to explore novel attention structures, and (ii) leverages convolution blocks to be the supplementary for attentions in the inter-layer level to better learn local dependency. To enhance the efficiency for finding promising architectures, we propose an Operation-Priority Neural Architecture Search (OP-NAS) algorithm, which optimizes both the search algorithm and evaluation of candidate models. Specifically, we propose Operation-Priority (OP) evolution strategy to facilitate model search via balancing exploration and exploitation. Furthermore, we design a Bi-branch Weight-Sharing (BIWS) training strategy for fast model evaluation. Extensive experiments show that the searched architecture (named AutoBERT-Zero) significantly outperforms BERT and its variants of different model capacities in various downstream tasks, proving the architecture's transfer and scaling abilities. Remarkably, AutoBERT-Zero-base outperforms RoBERTa-base (using much more data) and BERT-large (with much larger model size) by 2.4 and 1.4 higher score on GLUE test set. Jiahui Gao 0002, Hang Xu 0004, Xiaozhe Ren, Philip L. H. Yu, Xiaodan Liang, Xin Jiang 0002, Zhenguo Li |
AAAI | 7 |
| 2022 | UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationabstractWith the rapid increase of multimedia data, a large body of literature has emerged to work on multimodal summarization, the majority of which target at refining salient information from textual and image modalities to output a pictorial summary with the most relevant images. Existing methods mostly focus on either extractive or abstractive summarization and rely on the presence and quality of image captions to build image references. We are the first to propose a Unified framework for Multimodal Summarization grounding on BART, UniMS, that integrates extractive and abstractive objectives, as well as selecting the image output. Specially, we adopt knowledge distillation from a vision-language pretrained model to improve image selection, which avoids any requirement on the existence and quality of image captions. Besides, we introduce a visual guided decoder to better integrate textual and visual modalities in guiding abstractive text generation. Results show that our best model achieves a new state-of-the-art result on a large-scale benchmark dataset. The newly involved extractive objective as well as the knowledge distillation technique are proven to bring a noticeable improvement to the multimodal summarization task. Zhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 0002, Qun Liu 0001, Zhenglu Yang |
AAAI | 4 |
| 2022 | bert2BERT: Towards Reusable Pretrained Language ModelsabstractCheng Chen, Yichun Yin, Lifeng Shang, Xin Jiang, Yujia Qin, Fengyu Wang, Zhi Wang, Xiao Chen, Zhiyuan Liu, Qun Liu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Yujia Qin, Zhi Wang 0001, Xiao Chen 0012, Zhiyuan Liu 0001, Qun Liu 0001 |
ACL (1) | 4 |
| 2022 | Compression of Generative Pre-trained Language Models via QuantizationabstractChaofan Tao, Lu Hou, Wei Zhang, Lifeng Shang, Xin Jiang, Qun Liu, Ping Luo, Ngai Wong. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Chaofan Tao, Lu Hou 0002, Wei Zhang 0196, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Ping Luo 0002, Ngai Wong 0001 |
ACL (1) | 5 |
| 2022 | ClusterFormer: Neural Clustering Attention for Efficient and Effective TransformerabstractNingning Wang, Guobing Gan, Peng Zhang, Shuai Zhang, Junqiu Wei, Qun Liu, Xin Jiang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Guobing Gan, Peng Zhang 0002, Victor Junqiu Wei, Qun Liu 0001, Xin Jiang 0002 |
ACL (1) | 7 |
| 2022 | Hyperlink-induced Pre-training for Passage Retrieval in Open-domain Question AnsweringabstractJiawei Zhou, Xiaoguang Li, Lifeng Shang, Lan Luo, Ke Zhan, Enrui Hu, Xinyu Zhang, Hao Jiang, Zhao Cao, Fan Yu, Xin Jiang, Qun Liu, Lei Chen. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jiawei Zhou 0003, Lifeng Shang, Ke Zhan, Enrui Hu, Xinyu Zhang 0019, Hao Jiang 0022, Zhao Cao, Fan Yu 0004, Xin Jiang 0002, Qun Liu 0001, Lei Chen 0002 |
ACL (1) | 11 |
| 2022 | Pan More Gold from the Sand: Refining Open-domain Dialogue Training with Noisy Self-Retrieval GenerationabstractReal human conversation data are complicated, heterogeneous, and noisy, from which building open-domain dialogue systems remains a challenging task. In fact, such dialogue data still contains a wealth of information and knowledge, however, they are not fully explored. In this paper, we show existing open-domain dialogue generation methods that memorize context-response paired data with autoregressive or encode-decode language models underutilize the training data. Different from current approaches, using external knowledge, we explore a retrieval-generation training framework that can take advantage of the heterogeneous and noisy training data by considering them as “evidence”. In particular, we use BERTScore for retrieval, which gives better qualities of the evidence and generation. Experiments over publicly available datasets demonstrate that our method can help models generate better responses, even such training data are usually impressed as low-quality data. Such performance gain is comparable with those improved by enlarging the training set, even better. We also found that the model performance has a positive correlation with the relevance of the retrieved evidence. Moreover, our method performed well on zero-shot experiments, which indicates that our method can be more robust to real-world data. Yasheng Wang, Fei Mi, Pingyi Zhou, Xin Wang 0114, Jin Liu 0016, Xin Jiang 0002, Qun Liu 0001 |
COLING | 8 |
| 2022 | UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogabstractVisual Dialog aims to answer multi-round, interactive questions based on the dialog history and image content. Existing methods either consider answer ranking and generating individually or only weakly capture the relation across the two tasks implicitly by two separate models. The research on a universal framework that jointly learns to rank and generate answers in a single model is seldom explored. In this paper, we propose a contrastive learning-based framework UTC to unify and facilitate both discriminative and generative tasks in visual dialog with a single model. Specifically, considering the inherent limitation of the previous learning paradigm, we devise two inter-task contrastive losses i.e., context contrastive loss and answer contrastive loss to make the discriminative and generative tasks mutually reinforce each other. These two com-plementary contrastive losses exploit dialog context and target answer as anchor points to provide representation learning signals from different perspectives. We evaluate our proposed UTC on the VisDial v1.0 dataset, where our method outperforms the state-of-the-art on both discriminative and generative tasks and surpasses previous state-of-the-art generative methods by more than 2 absolute points on Recall@1. Zhenshan Tan, Qingrong Cheng, Xin Jiang 0002, Qun Liu 0001, Yudong Zhu, Xiaodong Gu 0001 |
CVPR | 4 |
| 2022 | LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal ModelingabstractRecent large-scale video-language pre-trained models have shown appealing performance on various downstream tasks.However, the pretraining process is computationally expensive due to the requirement of millions of videotext pairs and the redundant data structure of each video.To mitigate these problems, we propose LiteVL, which adapts a pre-trained image-language model BLIP into a video-text model directly on downstream tasks, without heavy pre-training.To enhance the temporal modeling lacking in the image-language model, we propose to add temporal attention modules in the image encoder of BLIP with dynamic temporal scaling.Besides the model-wise adaptation, we also propose a non-parametric pooling mechanism to adaptively reweight the fine-grained video embedding conditioned on the text.Experimental results on text-video retrieval and video question answering show that the proposed LiteVL even outperforms previous video-language pre-trained models by a clear margin, though without any videolanguage pre-training. Chaofan Tao, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 5 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 12 |
| 2022 | Pre-training Language Models with Deterministic Factual KnowledgeabstractPrevious works show that Pre-trained Language Models (PLMs) can capture factual knowledge.However, some analyses reveal that PLMs fail to perform it robustly, e.g., being sensitive to the changes of prompts when extracting factual knowledge.To mitigate this issue, we propose to let PLMs learn the deterministic relationship between the remaining context and the masked content.The deterministic relationship ensures that the masked factual content can be deterministically inferable based on the existing clues in the context.That would provide more stable patterns for PLMs to capture factual knowledge than randomly masking.Two pre-training tasks are further introduced to motivate PLMs to rely on the deterministic relationship when filling masks.Specifically, we use an external Knowledge Base (KB) to identify deterministic relationships and continuously pre-train PLMs with the proposed methods.The factual knowledge probing experiments indicate that the continuously pre-trained PLMs achieve better robustness in factual knowledge capturing.Further experiments on question-answering datasets show that trying to learn a deterministic relationship with the proposed methods can also help other knowledge-intensive tasks. Shaobo Li 0004, Lifeng Shang, Chengjie Sun, Bingquan Liu, Zhenzhou Ji, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 7 |
| 2022 | Hypoformer: Hybrid Decomposition Transformer for Edge-friendly Neural Machine TranslationabstractTransformer has been demonstrated effective in Neural Machine Translation (NMT).However, it is memory-consuming and time-consuming in edge devices, resulting in some difficulties for real-time feedback.To compress and accelerate Transformer, we propose a Hybrid Tensor-Train (HTT) decomposition, which retains full rank and meanwhile reduces operations and parameters.A Transformer using HTT, named Hypoformer, consistently and notably outperforms the recent lightweight SOTA methods on three standard translation tasks under different parameter and speed scales.In extreme low resource scenarios, Hypoformer has a 7.1 point absolute improvement in BLEU and 1.27× speedup than the vanilla Transformer on the IWSLT'14 De-En task. Sunzhu Li, Peng Zhang 0002, Guobing Gan, Xiuqing Lv, Benyou Wang, Victor Junqiu Wei, Xin Jiang 0002 |
EMNLP | 7 |
| 2022 | G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksabstractRecently, domain-specific PLMs have been proposed to boost the task performance of specific domains (e.g., biomedical and computer science) by continuing to pre-train general PLMs with domain-specific corpora.However, this Domain-Adaptive Pre-Training (DAPT; Gururangan et al. ( 2020)) tends to forget the previous general knowledge acquired by general PLMs, which leads to a catastrophic forgetting phenomenon and sub-optimal performance.To alleviate this problem, we propose a new framework of General Memory-Augmented Pre-trained Language Model (G-MAP), which augments the domain-specific PLM by a memory representation built from the frozen general PLM without losing any general knowledge.Specifically, we propose a new memory-augmented layer, and based on it, different augmented strategies are explored to build the memory representation and then adaptively fuse it into the domain-specific PLM.We demonstrate the effectiveness of G-MAP on various domains (biomedical and computer science publications, news, and reviews) and different kinds (text classification, QA, NER) of tasks, and the extensive results show that the proposed G-MAP 1 can achieve SOTA results on all tasks. Zhongwei Wan, Yichun Yin, Wei Zhang 0196, Jiaxin Shi, Lifeng Shang, Guangyong Chen, Xin Jiang 0002, Qun Liu 0001 |
EMNLP | 7 |
| 2022 | SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training
Wenyong Huang, Zhenhe Zhang, Yu Ting Yeung, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 4 |
| 2022 | Exploring extreme parameter compression for pre-trained language models
Benyou Wang, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001 |
ICLR | 4 |
| 2022 | FILIP: Fine-grained Interactive Language-Image Pre-Training
Lewei Yao, Runhui Huang, Lu Hou 0002, Guansong Lu, Minzhe Niu, Hang Xu 0004, Xiaodan Liang, Zhenguo Li, Xin Jiang 0002, Chunjing Xu |
ICLR | 9 |
| 2022 | Boosting Graph Structure Learning with Dummy NodesabstractWith the development of graph kernels and graph representation learning, many superior methods have been proposed to handle scalability and oversmoothing issues on graph structure learning. However, most of those strategies are designed based on practical experience rather than theoretical analysis. In this paper, we use a particular dummy node connecting to all existing vertices without affecting original vertex and edge properties. We further prove that such the dummy node can help build an efficient monomorphic edge-to-vertex transform and an epimorphic inverse to recover the original graph back. It also indicates that adding dummy nodes can preserve local and global structures for better graph representation learning. We extend graph kernels and graph neural networks with dummy nodes and conduct experiments on graph classification and subgraph isomorphism matching tasks. Empirical results demonstrate that taking graphs with dummy nodes as input significantly boosts graph structure learning, and using their edge-to-vertex graphs can also achieve similar results. We also discuss the gain of expressive power from the dummy in neural networks. Xin Liu 0039, Cheng Jiayang, Yangqiu Song, Xin Jiang 0002 |
ICML | 4 |
| 2022 | CoCA-MDD: A Coupled Cross-Attention based Framework for Streaming Mispronunciation Detection and DiagnosisabstractMispronunciation detection and diagnosis (MDD) is a popular research focus in computer-aided pronunciation training (CAPT) systems.End-to-end (e2e) approaches are becoming dominant in MDD.However an e2e MDD model usually requires entire speech utterances as input context, which leads to significant time latency especially for long paragraphs.We propose a streaming e2e MDD model called CoCA-MDD.We utilize conv-transformer structure to encode input speech in a streaming manner.A coupled cross-attention (CoCA) mechanism is proposed to integrate frame-level acoustic features with encoded reference linguistic features.CoCA also enables our model to perform mispronunciation classification with whole utterances.The proposed model allows system fusion between the streaming output and mispronunciation classification output for further performance enhancement.We evaluate CoCA-MDD on publicly available corpora.CoCA-MDD achieves F1 scores of 57.03% and 60.78% for streaming and fusion modes respectively on L2-ARCTIC.For phone-level pronunciation scoring, CoCA-MDD achieves 0.58 Pearson correlation coefficient (PCC) value on SpeechOcean762. Nianzu Zheng, Liqun Deng, Wenyong Huang, Yu Ting Yeung, Baohua Xu, Yasheng Wang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
INTERSPEECH | 9 |
| 2022 | Towards Efficient Post-training Quantization of Pre-trained Language ModelsabstractNetwork quantization has gained increasing attention with the rapid growth of large pre-trained language models~(PLMs). However, most existing quantization methods for PLMs follow quantization-aware training~(QAT) that requires end-to-end training with full access to the entire dataset. Therefore, they suffer from slow training, large memory overhead, and data accessibility issues. In this paper, we study post-training quantization~(PTQ) of PLMs, and propose module-wise quantization error minimization~(MREM), an efficient solution to mitigate these issues. By partitioning the PLM into multiple modules, we minimize the reconstruction error incurred by quantization for each module. In addition, we design a new model parallel training strategy such that each module can be trained locally on separate computing devices without waiting for preceding modules, which brings nearly the theoretical training speed-up (e.g., $4\times$ on $4$ GPUs). Experiments on GLUE and SQuAD benchmarks show that our proposed PTQ solution not only performs close to QAT, but also enjoys significant reductions in training time, memory overhead, and data consumption. Haoli Bai, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Irwin King, Michael R. Lyu |
NeurIPS | 4 |
| 2022 | Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training BenchmarkabstractVision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_\text{ViT-L}$ achieves an average accuracy of 73.03%. For the image-text retrieval task, it achieves a mean recall of 71.6% on AIC-ICC which is 12.9% higher than WenLan 2.0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e.g., Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to https://wukong-dataset.github.io/wukong-dataset/. Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 0002, Niu Minzhe, Xiaodan Liang, Lewei Yao, Runhui Huang, Wei Zhang 0196, Xin Jiang 0002, Chunjing Xu, Hang Xu 0004 |
NeurIPS | 10 |
| 2021 | HopRetriever: Retrieve Hops over Wikipedia to Answer Complex QuestionsabstractCollecting supporting evidence from large corpora of text (e.g., Wikipedia) is of great challenge for open-domain Question Answering (QA). Especially, for multi-hop open-domain QA, scattered evidence pieces are required to be gathered together to support the answer extraction. In this paper, we propose a new retrieval target, hop, to collect the hidden reasoning evidence from Wikipedia for complex question answering. Specifically, the hop in this paper is defined as the combination of a hyperlink and the corresponding outbound link document. The hyperlink is encoded as the mention embedding which models the structured knowledge of how the outbound link entity is mentioned in the textual context, and the corresponding outbound link document is encoded as the document embedding representing the unstructured knowledge within it. Accordingly, we build HopRetriever which retrieves hops over Wikipedia to answer complex questions. Experiments on the HotpotQA dataset demonstrate that HopRetriever outperforms previously published evidence retrieval methods by large margins. Moreover, our approach also yields quantifiable interpretations of the evidence collection process. Shaobo Li 0004, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Chengjie Sun, Zhenzhou Ji, Bingquan Liu |
AAAI | 4 |
| 2021 | Continuous Self-Attention Models with Neural ODE NetworksabstractStacked self-attention models receive widespread attention, due to its ability of capturing global dependency among words. However, the stacking of many layers and components generates huge parameters, leading to low parameter efficiency. In response to this issue, we propose a lightweight architecture named Continuous Self-Attention models with neural ODE networks (CSAODE). In CSAODE, continuous dynamical models (i.e., neural ODEs) are coupled with our proposed self-attention block to form a self-attention ODE solver. This solver continuously calculates and optimizes the hidden states via only one layer of parameters to improve the parameter efficiency. In addition, we design a novel accelerated continuous dynamical model to reduce computing costs, and integrate it in CSAODE. Moreover, since the original self-attention ignores local information, CSAODE makes use of N-gram convolution to encode local representations, and a fusion layer with only two trainable scalars are designed for generating sentence vectors. We perform a series of experiments on text classification, neural language inference (NLI) and text matching tasks. With fewer parameters, CSAODE outperforms state-of-the-art models on text classification tasks (e.g., 1.3% accuracy improved on SUBJ task), and has competitive performances for NLI and text matching tasks as well. Peng Zhang 0002, Baiwen Kong, Victor Junqiu Wei, Xin Jiang 0002 |
AAAI | 5 |
| 2021 | BinaryBERT: Pushing the Limit of BERT QuantizationabstractHaoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jin Jin, Xin Jiang, Qun Liu, Michael Lyu, Irwin King. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haoli Bai, Wei Zhang 0196, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Michael R. Lyu, Irwin King |
ACL/IJCNLP (1) | 6 |
| 2021 | GhostBERT: Generate More Features with Cheap Operations for BERTabstractZhiqi Huang, Lu Hou, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhiqi Huang 0001, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Exploring Discourse Structures for Argument Impact ClassificationabstractXin Liu, Jiefu Ou, Yangqiu Song, Xin Jiang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Xin Liu 0039, Jiefu Ou, Yangqiu Song, Xin Jiang 0002 |
ACL/IJCNLP (1) | 4 |
| 2021 | AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language ModelsabstractYichun Yin, Cheng Chen, Lifeng Shang, Xin Jiang, Xiao Chen, Qun Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | EditSpeech: A Text Based Speech Editing System Using Partial Inference and Bidirectional FusionabstractThis paper presents the design, implementation and evaluation of a speech editing system, named EditSpeech, which allows a user to perform deletion, insertion and replacement of words in a given speech utterance, without causing audible degradation in speech quality and naturalness. The EditSpeech system is developed upon a neural text-to-speech (NTTS) synthesis framework. Partial inference and bidirectional fusion are proposed to effectively incorporate the contextual information related to the edited region and achieve smooth transition at both left and right boundaries. Distortion introduced to the unmodified parts of the utterance is alleviated. The EditSpeech system is developed and evaluated on English and Chinese in multi-speaker scenarios. Objective and subjective evaluation demonstrate that EditSpeech outperforms a few baseline systems in terms of low spectral distortion and preferred speech quality. Audio samples are available online for demonstration11https://daxintan-cuhk.github.io/EditSpeech/. Daxin Tan, Liqun Deng, Yu Ting Yeung, Xin Jiang 0002, Xiao Chen 0012, Tan Lee |
ASRU | 4 |
| 2021 | Improving Unsupervised Question Answering via Summarization-Informed Question GenerationabstractQuestion Generation (QG) is the task of generating a plausible question for a given pair.Template-based QG uses linguistically-informed heuristics to transform declarative sentences into interrogatives, whereas supervised QG uses existing Question Answering (QA) datasets to train a system to generate a question given a passage and an answer.A disadvantage of the heuristic approach is that the generated questions are heavily tied to their declarative counterparts.A disadvantage of the supervised approach is that they are heavily tied to the domain/language of the QA dataset used as training data.In order to overcome these shortcomings, we propose an unsupervised QG method which uses questions generated heuristically from summaries as a source of training data for a QG system.We make use of freely available news summary data, transforming declarative summary sentences into appropriate questions using heuristics informed by dependency parsing, named entity recognition and semantic role labeling.The resulting questions are then combined with the original news articles to train an end-to-end neural QG model.We extrinsically evaluate our approach using unsupervised QA: our QG model is used to generate synthetic QA pairs for training a QA model.Experimental results show that, trained with only 20k English Wikipedia-based synthetic QA pairs, the QA model substantially outperforms previous unsupervised models on three in-domain datasets (SQuAD1.1,Natural Questions, TriviaQA) and three out-of-domain datasets (NewsQA, BioASQ, DuoRC), demonstrating the transferability of the approach. Chenyang Lyu, Lifeng Shang, Yvette Graham, Jennifer Foster, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 5 |
| 2021 | DyLex: Incorporating Dynamic Lexicons into BERT for Sequence LabelingabstractBaojun Wang, Zhao Zhang, Kun Xu, Guang-Yuan Hao, Yuyang Zhang, Lifeng Shang, Linlin Li, Xiao Chen, Xin Jiang, Qun Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Baojun Wang, Guang-Yuan Hao, Lifeng Shang, Linlin Li 0001, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 9 |
| 2021 | Extract then Distill: Efficient and Effective Task-Agnostic BERT Distillation
Yichun Yin, Lifeng Shang, Zhi Wang 0001, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
ICANN (3) | 5 |
| 2021 | On Position Embeddings in BERT
Benyou Wang, Lifeng Shang, Christina Lioma, Xin Jiang 0002, Hao Yang 0006, Qun Liu 0001, Jakob Grue Simonsen |
ICLR | 4 |
| 2021 | Reweighting Augmented Samples by Minimizing the Maximal Expected Loss
Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma |
ICLR | 4 |
| 2021 | Improved OOD Generalization via Adversarial Training and PretraingabstractRecently, learning a model that generalizes well on out-of-distribution (OOD) data has attracted great attention in the machine learning community. In this paper, after defining OOD generalization by Wasserstein distance, we theoretically justify that a model robust to input perturbation also generalizes well on OOD data. Inspired by previous findings that adversarial training helps improve robustness, we show that models trained by adversarial training have converged excess risk on OOD data. Besides, in the paradigm of pre-training then fine-tuning, we theoretically justify that the input perturbation robust model in the pre-training stage provides an initialization that generalizes well on downstream OOD data. Finally, various experiments conducted on image classification and natural language understanding tasks verify our theoretical findings. Mingyang Yi, Lu Hou 0002, Lifeng Shang, Xin Jiang 0002, Qun Liu 0001, Zhiming Ma |
ICML | 5 |
| 2021 | Improving task-agnostic BERT distillation with layer mapping search
Xiaoqi Jiao, Huating Chang, Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Linlin Li 0001, Fang Wang 0001, Qun Liu 0001 |
Neurocomputing | 5 |
| 2021 | EyelashNet: a dataset and a baseline method for eyelash mattingabstractEyelashes play a crucial part in the human facial structure and largely affect the facial attractiveness in modern cosmetic design. However, the appearance and structure of eyelashes can easily induce severe artifacts in high-fidelity multi-view 3D face reconstruction. Unfortunately it is highly challenging to remove eyelashes from portrait images using both traditional and learning-based matting methods due to the delicate nature of eyelashes and the lack of eyelash matting dataset. To this end, we present EyelashNet, the first eyelash matting dataset which contains 5,400 high-quality eyelash matting data captured from real world and 5,272 virtual eyelash matting data created by rendering avatars. Our work consists of a capture stage and an inference stage to automatically capture and annotate eyelashes instead of tedious manual efforts. The capture is based on a specifically-designed fluorescent labeling system. By coloring the eyelashes with a safe and invisible fluorescent substance, our system takes paired photos with colored and normal eyelashes by turning the equipped ultraviolet (UVA) flash on and off. We further correct the alignment between each pair of photos and use a novel alpha matte inference network to extract the eyelash alpha matte. As there is no prior eyelash dataset, we propose a progressive training strategy that progressively fuses captured eyelash data with virtual eyelash data to learn the latent semantics of real eyelashes. As a result, our method can accurately extract eyelash alpha mattes from fuzzy and self-shadow regions such as pupils, which is almost impossible by manual annotations. To validate the advantage of EyelashNet, we present a baseline method based on deep learning that achieves state-of-the-art eyelash matting performance with RGB portrait images as input. We also demonstrate that our work can largely benefit important real applications including high-fidelity personalized avatar and cosmetic design. Qinjie Xiao, Luyuan Wang, Xiaogang Jin 0001, Xin Jiang 0002, Tianjia Shao, Kun Zhou 0001 |
ACM Trans. Graph. | 7 |
| 2020 | Dialog State Tracking with Reinforced Data AugmentationabstractNeural dialog state trackers are generally limited due to the lack of quantity and diversity of annotated training data. In this paper, we address this difficulty by proposing a reinforcement learning (RL) based framework for data augmentation that can generate high-quality data to improve the neural state tracker. Specifically, we introduce a novel contextual bandit generator to learn fine-grained augmentation policies that can generate new effective instances by choosing suitable replacements for specific context. Moreover, by alternately learning between the generator and the state tracker, we can keep refining the generative policies to generate more high-quality training data for neural state tracker. Experimental results on the WoZ and MultiWoZ (restaurant) datasets demonstrate that the proposed framework significantly improves the performance over the state-of-the-art models, especially with limited training data. Yichun Yin, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
AAAI | 3 |
| 2020 | Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderabstractMasked language model and autoregressive language model are two types of language models.While pretrained masked language models such as BERT (Devlin et al., 2019) overwhelm the line of natural language understanding (NLU) tasks, autoregressive language models such as GPT (Radford et al., 2018) are especially capable in natural language generation (NLG).In this paper, we propose a probabilistic masking scheme for the masked language model, which we call probabilistically masked language model (PMLM).We implement a specific PMLM with a uniform prior distribution on the masking ratio named u-PMLM.We prove that u-PMLM is equivalent to an autoregressive permutated language model.One main advantage of the model is that it supports text generation in arbitrary order with surprisingly good quality, which could potentially enable new applications over traditional unidirectional generation.Besides, the pretrained u-PMLM also outperforms BERT on a set of downstream NLU tasks. Xin Jiang 0002, Qun Liu 0001 |
ACL | 2 |
| 2020 | Accurate Word Alignment Induction from Neural Machine TranslationabstractDespite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism.In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET.The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work.SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change.SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments.Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8AER points. 1 Yun Chen 0007, Yang Liu 0005, Guanhua Chen 0001, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 4 |
| 2020 | TernaryBERT: Distillation-aware Ultra-low Bit BERTabstractTransformer-based pre-training models like BERT have achieved remarkable performance in many natural language processing tasks.However, these models are both computation and memory expensive, hindering their deployment to resource-constrained devices.In this work, we propose TernaryBERT, which ternarizes the weights in a fine-tuned BERT model.Specifically, we use both approximation-based and loss-aware ternarization methods and empirically investigate the ternarization granularity of different parts of BERT.Moreover, to reduce the accuracy degradation caused by the lower capacity of low bits, we leverage the knowledge distillation technique (Jiao et al., 2019) in the training process.Experiments on the GLUE benchmark and SQuAD show that our proposed TernaryBERT outperforms the other BERT quantization methods, and even achieves comparable performance as the fullprecision model while being 14.9x smaller. Wei Zhang 0196, Lu Hou 0002, Yichun Yin, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qun Liu 0001 |
EMNLP (1) | 6 |
| 2020 | Progressive Memory Banks for Incremental Domain Adaptation
Nabiha Asghar, Lili Mou, Kira A. Selby, Kevin D. Pantasdo, Pascal Poupart, Xin Jiang 0002 |
ICLR | 6 |
| 2020 | On the Importance of Word and Sentence Representation Learning in Implicit Discourse Relation ClassificationabstractImplicit discourse relation classification is one of the most difficult parts in shallow discourse parsing as the relation prediction without explicit connectives requires the language understanding at both the text span level and the sentence level. Previous studies mainly focus on the interactions between two arguments. We argue that a powerful contextualized representation module, a bilateral multi-perspective matching module, and a global information fusion module are all important to implicit discourse analysis. We propose a novel model to combine these modules together. Extensive experiments show that our proposed model outperforms BERT and other state-of-the-art systems on the PDTB dataset by around 8% and CoNLL 2016 datasets around 16%. We also analyze the effectiveness of different modules in the implicit discourse relation classification task and demonstrate how different levels of representation learning can affect the results. Xin Liu 0039, Jiefu Ou, Yangqiu Song, Xin Jiang 0002 |
IJCAI | 4 |
| 2020 | An Investigation of Few-Shot Learning in Spoken Term Classificationabstract202402 bcch Yangbin Chen, Tom Ko, Lifeng Shang, Xiao Chen 0012, Xin Jiang 0002, Qing Li 0001 |
INTERSPEECH | 5 |
| 2020 | Neural Subgraph Isomorphism CountingabstractIn this paper, we study a new graph learning problem: learning to count subgraph isomorphisms. Different from other traditional graph learning problems such as node classification and link prediction, subgraph isomorphism counting is NP-complete and requires more global inference to oversee the whole graph. To make it scalable for large-scale graphs and patterns, we propose a learning framework that augments different representation learning architectures and iteratively attends pattern and target data graphs to memorize intermediate states of subgraph isomorphism searching for global counting. We develop both small graphs (<= 1,024 subgraph isomorphisms in each) and large graphs (<= 4,096 subgraph isomorphisms in each) sets to evaluate different representation and interaction modules. A mutagenic compound dataset, MUTAG, is also used to evaluate neural models and demonstrate the success of transfer learning. While the learning based approach is inexact, we are able to generalize to count large patterns and data graphs in linear time compared to the exponential time of the original NP-complete problem. Experimental results show that learning based subgraph isomorphism counting can speed up the traditional algorithm, VF2, 10-1,000 times with acceptable errors. Domain adaptation based on fine-tuning also shows the usefulness of our approach in real-world applications. Xin Liu 0039, Haojie Pan, Mutian He 0001, Yangqiu Song, Xin Jiang 0002, Lifeng Shang |
KDD | 5 |
| 2020 | DynaBERT: Dynamic BERT with Adaptive Width and DepthabstractThe pre-trained language models like BERT, though powerful in many natural language processing tasks, are both computation and memory expensive. To alleviate this problem, one approach is to compress them for specific tasks before deployment. However, recent works on BERT compression usually compress the large BERT model to a fixed smaller size, and can not fully satisfy the requirements of different edge devices with various hardware performances. In this paper, we propose a novel dynamic BERT model (abbreviated as DynaBERT), which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. Comprehensive experiments under various efficiency constraints demonstrate that our proposed dynamic BERT (or RoBERTa) at its largest size has comparable performance as BERT-base (or RoBERTa-base), while at smaller widths and depths consistently outperforms existing BERT compression methods. Code is available at https://github.com/huawei-noah/Pretrained-Language-Model/tree/master/DynaBERT. Lu Hou 0002, Zhiqi Huang 0001, Lifeng Shang, Xin Jiang 0002, Xiao Chen 0012, Qun Liu 0001 |
NeurIPS | 4 |
| 2020 | Unsupervised Text Generation by Learning from SearchabstractIn this work, we propose TGLS, a novel framework for unsupervised Text Generation by Learning from Search. We start by applying a strong search algorithm (in particular, simulated annealing) towards a heuristically defined objective that (roughly) estimates the quality of sentences. Then, a conditional generative model learns from the search results, and meanwhile smooth out the noise of search. The alternation between search and learning can be repeated for performance bootstrapping. We demonstrate the effectiveness of TGLS on two real-world natural language generation tasks, unsupervised paraphrasing and text formalization. Our model significantly outperforms unsupervised baseline methods in both tasks. Especially, it achieves comparable performance to strong supervised methods for paraphrase generation. Jingjing Li 0007, Zichao Li 0001, Lili Mou, Xin Jiang 0002, Michael R. Lyu, Irwin King |
NeurIPS | 4 |
| 2019 | Decomposable Neural Paraphrase GenerationabstractParaphrasing exists at different granularity levels, such as lexical level, phrasal level and sentential level.This paper presents Decomposable Neural Paraphrase Generator (DNPG), a Transformer-based model that can learn and generate paraphrases of a sentence at different levels of granularity in a disentangled way.Specifically, the model is composed of multiple encoders and decoders with different structures, each of which corresponds to a specific granularity.The empirical study shows that the decomposition mechanism of DNPG makes paraphrase generation more interpretable and controllable.Based on DNPG, we further develop an unsupervised domain adaptation method for paraphrase generation.Experimental results show that the proposed model achieves competitive in-domain performance compared to the state-of-the-art neural models, and significantly better performance when adapting to a new domain.What is the population of New York?How many people is there in NYC?Who wrote the Winnie the Pooh books?Who is the author of winnie the pooh?What is the best phone to buy below 15k?Which are best mobile phones to buy under 15000?How can I be a good geologist?What should I do to be a great geologist?How do I reword a sentence to avoid plagiarism?How can I paraphrase my essay and avoid plagiarism? Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Qun Liu 0001 |
ACL (1) | 2 |
| 2019 | ERNIE: Enhanced Language Representation with Informative EntitiesabstractNeural language representation models such as BERT pre-trained on large-scale corpora can well capture rich semantic patterns from plain text, and be fine-tuned to consistently improve the performance of various NLP tasks.However, the existing pre-trained language models rarely consider incorporating knowledge graphs (KGs), which can provide rich structured knowledge facts for better language understanding.We argue that informative entities in KGs can enhance language representation with external knowledge.In this paper, we utilize both large-scale textual corpora and KGs to train an enhanced language representation model (ERNIE), which can take full advantage of lexical, syntactic, and knowledge information simultaneously.The experimental results have demonstrated that ERNIE achieves significant improvements on various knowledge-driven tasks, and meanwhile is comparable with the state-of-the-art model BERT on other common NLP tasks.The source code and experiment details of this paper can be obtained from https:// github.com/thunlp/ERNIE. Zhengyan Zhang, Xu Han 0007, Zhiyuan Liu 0001, Xin Jiang 0002, Maosong Sun 0001, Qun Liu 0001 |
ACL (1) | 4 |
| 2019 | Exploring Diverse Expressions for Paraphrase GenerationabstractLihua Qian, Lin Qiu, Weinan Zhang, Xin Jiang, Yong Yu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lihua Qian, Weinan Zhang 0001, Xin Jiang 0002, Yong Yu 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Triple-to-Text: Converting RDF Triples into High-Quality Natural Languages via Optimizing an Inverse KL DivergenceabstractKnowledge base is one of the main forms to represent information in a structured way. A knowledge base typically consists of Resource Description Frameworks (RDF) triples which describe the entities and their relations. Generating natural language description of the knowledge base is an important task in NLP, which has been formulated as a conditional language generation task and tackled using the sequence-to-sequence framework. Current works mostly train the language models by maximum likelihood estimation, which tends to generate lousy sentences. In this paper, we argue that such a problem of maximum likelihood estimation is intrinsic, which is generally irrevocable via changing network structures. Accordingly, we propose a novel Triple-to-Text (T2T) framework, which approximately optimizes the inverse Kullback-Leibler (KL) divergence between the distributions of the real and generated sentences. Due to the nature that inverse KL imposes large penalty on fake-looking samples, the proposed method can significantly reduce the probability of generating low-quality sentences. Our experiments on three real-world datasets demonstrate that T2T can generate higher-quality sentences and outperform baseline models in several evaluation metrics. Yaoming Zhu, Juncheng Wan, Zhiming Zhou 0001, Weinan Zhang 0001, Xin Jiang 0002, Yong Yu 0001 |
SIGIR | 7 |
| 2018 | Affective Neural Response Generation
Nabiha Asghar, Pascal Poupart, Jesse Hoey, Xin Jiang 0002, Lili Mou |
ECIR | 4 |
| 2018 | Paraphrase Generation with Deep Reinforcement LearningabstractAutomatic generation of paraphrases from a given sentence is an important yet challenging task in natural language processing (NLP).In this paper, we present a deep reinforcement learning approach to paraphrase generation.Specifically, we propose a new framework for the task, which consists of a generator and an evaluator, both of which are learned from data.The generator, built as a sequenceto-sequence learning model, can produce paraphrases given a sentence.The evaluator, constructed as a deep matching model, can judge whether two sentences are paraphrases of each other.The generator is first trained by deep learning and then further fine-tuned by reinforcement learning in which the reward is given by the evaluator.For the learning of the evaluator, we propose two methods based on supervised learning and inverse reinforcement learning respectively, depending on the type of available training data.Experimental results on two datasets demonstrate the proposed models (the generators) can produce more accurate paraphrases and outperform the stateof-the-art methods in paraphrase generation in both automatic evaluation and human evaluation. Zichao Li 0001, Xin Jiang 0002, Lifeng Shang, Hang Li 0001 |
EMNLP | 2 |
| 2016 | Neural Generative Question Answering
Xin Jiang 0002, Zhengdong Lu, Lifeng Shang, Hang Li 0001 |
IJCAI | 2 |
| 2014 | Ranking Optimization with ConstraintsabstractThis paper addresses the problem of post-processing of ranking in search, referred to as post ranking. Although important, no research seems to have been conducted on the problem, particularly with a principled approach, and in practice ad-hoc ways of performing the task are being adopted. This paper formalizes the problem as constrained optimization in which the constraints represent the post-processing rules and the objective function represents the trade-off between adherence to the original ranking and satisfaction of the rules. The optimization amounts to refining the original ranking result based on the rules. We further propose a specific probabilistic implementation of the general formalization on the basis of the Bradley-Terry model, which is theoretically sound, effective, and efficient. Our experimental results, using benchmark datasets and enterprise search dataset, show that the proposed method works much better than several baseline methods of utilizing rules. Fangzhao Wu, Jun Xu 0001, Hang Li 0001, Xin Jiang 0002 |
CIKM | 4 |
| 2009 | A ranking approach to keyphrase extractionabstractThis paper addresses the issue of automatically extracting keyphrases from a document. Previously, this problem was formalized as classification and learning methods for classification were utilized. This paper points out that it is more essential to cast the problem as ranking and employ a learning to rank method to perform the task. Specifically, it employs Ranking SVM, a state-of-art method of learning to rank, in keyphrase extraction. Experimental results on three datasets show that Ranking SVM significantly outperforms the baseline methods of SVM and Naive Bayes, indicating that it is better to exploit learning to rank techniques in keyphrase extraction. Xin Jiang 0002, Yunhua Hu, Hang Li 0001 |
SIGIR | 1 |