Liang Ding 0006

dblp:88/3340-6 · DBLP profile ↗
← Back
83ranked-venue papers
7as first author
81since 2021 · last 2026
0000-0001-8976-2084ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 67 · 7 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 18 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model
abstract
Long Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions: 1) The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient.
Rong Bao, Bo Wang 0084, Xiao Wang 0014, Hongyu Li 0004, Leszek Rutkowski, Qi Zhang 0001, Liang Ding 0006, Dacheng Tao
AAAI8
2026 The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Check
abstract
The pursuit of real-time agentic interaction has driven interest in Diffusion-based Large Language Models (dLLMs) as alternatives to autoregressive backbones, promising to break the sequential latency bottleneck.However, does such efficiency gains translate into effective agentic behavior?In this work, we present a comprehensive evaluation of dLLMs (e.g., LLaDA, Dream) across two distinct agentic paradigms: Embodied Agents (requiring longhorizon planning) and Tool-Calling Agents (requiring precise formatting).Contrary to the efficiency hype, our results on Agentboard and BFCL reveal a "bitter lesson": current dLLMs fail to serve as reliable agentic backbones, frequently leading to systematic failure.(1) In Embodied settings, dLLMs suffer repeated attempts, failing to branch under temporal feedback.(2) In Tool-Calling settings, dLLMs fail to maintain symbolic precision (e.g.strict JSON schemas) under diffusion noise.To assess the potential of dLLMs in agentic workflows, we introduce DiffuAgent, a multi-agent evaluation framework that integrates dLLMs as plug-and-play cognitive cores.Our analysis shows that dLLMs are effective in non-causal roles (e.g., memory summarization and tool selection) but require the incorporation of causal, precise, and logically grounded reasoning mechanisms into the denoising process to be viable for agentic tasks.
Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Jinxia Zhang, Dacheng Tao
ACL (1)2
2026 Rethinking the Hidden Risk of Reranking: Achieving Risk-aware Reranking with Information Gain for RAG with LLMs
abstract
Retrieval-augmented generation (RAG) has become a cornerstone for enhancing large language models (LLMs) with real-time information from the Web, but its performance often heavily depends on the quality of the retrieved documents. Given that RAG systems frequently draw from vast and often noisy Web corpora, ensuring the reliability of retrieved content is paramount. While rerankers improve the factual accuracy of the RAG system by elevating the proportion of ground-truth documents (GD) in high-ranked results, the shifts of document type distributions during reranking remain unclear, hindering the understanding of the reranker's behavior. To bridge this gap, we conduct an empirical study to categorize documents and compare their distribution before and after reranking. We reveal a counterintuitive finding: though rerankers improve the proportion of GD, they also significantly increase the proportion of harmful documents (HD) in top-ranked retrieved documents. It not only narrows the potential context window for ranking the GD higher but also increases the risk of HD misleading the LLMs, potentially leading to the generation and propagation of misinformation across Web platforms. Motivated by this finding, we propose a risk-aware reranking method for RAG with LLMs, which balances the risk and benefit during reranking. Given a query, the RAG framework first retrieves relevant documents. Then, our approach quantifies the potential beneficial and harmful impacts of various documents on the LLMs' generation. To estimate the impacts, we conduct a dual-aspect document impact assessment via information gain, which employs a risk clipping to avoid the numerical fluctuations in the estimation. Finally, we conduct the reranking according to the potential impact of each document, enabling the reranker to significantly reduce the HD proportion. Experiments and analysis across multiple models and datasets, including Wikipedia, web news, and research papers, show the effectiveness of our method. Our code is available at https://github.com/lzz335/hidden_risk_of_reranking.
Zhizhao Liu, Zhihua Wen, Zhiliang Tian, Zhen Huang 0006, Miaorong Zhu, Zimian Wei, Yifu Gao, Liang Ding 0006, Dongsheng Li 0001
WWW8
2026 Improving zero-shot translation with the navigation ability-enhanced language tags
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibin Lei, Yibing Zhan, Weifeng Liu 0001
Eng. Appl. Artif. Intell.2
2026 Achieving >97% on GSM8K: deeply understanding the problems makes LLMs better solvers for math word problems
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001
Frontiers Comput. Sci.4
2026 Towards alleviating hallucination in text-to-image retrieval for CLIP in zero-shot learning
Hanyao Wang, Yibing Zhan, Liu Liu 0014, Liang Ding 0006, Jun Yu 0002
Neurocomputing4
2026 Exploring and enhancing the transfer of distribution in knowledge distillation for autoregressive language models
Jun Rao, Xuebo Liu 0002, Zepeng Lin, Liang Ding 0006, Jing Li 0034, Min Zhang 0005
Knowl. Based Syst.4
2026 Deep Model Fusion: A Survey
abstract
Deep model fusion/merging is an emerging technique that integrates parameters or predictions from multiple deep learning (DL) models into a unified framework. It combines the abilities of different models to compensate for the biases and errors of an individual model, improving overall performance. However, deep model fusion, especially on large-scale DL models such as large language models (LLMs) and foundation models, faces several challenges, including high computational cost and interference between different heterogeneous models. In order to understand it better, we present a comprehensive survey to summarize the recent progress. We categorize existing model fusion methods as fourfold: 1) weight average (WA) averages the parameters of multiple models to obtain results closer to the optimal solution; 2) considering that direct averaging of models often yields suboptimal results, "mode connectivity" connects networks via paths of nonincreasing loss in weight spaces before the fusion. Along these paths, initial models are transformed into forms with consistent functions and better fusion effects; 3) similarly, for models with poor direct fusion results, "alignment" matches the corresponding units and merges these models, thus fully exploiting the corresponding relationships between the models; and 4) in addition to the above-mentioned methods of parameter fusion, "ensemble learning" fuses the outputs of multiple models in the inference stage to improve the accuracy and robustness of networks. In addition, we analyze the challenges of deep model fusion and illuminate the possible research directions in the future.
Yong Peng 0006, Miao Zhang 0037, Liang Ding 0006, Han Hu 0003, Li Shen 0008
IEEE Trans. Neural Networks Learn. Syst.4
2025 Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
abstract
Multimodal large language models (MLLMs) have experienced significant advancements recently, but still struggle to recognize and interpret intricate details in high-resolution (HR) images effectively. While state-of-the-art (SOTA) MLLMs claim to process images at 4K resolution, existing MLLM benchmarks only support up to 2K, leaving the capabilities of SOTA models on true HR images largely untested. Furthermore, existing methods for enhancing HR image perception in MLLMs rely on computationally expensive visual instruction tuning. To address these limitations, we introduce HR-Bench, the first deliberately designed benchmark to rigorously evaluate MLLM performance on 4K & 8K images. Through extensive experiments, we demonstrate that while downsampling HR images leads to vision information loss, leveraging complementary modalities, e.g., text, can effectively compensate for this loss. Building upon this insight, we propose Divide, Conquer and Combine, a novel training-free framework for enhancing MLLM perception of HR images. Our method follows a three-staged approach: 1) Divide: recursively partitioning the HR image into patches and merging similar patches to minimize computational overhead, 2) Conquer: leveraging the MLLM to generate accurate textual descriptions for each image patch, and 3) Combine: utilizing the generated text descriptions to enhance the MLLM's understanding of the overall HR image. Extensive experiments show that: 1) the SOTA MLLM achieves 63% accuracy, which is markedly lower than the 87% accuracy achieved by humans on HR-Bench; 2) our method brings consistent and significant improvements (a relative increase of +6% on HR-Bench and +8% on general multimodal benchmarks).
Liang Ding 0006, Minyan Zeng, Xiabin Zhou, Li Shen 0008, Yong Luo 0002, Wei Yu 0004, Dacheng Tao
AAAI2
2025 Improving Complex Reasoning over Knowledge Graph with Logic-Aware Curriculum Tuning
abstract
Answering complex queries over incomplete knowledge graphs (KGs) is a challenging job. Most previous works have focused on learning entity/relation embeddings and simulating first-order logic operators with various neural networks. However, they are bottlenecked by the inability to share world knowledge to improve logical reasoning, thus resulting in suboptimal performance. In this paper, we propose a complex reasoning schema over KG upon large language models (LLMs), containing a curriculum-based logical-aware instruction tuning framework, named LACT. Specifically, we augment the arbitrary first-order logical queries via binary tree decomposition, to stimulate the reasoning capability of LLMs. To address the difficulty gap among different types of complex queries, we design a simple and flexible logic-aware curriculum learning framework. Experiments across widely used datasets demonstrate that LACT has substantial improvements~(brings an average +5.5% MRR score) over advanced methods, achieving the new state-of-the-art.
Tianle Xia, Liang Ding 0006, Guojia Wan, Yibing Zhan, Bo Du 0001, Dacheng Tao
AAAI2
2025 AgentDropout: Dynamic Agent Elimination for Token-Efficient and High-Performance LLM-Based Multi-Agent Collaboration
abstract
Multi-agent systems (MAS) based on large language models (LLMs) have demonstrated significant potential in collaborative problemsolving.However, they still face substantial challenges of low communication efficiency and suboptimal task performance, making the careful design of the agents' communication topologies particularly important.Inspired by the management theory that roles in an efficient team are often dynamically adjusted, we propose AgentDropout, which identifies redundant agents and communication across different communication rounds by optimizing the adjacency matrices of the communication graphs and eliminates them to enhance both token efficiency and task performance.Compared to state-of-the-art methods, AgentDropout achieves an average reduction of 21.6% in prompt token consumption and 18.4% in completion token consumption, along with a performance improvement of 1.14 on the tasks.Furthermore, the extended experiments demonstrate that AgentDropout achieves notable domain transferability and structure robustness, revealing its reliability and effectiveness.We release our code at https://github. com/wangzx1219/AgentDropout.
Zhexuan Wang, Xuebo Liu 0002, Liang Ding 0006, Miao Zhang 0037, Jie Liu 0001, Min Zhang 0005
ACL (1)4
2025 VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
abstract
Recent advancements in Large Vision-Language Models have showcased remarkable capabilities. However, they often falter when confronted with complex reasoning tasks that humans typically address through visual aids and deliberate, step-by-step thinking. While existing methods have explored text-based slow thinking or rudimentary visual assistance, they fall short of capturing the intricate, interleaved nature of human visual-verbal reasoning processes. To overcome these limitations and inspired by the mechanisms of slow thinking in human cognition, we introduce VisuoThink, a novel framework that seamlessly integrates visuospatial and linguistic domains. VisuoThink facilitates multimodal slow thinking by enabling progressive visual-textual reasoning and incorporates test-time scaling through look-ahead tree search. Extensive experiments demonstrate that VisuoThink significantly enhances reasoning capabilities via inference-time scaling, even without fine-tuning, achieving state-of-the-art performance in tasks involving geometry and spatial reasoning.
Yikun Wang 0001, Siyin Wang, Qinyuan Cheng, Zhaoye Fei, Liang Ding 0006, Qipeng Guo, Dacheng Tao, Xipeng Qiu
ACL (1)5
2025 MQM-APE: Toward High-Quality Error Annotation Predictors with Automatic Post-Editing in LLM Translation Evaluators
abstract
Large Language Models (LLMs) have shown significant potential as judges for Machine Translation (MT) quality assessment, providing both scores and fine-grained feedback. Although approaches such as GEMBA-MQM have shown state-of-the-art performance on reference-free evaluation, the predicted errors do not align well with those annotated by human, limiting their interpretability as feedback signals. To enhance the quality of error annotations predicted by LLM evaluators, we introduce a universal and training-free framework, MQM-APE, based on the idea of filtering out non-impactful errors by Automatically Post-Editing (APE) the original translation based on each error, leaving only those errors that contribute to quality improvement. Specifically, we prompt the LLM to act as 1) evaluator to provide error annotations, 2) post-editor to determine whether errors impact quality improvement and 3) pairwise quality verifier as the error filter. Experiments show that our approach consistently improves both the reliability and quality of error spans against GEMBA-MQM, across eight LLMs in both high- and low-resource languages. Orthogonal to trained approaches, MQM-APE complements translation-specific evaluators such as Tower, highlighting its broad applicability. Further analysis confirms the effectiveness of each module and offers valuable insights into evaluator design and LLMs selection.
Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Jinxia Zhang, Dacheng Tao
COLING2
2025 Self-Evolution Knowledge Distillation for LLM-based Machine Translation
abstract
Knowledge distillation (KD) has shown great promise in transferring knowledge from larger teacher models to smaller student models. However, existing KD strategies for large language models often minimize output distributions between student and teacher models indiscriminately for each token. This overlooks the imbalanced nature of tokens and their varying transfer difficulties. In response, we propose a distillation strategy called Self-Evolution KD. The core of this approach involves dynamically integrating teacher distribution and one-hot distribution of ground truth into the student distribution as prior knowledge, which promotes the distillation process. It adjusts the ratio of prior knowledge based on token learning difficulty, fully leveraging the teacher model’s potential. Experimental results show our method brings an average improvement of approximately 1.4 SacreBLEU points across four translation directions in the WMT22 test sets. Further analysis indicates that the improvement comes from better knowledge transfer from teachers, confirming our hypothesis.
Yuncheng Song, Liang Ding 0006, Changtong Zan, Shujian Huang
COLING2
2025 Intention Analysis Makes LLMs A Good Jailbreak Defender
abstract
Aligning large language models (LLMs) with human values, particularly when facing complex and stealthy jailbreak attacks, presents a formidable challenge. Unfortunately, existing methods often overlook this intrinsic nature of jailbreaks, which limits their effectiveness in such complex scenarios. In this study, we present a simple yet highly effective defense strategy, i.e., Intention Analysis (IA). IA works by triggering LLMs’ inherent self-correct and improve ability through a two-stage process: 1) analyzing the essential intention of the user input, and 2) providing final policy-aligned responses based on the first round conversation. Notably,IA is an inference-only method, thus could enhance LLM safety without compromising their helpfulness. Extensive experiments on varying jailbreak benchmarks across a wide range of LLMs show that IA could consistently and significantly reduce the harmfulness in responses (averagely -48.2% attack success rate). Encouragingly, with our IA, Vicuna-7B even outperforms GPT-3.5 regarding attack success rate. We empirically demonstrate that, to some extent, IA is robust to errors in generated intentions. Further analyses reveal the underlying principle of IA: suppressing LLM’s tendency to follow jailbreak prompts, thereby enhancing safety.
Yuqi Zhang 0002, Liang Ding 0006, Lefei Zhang, Dacheng Tao
COLING2
2025 Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
abstract
Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions.Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent.This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context.We present TOXIREWRITECN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity.The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans.It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues.We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity.Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs.We release TOXIREWRITECN to support future research on controllable, sentiment-aware detoxification for Chinese.Caution: This paper contains examples of violent or offensive language that may be disturbing to some readers.
Xintong Wang 0001, Jingheng Pan, Liang Ding 0006, Longyue Wang, Chris Biemann
EMNLP4
2025 The Energy Loss Phenomenon in RLHF: A New Perspective on Mitigating Reward Hacking
abstract
This work identifies the *Energy Loss Phenomenon* in Reinforcement Learning from Human Feedback (RLHF) and its connection to reward hacking. Specifically, energy loss in the final layer of a Large Language Model (LLM) gradually increases during the RL process, with an *excessive* increase in energy loss characterizing reward hacking. Beyond empirical analysis, we further provide a theoretical foundation by proving that, under mild conditions, the increased energy loss reduces the upper bound of contextual relevance in LLMs, which is a critical aspect of reward hacking as the reduced contextual relevance typically indicates overfitting to reward model-favored patterns in RL. To address this issue, we propose an *Energy loss-aware PPO algorithm (EPPO)* which penalizes the increase in energy loss in the LLM's final layer during reward calculation to prevent excessive energy loss, thereby mitigating reward hacking. We theoretically show that EPPO can be conceptually interpreted as an entropy-regularized RL algorithm, which provides deeper insights into its effectiveness. Extensive experiments across various LLMs and tasks demonstrate the commonality of the energy loss phenomenon, as well as the effectiveness of EPPO in mitigating reward hacking and improving RLHF performance.
Yuchun Miao, Sen Zhang 0006, Liang Ding 0006, Yuqi Zhang 0002, Lefei Zhang, Dacheng Tao
ICML3
2025 Retrieval-Augmented Perception: High-resolution Image Perception Meets Visual RAG
abstract
High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs). To drive progress beyond the limits of heuristic methods, this paper advances HR perception capabilities of MLLMs by harnessing cutting-edge long-context techniques such as retrieval-augmented generation (RAG). Towards this end, this paper presents the first study exploring the use of RAG to address HR perception challenges. Specifically, we propose Retrieval-Augmented Perception (RAP), a training-free framework that retrieves and fuses relevant image crops while preserving spatial context using the proposed Spatial-Awareness Layout. To accommodate different tasks, the proposed Retrieved-Exploration Search (RE-Search) dynamically selects the optimal number of crops based on model confidence and retrieval scores. Experimental results on HR benchmarks demonstrate the significant effectiveness of RAP, with LLaVA-v1.5-13B achieving a 43% improvement on $V^*$ Bench and 19% on HR-Bench. Code is available at https://github.com/DreamMr/RAP.
Yongcheng Jing, Liang Ding 0006, Li Shen 0008, Yong Luo 0002, Bo Du 0001, Dacheng Tao
ICML3
2025 Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity in LLMs
abstract
Continual learning for large language models (LLMs) demands a precise balance between $\textbf{plasticity}$ - the ability to absorb new tasks - and $\textbf{stability}$ - the preservation of previously learned knowledge. Conventional rehearsal methods, which replay stored examples, are limited by long-term data inaccessibility; earlier pseudo-rehearsal methods require additional generation modules, while self-synthesis approaches often generate samples that poorly align with real tasks, suffer from unstable outputs, and ignore task relationships. We present $\textbf{\textit{Self-Evolving Pseudo-Rehearsal for Catastrophic Forgetting with Task Similarity}}(\textbf{SERS})$, a lightweight framework that 1) decouples pseudo-input synthesis from label creation, using semantic masking and template guidance to produce diverse, task-relevant prompts without extra modules; 2) applies label self-evolution, blending base-model priors with fine-tuned outputs to prevent over-specialization; and 3) introduces a dynamic regularizer driven by the Wasserstein distance between task distributions, automatically relaxing or strengthening constraints in proportion to task similarity. Experiments across diverse tasks on different LLMs show that our SERS reduces forgetting by over 2\% points against strong pseudo-rehearsal baselines, by ensuring efficient data utilization and wisely transferring knowledge. The code will be released at https://github.com/JerryWangJun/LLM_CL_SERS/.
Liang Ding 0006, Shuai Wang 0011, Hongyu Li 0004, Yong Luo 0002, Huangxuan Zhao, Han Hu 0003, Bo Du 0001
NeurIPS2
2025 Layer as Puzzle Pieces: Compressing Large Language Models through Layer Concatenation
abstract
Large Language Models (LLMs) excel at natural language processing tasks, but their massive size leads to high computational and storage demands. Recent works have sought to reduce their model size through layer-wise structured pruning. However, they tend to ignore retaining the capabilities in the pruned part. In this work, we re-examine structured pruning paradigms and uncover several key limitations: 1) notable performance degradation due to direct layer removal, 2) incompetent linear weighted layer aggregation, and 3) the lack of effective post-training recovery mechanisms. To address these limitations, we propose CoMe, including a progressive layer pruning framework with a Concatenation-based Merging technology and a hierarchical distillation post-training process. Specifically, we introduce a channel sensitivity metric that utilizes activation intensity and weight norms for fine-grained channel selection. Subsequently, we employ a concatenation-based layer merging method to fuse the most critical channels in the adjacent layers, enabling a progressive model size reduction. Finally, we propose a hierarchical distillation protocol, which leverages the correspondences between the original and pruned model layers established during pruning, enabling efficient knowledge transfer. Experiments on seven benchmarks show that CoMe achieves state-of-the-art performance; when pruning 30% of LLaMA-2-7b's parameters, the pruned model retains 83% of its original average accuracy.
Fei Wang 0032, Li Shen 0008, Liang Ding 0006, Chao Xue 0003, Ye Liu 0014, Changxing Ding
NeurIPS3
2025 Widening the bottleneck of lexical choice for non-autoregressive translation
Liang Ding 0006, Longyue Wang, Siyou Liu, Weihua Luo, Kaifu Zhang
Comput. Speech Lang.1
2025 Code-switching finetuning: Bridging multilingual pretrained language models for enhanced cross-lingual performance
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001
Eng. Appl. Artif. Intell.2
2025 Recursively summarizing enables long-term dialogue memory in large language models
Qingyue Wang, Yanhe Fu, Yanan Cao 0001, Shuai Wang 0011, Zhiliang Tian, Liang Ding 0006
Neurocomputing6
2025 Hypnos: A domain-specific large language model for anesthesiology
Zhonghai Wang, Yibing Zhan, Bohao Zhou, Chong Zhang 0013, Baosheng Yu, Liang Ding 0006, Weifeng Liu 0001
Neurocomputing8
2025 Building accurate translation-tailored large language models with language-aware instruction tuning
abstract
Large language models (LLMs) exhibit remarkable capabilities in various natural language processing tasks, such as machine translation. However, the large number of LLM parameters incurs significant costs during inference. Previous studies have attempted to train translation-tailored LLMs with moderately sized models by fine-tuning them on the translation data. Nevertheless, when performing translations in zero-shot directions that are absent from the fine-tuning data, the problem of ignoring instructions and thus producing translations in the wrong language (i.e., the off-target translation issue) remains unresolved. In this work, we design a two-stage fine-tuning algorithm to improve the instruction-following ability of translation-tailored LLMs, particularly for maintaining accurate translation directions. We first fine-tune LLMs on the translation data to elicit basic translation capabilities. At the second stage, we construct instruction-conflicting samples by randomly replacing the instructions with the incorrect ones. Then, we introduce an extra unlikelihood loss to reduce the probability assigned to those samples. Experiments on two benchmarks using the LLaMA 2 and LLaMA 3 models, spanning 16 zero-shot directions, demonstrate that, compared to the competitive baseline—translation-finetuned LLaMA, our method could effectively reduce the off-target translation ratio (up to −62.4 percentage points), thus improving translation quality (up to +9.7 bilingual evaluation understudy). Analysis shows that our method can preserve the model’s performance on other tasks, such as supervised translation and general tasks. Code is released at https://github.com/alphadl/LanguageAware_Tuning .
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibing Zhan, Xinghao Yang, Weifeng Liu 0001
Frontiers Inf. Technol. Electron. Eng.2
2025 Fuzzy-Assisted Contrastive Decoding Improving Code Generation of Large Language Models
abstract
Large Language Models (LLMs) play a crucial role in intelligent code generation tasks. Most existing work focuses on pre-training or fine-tuning specialized code LLMs, e.g., CodeLlama. However, pre-training or fine-tuning a code LLM requires a vast corpus of data, significant computational resources, and considerable human effort. Compared to pre-training or fine-tuning LLMs, a simple and flexible method of contrastive decoding has garnered widespread attention to improve the text generation quality of LLMs. While contrastive decoding can indeed improve the text generation quality of LLMs, our research has found that directly using contrastive decoding: 1) introduces erroneous information into the logit distribution generated from normal prompts (i.e., user's input), particularly in the code generation of LLMs; 2) significantly impedes the inference and decoding time of LLMs. In this work, the limitations of using contrastive decoding directly are systematically highlighted, and a novel real-time fuzzy-assisted contrastive decoding (FCD) mechanism is proposed to improve the code generation quality of LLMs. The proposed FCD mechanism initially categorises prompts into high-quality and low-quality groups based on the results of the evaluator (i.e., unit test) before integrating the LLM. Next, feature values (e.g., standard deviation, peak value, etc.) related to the logit distribution of predicted tokens during the LLM's inference process for both high-quality and low-quality prompts are extracted. Finally, the extracted feature values are used to train the fuzzy neural network (i.e, fuzzy min-max neural network) offline, allowing for the prejudgement of the reliability of the logit distribution for normal prompt outputs. This prevents the direct use of erroneous information from contrastive decoding and improves the code generation quality of LLMs. Through extensive experiments, it has been demonstrated that the proposed FCD mechanism can significantly improve the code generation quality of LLMs through fuzzy-assisted contrastive decoding. Moreover, the FCD mechanism can also reduce the time required for inference and contrastive decoding. The code and data are publicly available on GitHub11https://github.com/LLMcodegen/Fuzzy_contrastive_decoding.and HuggingFace22https://huggingface.co/wangle123/Fuzzy_contrastive_decoding..
Shuai Wang 0011, Liang Ding 0006, Yibing Zhan, Yong Luo 0002, Shuai Liu 0002, Weiping Ding 0001
IEEE Trans. Fuzzy Syst.2
2025 SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image Classification
abstract
Recently, Mix-style data augmentation methods (e.g., Mixup and CutMix) have shown promising performance in various visual tasks. However, these methods are primarily designed for single-label images, ignoring the considerable discrepancies between single- and multi-label images,i.e., a multi-label image involves multiple co-occurred categories and fickle object scales. On the other hand, previous multi-label image classification (MLIC) methods tend to design elaborate models, bringing expensive computation. In this article, we introduce a simple but effective augmentation strategy for multi-label image classification, namely SpliceMix. The “splice” in our method is two-fold:1)Each mixed image is a splice of several downsampled images in the form of a grid, where the semantics of images attending to mixing are blended without object deficiencies for alleviating co-occurred bias;2)We splice mixed images and the original mini-batch to form a new SpliceMixed mini-batch, which allows an image with different scales to contribute to training together. Furthermore, such splice in our SpliceMixed mini-batch enables interactions between mixed images and original regular images. We also provide a simple and non-parametric extension based on consistency learning (SpliceMix-CL) to show the potential of extending our SpliceMix. Extensive experiments on various tasks demonstrate that only using SpliceMix with a baseline model (e.g., ResNet) achieves better performance than state-of-the-art methods. Moreover, the generalizability of our SpliceMix is further validated by the improvements in current MLIC methods when married with our SpliceMix.
Lei Wang 0095, Yibing Zhan, Leilei Ma 0002, Dapeng Tao, Liang Ding 0006, Chen Gong 0002
IEEE Trans. Multim.5
2024 Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion Models
abstract
Denoising Diffusion Probabilistic Models (DDPMs) have achieved significant success in generation tasks. Nevertheless, the exposure bias issue, i.e., the natural discrepancy between the training (the output of each step is calculated individually by a given input) and inference (the output of each step is calculated based on the input iteratively obtained based on the model), harms the performance of DDPMs. To our knowledge, few works have tried to tackle this issue by modifying the training process for DDPMs, but they still perform unsatisfactorily due to 1) partially modeling the discrepancy and 2) ignoring the prediction error accumulation. To address the above issues, in this paper, we propose a multi-step denoising scheduled sampling (MDSS) strategy to alleviate the exposure bias for DDPMs. Analyzing the formulations of the training and inference of DDPMs, MDSS 1) comprehensively considers the discrepancy influence of prediction errors on the output of the model (the Gaussian noise) and the output of the step (the calculated input signal of the next step), and 2) efficiently models the prediction error accumulation by using multiple iterations of a mathematical formulation initialized from one-step prediction error obtained from the model. The experimental results, compared with previous works, demonstrate that our approach is more effective in mitigating exposure bias in DDPM, DDIM, and DPM-solver. In particular, MDSS achieves an FID score of 3.86 in 100 sample steps of DDIM on the CIFAR-10 dataset, whereas the second best obtains 4.78. The code will be available on GitHub.
Zhiyao Ren, Yibing Zhan, Liang Ding 0006, Gaoang Wang, Zhongyi Fan, Dacheng Tao
AAAI3
2024 POMP: Probability-driven Meta-graph Prompter for LLMs in Low-resource Unsupervised Neural Machine Translation
abstract
Shilong Pan, Zhiliang Tian, Liang Ding, Haoqi Zheng, Zhen Huang, Zhihua Wen, Dongsheng Li. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Shilong Pan, Zhiliang Tian, Liang Ding 0006, Haoqi Zheng, Zhen Huang 0006, Zhihua Wen, Dongsheng Li 0001
ACL (1)3
2024 Revisiting Demonstration Selection Strategies in In-Context Learning
abstract
Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, Dacheng Tao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Keqin Peng, Liang Ding 0006, Yancheng Yuan, Xuebo Liu 0002, Min Zhang 0005, Yuanxin Ouyang, Dacheng Tao
ACL (1)2
2024 Uncertainty Aware Learning for Language Model Alignment
abstract
As instruction-tuned large language models (LLMs) evolve, aligning pretrained foundation models presents increasing challenges.Existing alignment strategies, which typically leverage diverse and high-quality data sources, often overlook the intrinsic uncertainty of tasks, learning all data samples equally.This may lead to suboptimal data efficiency and model performance.In response, we propose uncertainty-aware learning (UAL) to improve the model alignment of different task scenarios, by introducing the sample uncertainty (elicited from more capable LLMs).We implement UAL by a simple fashion -adaptively setting the label smoothing value of training according to the uncertainty of individual samples.Analysis shows that our UAL indeed facilitates better token clustering in the feature space, validating our hypothesis.Extensive experiments on widely used benchmarks demonstrate that our UAL significantly and consistently outperforms standard supervised fine-tuning.Notably, LLMs aligned in a mixed scenario have achieved an average improvement of 10.62% on high-entropy tasks (i.e., AlpacaEval leaderboard), and 1.81% on complex low-entropy tasks (i.e., MetaMath and GSM8K).
Yikun Wang 0001, Liang Ding 0006, Qi Zhang 0001, Dahua Lin, Dacheng Tao
ACL (1)3
2024 Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation
abstract
End-to-end speech translation (ST) presents notable disambiguation challenges as it necessitates simultaneous cross-modal and crosslingual transformations.While word sense disambiguation is an extensively investigated topic in textual machine translation, the exploration of disambiguation strategies for ST models remains limited.Addressing this gap, this paper introduces the concept of speech sense disambiguation (SSD), specifically emphasizing homophones -words pronounced identically but with different meanings.To facilitate this, we first create a comprehensive homophone dictionary and an annotated dataset rich with homophone information established based on speech-text alignment.Building on this unique dictionary, we introduce AmbigST, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy.Our experiments on different MuST-C and CoVoST ST benchmarks demonstrate that AmbigST sets new performance standards.Specifically, it achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity.
Tengfei Yu, Xuebo Liu 0002, Liang Ding 0006, Kehai Chen, Dacheng Tao, Min Zhang 0005
ACL (1)3
2024 Revisiting Knowledge Distillation for Autoregressive Language Models
abstract
Knowledge distillation (KD) is a common approach to compress a teacher model to reduce its inference cost and memory footprint, by training a smaller student model.However, in the context of autoregressive language models (LMs), we empirically find that larger teachers might dramatically result in a poorer student.In response to this problem, we conduct a series of analyses and reveal that different tokens have different teaching modes, neglecting which will lead to performance degradation.Motivated by this, we propose a simple yet effective adaptive teaching approach (ATKD) to improve the KD.The core of ATKD is to reduce rote learning and make teaching more diverse and flexible.Extensive experiments on 8 LM tasks show that, with the help of ATKD, various baseline KD methods can achieve consistent and significant performance gains (up to +3.04% average score) across all model types and sizes.More encouragingly, ATKD can improve the student model generalization effectively.
Qihuang Zhong, Liang Ding 0006, Li Shen 0008, Juhua Liu, Bo Du 0001, Dacheng Tao
ACL (1)2
2024 3AM: An Ambiguity-Aware Multi-Modal Machine Translation Dataset
abstract
Multimodal machine translation (MMT) is a challenging task that seeks to improve translation quality by incorporating visual information. However, recent studies have indicated that the visual information provided by existing MMT datasets is insufficient, causing models to disregard it and overestimate their capabilities. This issue presents a significant obstacle to the development of MMT research. This paper presents a novel solution to this issue by introducing 3AM, an ambiguity-aware MMT dataset comprising 26,000 parallel sentence pairs in English and Chinese, each with corresponding images. Our dataset is specifically designed to include more ambiguity and a greater variety of both captions and images than other MMT datasets. We utilize a word sense disambiguation model to select ambiguous data from vision-and-language datasets, resulting in a more challenging dataset. We further benchmark several state-of-the-art MMT models on our proposed dataset. Experimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets. Our work provides a valuable resource for researchers in the field of multimodal learning and encourages further exploration in this area. The data, code and scripts are freely available at https://github.com/MaxyLee/3AM.
Xuebo Liu 0002, Derek F. Wong, Jun Rao, Liang Ding 0006, Lidia S. Chao, Dacheng Tao, Min Zhang 0005
LREC/COLING6
2024 Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
abstract
Recent research shows that pre-trained language models (PLMs) suffer from “prompt bias” in factual knowledge extraction, i.e., prompts tend to introduce biases toward specific labels. Prompt bias presents a significant challenge in assessing the factual knowledge within PLMs. Therefore, this paper aims to improve the reliability of existing benchmarks by thoroughly investigating and mitigating prompt bias. We show that: 1) all prompts in the experiments exhibit non-negligible bias, with gradient-based prompts like AutoPrompt and OptiPrompt displaying significantly higher levels of bias; 2) prompt bias can amplify benchmark accuracy unreasonably by overfitting the test datasets, especially on imbalanced datasets like LAMA. Based on these findings, we propose a representation-based approach to mitigate the prompt bias during inference time. Specifically, we first estimate the biased representation using prompt-only querying, and then remove it from the model’s internal representations to generate the debiased representations, which are used to produce the final debiased outputs. Experiments across various prompts, PLMs, and benchmarks show that our approach can not only correct the overfitted performance caused by prompt bias, but also significantly improve the prompt retrieval capability (up to 10% absolute performance gain). These results indicate that our approach effectively alleviates prompt bias in knowledge evaluation, thereby enhancing the reliability of benchmark assessments. Hopefully, our plug-and-play approach can be a golden standard to strengthen PLMs toward reliable knowledge bases. Code and data are released in https://github.com/FelliYang/PromptBias.
Keqin Peng, Liang Ding 0006, Dacheng Tao, Xiliang Lu
LREC/COLING3
2024 Sheared Backpropagation for Fine-Tuning Foundation Models
abstract
Fine-tuning is the process of extending the training of pre-trained models on specific target tasks, thereby significantly enhancing their performance across various applications. However, fine-tuning often demands large memory consumption, posing a challenge for low-memory devices that some previous memory-efficient fine-tuning methods attempted to mitigate by pruning activations for gradient computation, albeit at the cost of significant computational overhead from the pruning processes during training. To address these challenges, we introduce PreBackRazor; a novel activation pruning scheme offering both computational and memory efficiency through a sparsified back-propagation strategy, which judiciously avoids unnecessary activation pruning and storage and gradient computation. Before activation pruning, our approach samples a probability of selecting a portion of parameters to freeze, utilizing a bandit method for updates to prioritize impactful gradients on convergence. During the feed-forward pass, each model layer adjusts adaptively based on parameter activation status, obviating the need for sparsification and storage of redundant activations for subsequent backpropagation. Benchmarking on fine-tuning foundation models, our approach maintains baseline accuracy across diverse tasks, yielding over 20% speedup and around 10% memory reduction. Moreover, integrating with an advanced CUDA kernel achieves up to 60% speedup without extra memory costs or accuracy loss, significantly enhancing the efficiency of fine-tuning foundation models on memory-constrained devices.
Zhiyuan Yu 0004, Li Shen 0008, Liang Ding 0006, Xinmei Tian 0001, Yixin Chen 0001, Dacheng Tao
CVPR3
2024 Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
abstract
The Mixture of Experts (MoE) has emerged as a highly successful technique in deep learning, based on the principle of divide-and-conquer to maximize model capacity without significant additional computational cost. Even in the era of large-scale language models (LLMs), MoE continues to play a crucial role, as some researchers have indicated that GPT-4 adopts the MoE structure to ensure diverse inference results. However, MoE is susceptible to performance degeneracy, particularly evident in the issues of imbalance and homogeneous representation among experts. While previous studies have extensively addressed the problem of imbalance, the challenge of homogeneous representation remains unresolved. In this study, we shed light on the homogeneous representation problem, wherein experts in the MoE fail to specialize and lack diversity, leading to frustratingly high similarities in their representations (up to 99% in a well-performed MoE model). This problem restricts the expressive power of the MoE and, we argue, contradicts its original intention. To tackle this issue, we propose a straightforward yet highly effective solution: OMoE, an orthogonal expert optimizer. Additionally, we introduce an alternating training strategy that encourages each expert to update in a direction orthogonal to the subspace spanned by other experts. Our algorithm facilitates MoE training in two key ways: firstly, it explicitly enhances representation diversity, and secondly, it implicitly fosters interaction between experts during orthogonal weights computation. Through extensive experiments, we demonstrate that our proposed optimization algorithm significantly improves the performance of fine-tuning the MoE model on the GLUE benchmark, SuperGLUE benchmark, question-answering task, and name entity recognition tasks.
Boan Liu, Liang Ding 0006, Li Shen 0008, Keqin Peng, Yu Cao 0014, Dazhao Cheng, Dacheng Tao
ECAI2
2024 Context-aware Watermark with Semantic Balanced Green-red Lists for Large Language Models
abstract
Watermarking enables people to determine whether the text is generated by a specific model.It injects a unique signature based on the "green-red" list that can be tracked during detection, where the words in green lists are encouraged to be generated.Recent researchers propose to fix the green/red lists or increase the proportion of green tokens to defend against paraphrasing attacks.However, these methods cause degradation of text quality due to semantic disparities between the watermarked text and the unwatermarked text.In this paper, we propose a semantic-aware watermark method that considers contexts to generate a semantic-aware key to split a semantically balanced green/red list for watermark injection.The semantic balanced list reduces the performance drop due to adding bias on green lists.To defend against paraphrasing attacks, we generate the watermark key considering the semantics of contexts via locally sensitive hashing.To improve the text quality, we propose to split green/red lists considering semantics to enable the green list to cover almost all semantics.We also dynamically adapt the bias to balance text quality and robustness.The experiments show our advantages in both robustness and text quality comparable to existing baselines.
Zhiliang Tian, Yiping Song, Tianlun Liu, Liang Ding 0006, Dongsheng Li 0001
EMNLP5
2024 Self-Powered LLM Modality Expansion for Large Speech-Text Models
abstract
Large language models (LLMs) exhibit remarkable performance across diverse tasks, indicating their potential for expansion into large speech-text models (LSMs) by integrating speech capabilities.Although unified speech-text pre-training and multimodal data instruction-tuning offer considerable benefits, these methods generally entail significant resource demands and tend to overfit specific tasks.This study aims to refine the use of speech datasets for LSM training by addressing the limitations of vanilla instruction tuning.We explore the instruction-following dynamics within LSMs, identifying a critical issue termed speech anchor bias-a tendency for LSMs to over-rely on speech inputs, mistakenly interpreting the entire speech modality as directives, thereby neglecting textual instructions.To counteract this bias, we introduce a self-powered LSM that leverages augmented automatic speech recognition data generated by the model itself for more effective instruction tuning.Our experiments across a range of speech-based tasks demonstrate that selfpowered LSM mitigates speech anchor bias and improves the fusion of speech and text modalities in LSMs.
Tengfei Yu, Xuebo Liu 0002, Zhiyi Hou, Liang Ding 0006, Dacheng Tao, Min Zhang 0005
EMNLP4
2024 WisdoM: Improving Multimodal Sentiment Analysis by Fusing Contextual World Knowledge
Liang Ding 0006, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dacheng Tao
ACM Multimedia2
2024 InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
abstract
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models with human values, reward hacking, also termed reward overoptimization, remains a critical challenge. This issue primarily arises from reward misgeneralization, where reward models (RMs) compute reward using spurious features that are irrelevant to human preferences. In this work, we tackle this problem from an information-theoretic perspective and propose a framework for reward modeling, namely InfoRM, by introducing a variational information bottleneck objective to filter out irrelevant information. Notably, we further identify a correlation between overoptimization and outliers in the IB latent space of InfoRM, establishing it as a promising tool for detecting reward overoptimization. Inspired by this finding, we propose the Cluster Separation Index (CSI), which quantifies deviations in the IB latent space, as an indicator of reward overoptimization to facilitate the development of online mitigation strategies. Extensive experiments on a wide range of settings and RM scales (70M, 440M, 1.4B, and 7B) demonstrate the effectiveness of InfoRM. Further analyses reveal that InfoRM's overoptimization detection mechanism is not only effective but also robust across a broad range of datasets, signifying a notable advancement in the field of RLHF. The code will be released upon acceptance.
Yuchun Miao, Sen Zhang 0006, Liang Ding 0006, Rong Bao, Lefei Zhang, Dacheng Tao
NeurIPS3
2024 Exploring sparsity in graph transformers
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001, Bo Du 0001
Neural Networks4
2024 AdaSAM: Boosting sharpness-aware minimization with adaptive learning rate and momentum for training deep neural networks
Hao Sun 0019, Li Shen 0008, Qihuang Zhong, Liang Ding 0006, Shixiang Chen, Jingwei Sun 0001, Jing Li 0047, Guangzhong Sun, Dacheng Tao
Neural Networks4
2024 Free-Form Composition Networks for Egocentric Action Recognition
abstract
Egocentric action recognition is gaining significant attention in the field of human action recognition. In this paper, we address data scarcity issue in egocentric action recognition from a compositional generalization perspective. To tackle this problem, we propose a free-form composition network (FFCN) that can simultaneously learn disentangled verb, preposition, and noun representations, and then use them to compose new samples in the feature space for rare classes of action videos. First, we use a graph to capture the spatial-temporal relations among different hand/object instances in each action video. We thus decompose each action into a set of verb and preposition spatial-temporal representations using the edge features in the graph. The temporal decomposition extracts verb and preposition representations from different video frames, while the spatial decomposition adaptively learns verb and preposition representations from action-related instances in each frame. With these spatial-temporal representations of verbs and prepositions, we can compose new samples for those rare classes in a free-form manner, which is not restricted to a rigid form of a verb and a noun. The proposed FFCN can directly generate new training data samples for rare classes, hence significantly improve action recognition performance. We evaluated our method on three popular egocentric action recognition datasets, Something-Something V2, H2O, and EPIC-KITCHENS-100, and the experimental results demonstrate the effectiveness of the proposed method for handling data scarcity problems, including long-tailed and few-shot egocentric action recognition.
Haoran Wang 0001, Qinghua Cheng, Baosheng Yu, Yibing Zhan, Dapeng Tao, Liang Ding 0006, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.6
2024 PanDa: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
abstract
Prompt Transfer (PoT) is a recently-proposed approach to improve prompt-tuning, by initializing the target prompt with the existing prompt trained on similar source tasks. However, such a vanilla PoT approach usually achieves sub-optimal performance, as (i) the PoT is sensitive to the similarity of source-target pair and (ii) directly fine-tuning the prompt initialized with source prompt on target task might lead to forgetting of the useful general knowledge learned from source task. To tackle these issues, we propose a new metric to accurately predict the prompt transferability (regarding (i)), and a novel PoT approach (namelyPanDa) that leverages the knowledge distillation technique to alleviate the knowledge forgetting effectively (regarding (ii)). Extensive and systematic experiments on 189 combinations of 21 source and 9 target datasets across 5 scales of PLMs demonstrate that: 1)our proposed metric works well to predict the prompt transferability; 2)ourPanDaconsistently outperforms the vanilla PoT approach by 2.3% average score (up to 24.1%) among all tasks and model sizes; 3)with ourPanDaapproach, prompt-tuning can achieve competitive and even better performance than model-tuning in various PLM scales scenarios.
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao
IEEE Trans. Knowl. Data Eng.2
2024 E2S2: Encoding-Enhanced Sequence-to-Sequence Pretraining for Language Understanding and Generation
abstract
Sequence-to-sequence (seq2seq) learning is a popular fashion for large-scale pretraining language models. However, the previous seq2seq pretraining models generally focus on reconstructive objectives on the decoder side and neglect the effect of encoder-side supervision, which we argue may lead to sub-optimal performance. To verify our hypothesis, we first empirically study the functionalities of the encoder and decoder in seq2seq pretrained language models, and find that the encoder takes an important but under-exploitation role than the decoder regarding the downstream performance and neuron activation. Therefore, we propose an encoding-enhanced seq2seq pretraining strategy, namelyE2S2, which improves the seq2seq models via integrating more efficient self-supervised information into the encoders. Specifically, E2S2 adopts two self-supervised objectives on the encoder side from two aspects: 1) locally denoising the corrupted sentence (denoising objective); and 2) globally learning better sentence representations (contrastive objective). With the help of both objectives, the encoder can effectively distinguish the noise tokens and capture high-level (i.e., syntactic and semantic) knowledge, thus strengthening the ability of seq2seq model to accurately achieve the conditional generation. On a large diversity of downstream natural language understanding and generation tasks, E2S2 dominantly improves the performance of its powerful backbone models, e.g., BART and T5. For example, upon BART backbone, we achieve +1.1% averaged gain on the general language understanding evaluation (GLUE) benchmark and +1.75%$F_{0.5}$score improvement on CoNLL2014 dataset. We also provide in-depth analyses to show the improvement stems from better linguistic representation. We hope that our work will foster future self-supervision research on seq2seq language model pretraining.
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao
IEEE Trans. Knowl. Data Eng.2
2024 Parameter-Efficient and Student-Friendly Knowledge Distillation
abstract
Pre-trained models are frequently employed in multimodal learning. However, these models have too many parameters and need too much effort to fine-tune the downstream tasks. Knowledge distillation (KD) is a method to transfer knowledge using the soft label from this pre-trained teacher model to a smaller student, where the parameters of the teacher are fixed (or partially) during training. Recent studies show that this mode may cause difficulties in knowledge transfer due to the mismatched model capacities. To alleviate the mismatch problem, adjustment of temperature parameters, label smoothing and teacher-student joint training methods (online distillation) to smooth the soft label of a teacher network, have been proposed. But those methods rarely explain the effect of smoothed soft labels to enhance the KD performance. The main contributions of our work are the discovery, analysis, and validation of the effect of the smoothed soft label and a less time-consuming and adaptive transfer of the pre-trained teacher's knowledge method, namely PESF-KD by adaptive tuning soft labels of the teacher network. Technically, we first mathematically formulate the mismatch as the sharpness gap between teacher's and student's predictive distributions, where we show such a gap can be narrowed with the appropriate smoothness of the soft label. Then, we introduce an adapter module for the teacher and only update the adapter to obtain soft labels with appropriate smoothness. Experiments on various benchmarks including CV and NLP show that PESF-KD can significantly reduce the training cost while obtaining competitive results compared to advanced online distillation methods.
Jun Rao, Xv Meng, Liang Ding 0006, Shuhan Qi, Xuebo Liu 0002, Min Zhang 0005, Dacheng Tao
IEEE Trans. Multim.3
2024 Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural Networks
abstract
Graph neural networks (GNNs) tend to suffer from high computation costs due to the exponentially increasing scale of graph data and a large number of model parameters, which restricts their utility in practical applications. To this end, some recent works focus on sparsifying GNNs (including graph structures and model parameters) with the lottery ticket hypothesis (LTH) to reduce inference costs while maintaining performance levels. However, the LTH-based methods suffer from two major drawbacks: 1) they require exhaustive and iterative training of dense models, resulting in an extremely large training computation cost, and 2) they only trim graph structures and model parameters but ignore the node feature dimension, where vast redundancy exists. To overcome the above limitations, we propose a comprehensive graph gradual pruning framework termed CGP. This is achieved by designing a during-training graph pruning paradigm to dynamically prune GNNs within one training process. Unlike LTH-based methods, the proposed CGP approach requires no retraining, which significantly reduces the computation costs. Furthermore, we design a cosparsifying strategy to comprehensively trim all the three core elements of GNNs: graph structures, node features, and model parameters. Next, to refine the pruning operation, we introduce a regrowth process into our CGP framework, to reestablish the pruned but important connections. The proposed CGP is evaluated over a node classification task across six GNN architectures, including shallow models [graph convolutional network (GCN) and graph attention network (GAT)], shallow-but-deep-propagation models [simple graph convolution (SGC) and approximate personalized propagation of neural predictions (APPNP)], and deep models [GCN via initial residual and identity mapping (GCNII) and residual GCN (ResGCN)], on a total of 14 real-world graph datasets, including large-scale graph datasets from the challenging Open Graph Benchmark (OGB). Experiments reveal that the proposed strategy greatly improves both training and inference efficiency while matching or even exceeding the accuracy of the existing methods.
Chuang Liu 0008, Xueqi Ma, Yibing Zhan, Liang Ding 0006, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001, Danilo P. Mandic
IEEE Trans. Neural Networks Learn. Syst.4
2024 Can Linguistic Knowledge Improve Multimodal Alignment in Vision-Language Pretraining?
abstract
The field of multimedia research has witnessed significant interest in leveraging multimodal pretrained neural network models to perceive and represent the physical world. Among these models, vision-language pretraining (VLP) has emerged as a captivating topic. Currently, the prevalent approach in VLP involves supervising the training process with paired image-text data. However, limited efforts have been dedicated to exploring the extraction of essential linguistic knowledge, such as semantics and syntax, during VLP and understanding its impact on multimodal alignment. In response, our study aims to shed light on the influence of comprehensive linguistic knowledge encompassing semantic expression and syntactic structure on multimodal alignment. To achieve this, we introduce SNARE , a large-scale multimodal alignment probing benchmark designed specifically for the detection of vital linguistic components, including lexical, semantic, and syntax knowledge. SNARE offers four distinct tasks: Semantic Structure, Negation Logic, Attribute Ownership, and Relationship Composition. Leveraging SNARE , we conduct holistic analyses of six advanced VLP models (BLIP, CLIP, Flava, X-VLM, BLIP2, and GPT-4), along with human performance, revealing key characteristics of the VLP model: (i) Insensitivity to complex syntax structures, relying primarily on content words for sentence comprehension. (ii) Limited comprehension of sentence combinations and negations. (iii) Challenges in determining actions or spatial relations within visual information, as well as difficulties in verifying the correctness of ternary relationships. Based on these findings, we propose the following strategies to enhance multimodal alignment in VLP: (1) Utilize a large generative language model as the language backbone in VLP to facilitate the understanding of complex sentences. (2) Establish high-quality datasets that emphasize content words and employ simple syntax, such as short-distance semantic composition, to improve multimodal alignment. (3) Incorporate more fine-grained visual knowledge, such as spatial relationships, into pretraining objectives. 1
Fei Wang 0032, Liang Ding 0006, Jun Rao, Ye Liu 0014, Li Shen 0008, Changxing Ding
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Improving Simultaneous Machine Translation with Monolingual Data
abstract
Simultaneous machine translation (SiMT) is usually done via sequence-level knowledge distillation (Seq-KD) from a full-sentence neural machine translation (NMT) model. However, there is still a significant performance gap between NMT and SiMT. In this work, we propose to leverage monolingual data to improve SiMT, which trains a SiMT student on the combination of bilingual data and external monolingual data distilled by Seq-KD. Preliminary experiments on En-Zh and En-Ja news domain corpora demonstrate that monolingual data can significantly improve translation quality (e.g., +3.15 BLEU on En-Zh). Inspired by the behavior of human simultaneous interpreters, we propose a novel monolingual sampling strategy for SiMT, considering both chunk length and monotonicity. Experimental results show that our sampling strategy consistently outperforms the random sampling strategy (and other conventional typical NMT monolingual sampling strategies) by avoiding the key problem of SiMT -- hallucination, and has better scalability. We achieve +0.72 BLEU improvements on average against random sampling on En-Zh and En-Ja. Data and codes can be found at https://github.com/hexuandeng/Mono4SiMT.
Hexuan Deng, Liang Ding 0006, Xuebo Liu 0002, Meishan Zhang, Dacheng Tao, Min Zhang 0005
AAAI2
2023 CASN: Class-Aware Score Network for Textual Adversarial Detection
abstract
Adversarial detection aims to detect adversarial samples that threaten the security of deep neural networks, which is an essential step toward building robust AI systems.Density-based estimation is widely considered as an effective technique by explicitly modeling the distribution of normal data and identifying adversarial ones as outliers.However, these methods suffer from significant performance degradation when the adversarial samples lie close to the non-adversarial data manifold.To address this limitation, we propose a score-based generative method to implicitly model the data distribution.Our approach utilizes the gradient of the log-density data distribution and calculates the distribution gap between adversarial and normal samples through multi-step iterations using Langevin dynamics.In addition, we use supervised contrastive learning to guide the gradient estimation using label information, which avoids collapsing to a single data manifold and better preserves the anisotropy of the different labeled data distributions.Experimental results on three text classification tasks upon four advanced attack algorithms show that our approach is a significant improvement (+15.2F1 score on average against previous SOTA) over previous detection methods.
Rong Bao, Liang Ding 0006, Qi Zhang 0001, Dacheng Tao
ACL (1)3
2023 PAD-Net: An Efficient Framework for Dynamic Networks
abstract
Dynamic networks, e.g., Dynamic Convolution (DY-Conv) and the Mixture of Experts (MoE), have been extensively explored as they can considerably improve the model's representation power with acceptable computational cost.The common practice in implementing dynamic networks is to convert the given static layers into fully dynamic ones where all parameters are dynamic (at least within a single layer) and vary with the input.However, such a fully dynamic setting may cause redundant parameters and high deployment costs, limiting the applicability of dynamic networks to a broader range of tasks and models.The main contributions of our work are challenging the basic commonsense in dynamic networks and proposing a partially dynamic network, namely PAD-Net, to transform the redundant dynamic parameters into static ones.Also, we further design Iterative Mode Partition to partition dynamic and static parameters efficiently.Our method is comprehensively supported by large-scale experiments with two typical advanced dynamic architectures, i.e., DY-Conv and MoE, on both image classification and GLUE benchmarks.Encouragingly, we surpass the fully dynamic networks by +0.7% top-1 acc with only 30% dynamic parameters for ResNet-50 and +1.9% average score in language understanding with only 50% dynamic parameters for BERT.Code will be released
Shwai He, Liang Ding 0006, Daize Dong, Boan Liu, Fuqiang Yu, Dacheng Tao
ACL (1)2
2023 Toward Human-Like Evaluation for Natural Language Generation with Error Analysis
abstract
The pretrained language model (PLM) based metrics have been successfully used in evaluating language generation tasks.Recent studies of the human evaluation community show that considering both major errors (e.g.mistranslated tokens) and minor errors (e.g.imperfections in fluency) can produce high-quality judgments.This inspires us to approach the final goal of the automatic metrics (human-like evaluations) by fine-grained error analysis.In this paper, we argue that the ability to estimate sentence confidence is the tip of the iceberg for PLM-based metrics.And it can be used to refine the generated sentence toward higher confidence and more reference-grounded, where the costs of refining and approaching reference are used to determine the major and minor errors, respectively.To this end, we take BARTScore as the testbed and present an innovative solution to marry the unexploited sentence refining capacity of BARTScore and human-like error analysis, where the final score consists of both the evaluations of major and minor errors.Experiments show that our solution consistently improves BARTScore, outperforming top-scoring metrics in 19/25 test settings.Analyses demonstrate our method robustly and efficiently approaches human-like evaluations, enjoying better interpretability.Our code and scripts will be publicly released in https: //github.com/Coldmist-Lu/ ErrorAnalysis_NLGEvaluation.
Qingyu Lu 0001, Liang Ding 0006, Kan-Jian Zhang, Derek F. Wong, Dacheng Tao
ACL (1)2
2023 Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State Tracking
abstract
Qingyue Wang, Liang Ding, Yanan Cao, Yibing Zhan, Zheng Lin, Shi Wang, Dacheng Tao, Li Guo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qingyue Wang, Liang Ding 0006, Yanan Cao 0001, Yibing Zhan, Zheng Lin 0001, Shi Wang 0002, Dacheng Tao, Li Guo 0001
ACL (1)2
2023 Revisiting Token Dropping Strategy in Efficient BERT Pretraining
abstract
Token dropping is a recently-proposed strategy to speed up the pretraining of masked language models, such as BERT, by skipping the computation of a subset of the input tokens at several middle layers.It can effectively reduce the training time without degrading much performance on downstream tasks.However, we empirically find that token dropping is prone to a semantic loss problem and falls short in handling semantic-intense tasks ( §2).Motivated by this, we propose a simple yet effective semantic-consistent learning method (SCTD) to improve the token dropping.SCTD aims to encourage the model to learn how to preserve the semantic information in the representation space.Extensive experiments on 12 tasks show that, with the help of our SCTD, token dropping can achieve consistent and significant performance gains across all task types and model sizes.More encouragingly, SCTD saves up to 57% of pretraining time and brings up to +1.56% average improvement over the vanilla token dropping.
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Xuebo Liu 0002, Min Zhang 0005, Bo Du 0001, Dacheng Tao
ACL (1)2
2023 Using Self-Supervised Dual Constraint Contrastive Learning for Cross-Modal Retrieval
abstract
In this work, we present a self-supervised dual constraint contrastive method for efficiently fine-tuning the vision-language pre-trained (VLP) models that have achieved great success on various cross-modal tasks, since full fine-tune these pre-trained models is computationally expensive and tend to result in catastrophic forgetting restricted by the size and quality of labeled datasets. Our approach freezes the pre-trained VLP models as the fundamental, generalized, and transferable multimodal representation and incorporates lightweight parameters to learn domain and task-specific features without labeled data. We demonstrated that our self-supervised dual contrastive model performs better than previous fine-tuning methods on MS COCO and Flickr 30K datasets on the cross-modal retrieval task, with an even more pronounced improvement in zero-shot performance. Furthermore, experiments on the MOTIF dataset prove that our self-supervised approach remains effective when trained on a small, out-of-domain dataset without overfitting. As a plug-and-play method, our proposed method is agnostic to the underlying models and can be easily integrated with different VLP models, allowing for the potential incorporation of future advancements in VLP models.
Xintong Wang 0001, Liang Ding 0006, Sanyuan Zhao, Chris Biemann
ECAI3
2023 Merging Experts into One: Improving Computational Efficiency of Mixture of Experts
abstract
Scaling the size of language models usually leads to remarkable advancements in NLP tasks.But it often comes with a price of growing computational cost.Although a sparse Mixture of Experts (MoE) can reduce the cost by activating a small subset of parameters (e.g., one expert) for each input, its computation escalates significantly if increasing the number of activated experts, limiting its practical utility.Can we retain the advantages of adding more experts without substantially increasing the computational costs?In this paper, we first demonstrate the superiority of selecting multiple experts and then propose a computation-efficient approach called Merging Experts into One (MEO), which reduces the computation cost to that of a single expert.Extensive experiments show that MEO significantly improves computational efficiency, e.g., FLOPS drops from 72.0G of vanilla MoE to 28.9G (MEO).Moreover, we propose a token-level attention block that further enhances the efficiency and performance of token-level MEO, e.g., 83.3% (MEO) vs. 82.6%(vanilla MoE) average score on the GLUE benchmark.Our code will be released upon acceptance.Code will be released at
Shwai He, Run-Ze Fan, Liang Ding 0006, Li Shen 0008, Tianyi Zhou 0001, Dacheng Tao
EMNLP3
2023 PromptST: Abstract Prompt Learning for End-to-End Speech Translation
abstract
An end-to-end speech-to-text (S2T) translation model is usually initialized from a pretrained speech recognition encoder and a pretrained text-to-text (T2T) translation decoder.Although this straightforward setting has been shown empirically successful, there do not exist clear answers to the research questions: 1) how are speech and text modalities fused in S2T model and 2) how to better fuse the two modalities?In this paper, we take the first step toward understanding the fusion of speech and text features in S2T model.We first design and release a 10GB linguistic probing benchmark, namely Speech-Senteval, to investigate the acoustic and linguistic behaviors of S2T models.Preliminary analysis reveals that the uppermost encoder layers of the S2T model can not learn linguistic knowledge efficiently, which is crucial for accurate translation.Based on the finding, we further propose a simple plug-in prompt-learning strategy on the uppermost encoder layers to broaden the abstract representation power of the encoder of S2T models.We call such a promptenhanced S2T model PromptST.Experimental results on four widely-used S2T datasets show that PromptST can deliver significant improvements over a strong baseline by capturing richer linguistic knowledge.Benchmarks,
Tengfei Yu, Liang Ding 0006, Xuebo Liu 0002, Kehai Chen, Meishan Zhang, Dacheng Tao, Min Zhang 0005
EMNLP2
2023 Self-Evolution Learning for Mixup: Enhance Data Augmentation on Few-Shot Text Classification Tasks
abstract
Text classification tasks often encounter fewshot scenarios with limited labeled data, and addressing data scarcity is crucial.Data augmentation with mixup merges sample pairs to generate new pseudos, which can relieve the data deficiency issue in text classification.However, the quality of pseudo-samples generated by mixup exhibits significant variations.Most of the mixup methods fail to consider the varying degree of learning difficulty in different stages of training.And mixup generates new samples with one-hot labels, which encourages the model to produce a high prediction score for the correct class that is much larger than other classes, resulting in the model's over-confidence.In this paper, we propose a self-evolution learning (SE) based mixup approach for data augmentation in text classification, which can generate more adaptive and model-friendly pseudo samples for the model training.SE caters to the growth of the model learning ability and adapts to the ability when generating training samples.To alleviate the model over-confidence, we introduce an instance-specific label smoothing regularization approach, which linearly interpolates the model's output and one-hot labels of the original samples to generate new soft labels for label mixing up.Through experimental analysis, experiments show that our SE brings consistent and significant improvements upon different mixup methods.In-depth analyses demonstrate that SE enhances the model's generalization ability.
Haoqi Zheng, Qihuang Zhong, Liang Ding 0006, Zhiliang Tian, Xin Niu 0002, Dongsheng Li 0001, Dacheng Tao
EMNLP3
2023 Zero-shot Sharpness-Aware Quantization for Pre-trained Language Models
abstract
Quantization is a promising approach for reducing memory overhead and accelerating inference, especially in large pre-trained language model (PLM) scenarios.While having no access to original training data due to security and privacy concerns has emerged the demand for zero-shot quantization.Most of the cuttingedge zero-shot quantization methods primarily ❶ apply to computer vision tasks, and ❷ neglect of overfitting problem in the generative adversarial learning process, leading to sub-optimal performance.Motivated by this, we propose a novel zero-shot sharpness-aware quantization (ZSAQ) framework for the zeroshot quantization of various PLMs.The key algorithm in solving ZSAQ is the SAM-SGA optimization, which aims to improve the quantization accuracy and model generalization via optimizing a minimax problem.We theoretically prove the convergence rate for the minimax optimization problem and this result can be applied to other nonconvex-PL minimax optimization frameworks.Extensive experiments on 11 tasks demonstrate that our method brings consistent and significant performance gains on both discriminative and generative PLMs, i.e., up to +6.98 average score.Furthermore, we empirically validate that our method can effectively improve the model generalization.
Miaoxi Zhu, Qihuang Zhong, Li Shen 0008, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao
EMNLP4
2023 FedSpeed: Larger Local Interval, Less Communication Round, and Higher Generalization Accuracy
Li Shen 0008, Tiansheng Huang, Liang Ding 0006, Dacheng Tao
ICLR4
2023 Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth Landscape
abstract
In federated learning (FL), a cluster of local clients are chaired under the coordination of the global server and cooperatively train one model with privacy protection. Due to the multiple local updates and the isolated non-iid dataset, clients are prone to overfit into their own optima, which extremely deviates from the global objective and significantly undermines the performance. Most previous works only focus on enhancing the consistency between the local and global objectives to alleviate this prejudicial client drifts from the perspective of the optimization view, whose performance would be prominently deteriorated on the high heterogeneity. In this work, we propose a novel and general algorithm FedSMOO by jointly considering the optimization and generalization targets to efficiently improve the performance in FL. Concretely, FedSMOO adopts a dynamic regularizer to guarantee the local optima towards the global objective, which is meanwhile revised by the global Sharpness Aware Minimization (SAM) optimizer to search for the consistent flat minima. Our theoretical analysis indicates that FedSMOO achieves fast $\mathcal{O}(1/T)$ convergence rate with low generalization bound. Extensive numerical studies are conducted on the real-world dataset to verify its peerless efficiency and excellent generality.
Li Shen 0008, Shixiang Chen, Liang Ding 0006, Dacheng Tao
ICML4
2023 Gapformer: Graph Transformer with Graph Pooling for Node Classification
abstract
Graph Transformers (GTs) have proved their advantage in graph-level tasks. However, existing GTs still perform unsatisfactorily on the node classification task due to 1) the overwhelming unrelated information obtained from a vast number of irrelevant distant nodes and 2) the quadratic complexity regarding the number of nodes via the fully connected attention mechanism. In this paper, we present Gapformer, a method for node classification that deeply incorporates Graph Transformer with Graph Pooling. More specifically, Gapformer coarsens the large-scale nodes of a graph into a smaller number of pooling nodes via local or global graph pooling methods, and then computes the attention solely with the pooling nodes rather than all other nodes. In such a manner, the negative influence of the overwhelming unrelated nodes is mitigated while maintaining the long-range information, and the quadratic complexity is reduced to linear complexity with respect to the fixed number of pooling nodes. Extensive experiments on 13 node classification datasets, including homophilic and heterophilic graph datasets, demonstrate the competitive performance of Gapformer over existing Graph Neural Networks and GTs.
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001
IJCAI4
2023 Prompt-Learning for Cross-Lingual Relation Extraction
abstract
Relation Extraction (RE) is a crucial task in Information Extraction, which entails predicting relationships between entities within a given sentence. However, extending pre-trained RE models to other languages is challenging, particularly in real-world scenarios where Cross-Lingual Relation Extraction (XRE) is required. Despite recent advancements in Prompt-Learning, which involves transferring knowledge from Multilingual Pre-trained Language Models (PLMs) to diverse downstream tasks, there is limited research on the effective use of multilingual PLMs with prompts to improve XRE. In this paper, we present a novel XRE algorithm based on Prompt-Tuning, referred to as Prompt-Xre. To evaluate its effectiveness, we design and implement several prompt templates, including hard, soft, and hybrid prompts, and empirically test their performance on competitive multilingual PLMs, specifically mBART. Our extensive experiments, conducted on the low-resource ACE05 benchmark across multiple languages, demonstrate that our Prompt-Xre algorithm significantly outperforms both vanilla multilingual PLMs and other existing models, achieving state-of-the-art performance in XRE. To further show the generalization of our Prompt-XRE on larger data scales, we construct and release a new XRE dataset-WMTI7-EnZh XRE, containing 0.9M English-Chinese pairs extracted from WMT 2017 parallel corpus. Experiments on WMTI7-EnZh XRE also show the effectiveness of our Prompt-XRE against other competitive baselines. The code and newly constructed dataset are freely available at httus://2ithub.com/HSU-CHIA-MING/Promut-XRE.
Chiaming Hsu, Changtong Zan, Liang Ding 0006, Longyue Wang, Weifeng Liu 0001, Wenbin Hu 0001
IJCNN3
2023 MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
abstract
Recently, Mixture-of-Experts (MoE) has become one of the most popular techniques to scale pre-trained models to extraordinarily large sizes. Dynamic activation of experts allows for conditional computation, increasing the number of parameters of neural networks, which is critical for absorbing the vast amounts of knowledge available in many deep learning areas. However, despite the existing system and algorithm optimizations, there are significant challenges to be tackled when it comes to the inefficiencies of communication and memory consumption.In this paper, we present the design and implementation of MPipeMoE, a high-performance library that accelerates MoE training with adaptive and memory-efficient pipeline parallelism. Inspired by that the MoE training procedure can be divided into multiple independent sub-stages, we design adaptive pipeline parallelism with an online algorithm to configure the granularity of the pipelining. Further, we analyze the memory footprint breakdown of MoE training and identify that activations and temporary buffers are the primary contributors to the overall memory footprint. Toward memory efficiency, we propose memory reusing strategies to reduce memory requirements by eliminating memory redundancies, and develop an adaptive selection component to determine the optimal strategy that considers both hardware capacities and model characteristics at runtime. We implement MPipeMoE upon PyTorch and evaluate it with common MoE models in a physical cluster consisting of 8 NVIDIA DGX A100 servers. Compared with the state-of-art approach, MPipeMoE achieves up to 2.8× speedup and reduces memory footprint by up to 47% in training large models.
Zheng Zhang 0036, Donglin Yang, Yaqi Xia, Liang Ding 0006, Dacheng Tao, Xiaobo Zhou 0002, Dazhao Cheng
IPDPS4
2023 SD-Conv: Towards the Parameter-Efficiency of Dynamic Convolution
abstract
Dynamic convolution achieves better performance for efficient CNNs at the cost of negligible FLOPs increase. However, the performance increase can not match the significantly expanded number of parameters, which is the main bottleneck in real-world applications. Contrastively, mask-based unstructured pruning obtains a lightweight network by removing redundancy in the heavy network. In this paper, we propose a new framework, Sparse Dynamic Convolution (SD-CONV), to naturally integrate these two paths such that it can inherit the advantage of dynamic mechanism and sparsity. We first design a binary mask derived from a learnable threshold to prune static kernels, significantly reducing the parameters and computational cost but achieving higher performance in Imagenet-1K. We further transfer pretrained models into a variety of downstream tasks, showing consistently better results than baselines. We hope our SD-Conv could be an efficient alternative to conventional dynamic convolutions.
Shwai He, Chenbo Jiang, Daize Dong, Liang Ding 0006
WACV4
2023 KE-X: Towards subgraph explanations of knowledge graph embedding based on knowledge information gain
Guojia Wan, Yibing Zhan, Zengmao Wang, Liang Ding 0006, Zhigao Zheng 0001, Bo Du 0001
Knowl. Based Syst.5
2023 Efficient Federated Learning Via Local Adaptive Amended Optimizer With Linear Speedup
abstract
Adaptive optimization has achieved notable success for distributed learning while extending adaptive optimizer to federated Learning (FL) suffers from severe inefficiency, including (i) rugged convergence due to inaccurate gradient estimation in global adaptive optimizer; (ii) client drifts exacerbated by local over-fitting with the local adaptive optimizer. In this work, we propose a novel momentum-based algorithm via utilizing the global gradient descent and locally adaptive amended optimizer to tackle these difficulties. Specifically, we incorporate a locally amended technique to the adaptive optimizer, named Federated Local ADaptive Amended optimizer (FedLADA), which estimates the global average offset in the previous communication round and corrects the local offset through a momentum-like term to further improve the empirical training speed and mitigate the heterogeneous over-fitting. Theoretically, we establish the convergence rate of FedLADA with a linear speedup property on the non-convex case under the partial participation settings. Moreover, we conduct extensive experiments on the real-world dataset to demonstrate the efficacy of our proposed FedLADA, which could greatly reduce the communication rounds and achieves higher accuracy than several baselines.
Li Shen 0008, Hao Sun 0019, Liang Ding 0006, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Unified Instance and Knowledge Alignment Pretraining for Aspect-Based Sentiment Analysis
abstract
The goal of aspect-based sentiment analysis (ABSA) is to determine the sentiment polarity towards an aspect. Because of the expensive and limited amounts of labelled data, the pretraining strategy has become the de facto standard for ABSA. However, there always exists a severe domain shift between the pretraining and downstream ABSA datasets, which hinders effective knowledge transfer when directly fine-tuning, making the downstream task suboptimal. To mitigate this domain shift, we introduce a unified alignment pretraining framework into the vanilla pretrain-finetune pipeline, that has both instance- and knowledge-level alignments. Specifically, we first devise a novel coarse-to-fine retrieval sampling approach to select target domain-related instances from the large-scale pretraining dataset, thus aligning the instances between pretraining and the target domains (First Stage). Then, we introduce a knowledge guidance-based strategy to further bridge the domain gap at the knowledge level. In practice, we formulate the model pretrained on the sampled instances into a knowledge guidance model and a learner model. On the target dataset, we design an on-the-fly teacher-student joint fine-tuning approach to progressively transfer the knowledge from the knowledge guidance model to the learner model (Second Stage). Therefore, the learner model can maintain more domain-invariant knowledge when learning new knowledge from the target dataset. In theThird Stage,the learner model is finetuned to better adapt its learned knowledge to the target dataset. Extensive experiments and analyses on several ABSA benchmarks demonstrate the effectiveness and universality of our proposed pretraining framework.
Juhua Liu, Qihuang Zhong, Liang Ding 0006, Bo Du 0001, Dacheng Tao
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Knowledge Graph Augmented Network Towards Multiview Representation Learning for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is a fine-grained task of sentiment analysis. To better comprehend long complicated sentences and obtain accurate aspect-specific information, linguistic and commonsense knowledge are generally required in this task. However, most current methods employ complicated and inefficient approaches to incorporate external knowledge, e.g., directly searching the graph nodes. Additionally, the complementarity between external knowledge and linguistic information has not been thoroughly studied. To this end, we propose a knowledge graph augmented network (KGAN), which aims to effectively incorporate external knowledge with explicitly syntactic and contextual information. In particular, KGAN captures the sentiment feature representations from multiple different perspectives,i.e., context-, syntax- and knowledge-based. First, KGAN learns the contextual and syntactic representations in parallel to fully extract the semantic features. Then, KGAN integrates the knowledge graphs into the embedding space, based on which the aspect-specific knowledge representations are further obtained via an attention mechanism. Last, we propose a hierarchical fusion module to complement these multi-view representations in alocal-to-globalmanner. Extensive experiments on five popular ABSA benchmarks demonstrate the effectiveness and robustness of our KGAN. Notably, with the help of the pretrained model of RoBERTa, KGAN achieves a new record of state-of-the-art performance among all datasets.
Qihuang Zhong, Liang Ding 0006, Juhua Liu, Bo Du 0001, Dacheng Tao
IEEE Trans. Knowl. Data Eng.2
2023 Dynamic Contrastive Distillation for Image-Text Retrieval
abstract
Although the vision-and-language pretraining (VLP) equipped cross-modal image-text retrieval (ITR) has achieved remarkable progress in the past two years, it suffers from a major drawback: the ever-increasing size of VLP models restrict its deployment to real-world search scenarios (where the high latency is unacceptable). To alleviate this problem, we present a novel plug-in dynamic contrastive distillation (DCD) framework to compress the large VLP models for the ITR task. Technically, we face the following two challenges: 1) the typical uni-modal metric learning approach is difficult to directly apply to cross-modal task, due to the limited GPU memory to optimize too many negative samples during handling cross-modal fusion features. 2) it is inefficient to static optimize the student network from different hard samples, which have different effects on distillation learning and student network optimization. We try to overcome these challenges from two points. First, to achieve multi-modal contrastive learning, and balance the training costs and effects, we propose to use a teacher network to estimate the difficult samples for students, making the students absorb the powerful knowledge from pre-trained teachers, and master the knowledge from hard samples. Second, to dynamic learn from hard sample pairs, we propose dynamic distillation to dynamically learn samples of different difficulties, from the perspective of better balancing the difficulty of knowledge and students' self-learning ability. We successfully apply our proposed DCD strategy on two state-of-the-art vision-language pretrained models, i.e. ViLT and METER. Extensive experiments on MS-COCO and Flickr 30 K benchmarks show the effectiveness and efficiency of our DCD framework. Encouragingly, we can speed up the inference at least 129 × compared to the existing ITR models. We further provide in-depth analyses and discussions that explain where the performance improvement comes from. We hope our work can shed light on other tasks that require distillation and contrastive learning.
Jun Rao, Liang Ding 0006, Shuhan Qi, Yang Liu 0039, Li Shen 0008, Dacheng Tao
IEEE Trans. Multim.2
2022 Redistributing Low-Frequency Words: Making the Most of Monolingual Data in Non-Autoregressive Translation
abstract
Knowledge distillation (KD) is the preliminary step for training non-autoregressive translation (NAT) models, which eases the training of NAT models at the cost of losing important information for translating low-frequency words.In this work, we provide an appealing alternative for NAT -monolingual KD, which trains NAT student on external monolingual data with AT teacher trained on the original bilingual data.Monolingual KD is able to transfer both the knowledge of the original bilingual data (implicitly encoded in the trained AT teacher model) and that of the new monolingual data to the NAT student model.Extensive experiments on eight WMT benchmarks over two advanced NAT models show that monolingual KD consistently outperforms the standard KD by improving lowfrequency word translation, without introducing any computational cost.Monolingual KD enjoys desirable expandability, which can be further enhanced (when given more computational budget) by combining with the standard KD, a reverse monolingual KD, or enlarging the scale of monolingual data.Extensive analyses demonstrate that these techniques can be used together profitably to further recall the useful information lost in the standard KD.Encouragingly, combining with standard KD, our approach achieves 30.4 and 34.1 BLEU points on the WMT14 English-German and German-English datasets, respectively.Our code and trained models are freely available at https://github.com/ alphadl/RLFW-NAT.mono.
Liang Ding 0006, Longyue Wang, Shuming Shi 0001, Dacheng Tao, Zhaopeng Tu
ACL (1)1
2022 A Contrastive Cross-Channel Data Augmentation Framework for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) is a fine-grained sentiment analysis task, which focuses on detecting the sentiment polarity towards the aspect in a sentence. However, it is always sensitive to the multi-aspect challenge, where features of multiple aspects in a sentence will affect each other. To mitigate this issue, we design a novel training framework, called Contrastive Cross-Channel Data Augmentation (C3 DA), which leverages an in-domain generator to construct more multi-aspect samples and then boosts the robustness of ABSA models via contrastive learning on these generated data. In practice, given a generative pretrained language model and some limited ABSA labeled data, we first employ some parameter-efficient approaches to perform the in-domain fine-tuning. Then, the obtained in-domain generator is used to generate the synthetic sentences from two channels, i.e., Aspect Augmentation Channel and Polarity Augmentation Channel, which generate the sentence condition on a given aspect and polarity respectively. Specifically, our C3 DA performs the sentence generation in a cross-channel manner to obtain more sentences, and proposes an Entropy-Minimization Filter to filter low-quality generated samples. Extensive experiments show that our C3 DA can outperform those baselines without any augmentations by about 1% on accuracy and Macro- F1. Code and data are released in https://github.com/wangbing1416/C3DA.
Bing Wang 0018, Liang Ding 0006, Qihuang Zhong, Ximing Li 0002, Dacheng Tao
COLING2
2022 On the Complementarity between Pre-Training and Random-Initialization for Resource-Rich Machine Translation
abstract
Pre-Training (PT) of text representations has been successfully applied to low-resource Neural Machine Translation (NMT). However, it usually fails to achieve notable gains (some- times, even worse) on resource-rich NMT on par with its Random-Initialization (RI) counterpart. We take the first step to investigate the complementarity between PT and RI in resource-rich scenarios via two probing analyses, and find that: 1) PT improves NOT the accuracy, but the generalization by achieving flatter loss landscapes than that of RI; 2) PT improves NOT the confidence of lexical choice, but the negative diversity by assigning smoother lexical probability distributions than that of RI. Based on these insights, we propose to combine their complementarities with a model fusion algorithm that utilizes optimal transport to align neurons between PT and RI. Experiments on two resource-rich translation benchmarks, WMT’17 English-Chinese (20M) and WMT’19 English-German (36M), show that PT and RI could be nicely complementary to each other, achieving substantial improvements considering both translation accuracy, generalization, and negative diversity. Probing tools and code are released at: https://github.com/zanchangtong/PTvsRI.
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yu Cao 0014, Weifeng Liu 0001, Dacheng Tao
COLING2
2022 Fine-tuning Global Model via Data-Free Knowledge Distillation for Non-IID Federated Learning
abstract
Federated Learning (FL) is an emerging distributed learning paradigm under privacy constraint. Data heterogeneity is one of the main challenges in FL, which results in slow convergence and degraded performance. Most existing approaches only tackle the heterogeneity challenge by restricting the local model update in client, ignoring the performance drop caused by direct global model aggregation. Instead, we propose a data-free knowledge distillation method to fine-tune the global model in the server (FedFTG), which relieves the issue of direct model aggregation. Concretely, FedFTG explores the input space of local models through a generator, and uses it to transfer the knowledge from local models to the global model. Besides, we propose a hard sample mining scheme to achieve effective knowledge distillation throughout the training. In addition, we develop customized label sampling and class-level ensemble to derive maximum utilization of knowledge, which implicitly mitigates the distribution discrepancy across clients. Extensive experiments show that our FedFTG significantly outperforms the state-of-the-art (SOTA) FL algorithms and can serve as a strong plugin for enhancing FedAvg, FedProx, FedDyn, and SCAFFOLD.
Lin Zhang 0014, Li Shen 0008, Liang Ding 0006, Dacheng Tao, Ling-Yu Duan
CVPR3
2022 Where Does the Performance Improvement Come From?: - A Reproducibility Concern about Image-Text Retrieval
abstract
This article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the last decade, image-text retrieval has steadily become a major research direction in the field of information retrieval. Numerous researchers train and evaluate image-text retrieval algorithms using benchmark datasets such as MS-COCO and Flickr30k. Research in the past has mostly focused on performance, with multiple state-of-the-art methodologies being suggested in a variety of ways. According to their assertions, these techniques provide improved modality interactions and hence more precise multimodal representations. In contrast to previous works, we focus on the reproducibility of the approaches and the examination of the elements that lead to improved performance by pretrained and nonpretrained models in retrieving images and text.
Jun Rao, Fei Wang 0032, Liang Ding 0006, Shuhan Qi, Yibing Zhan, Weifeng Liu 0001, Dacheng Tao
SIGIR3
2021 Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
abstract
Liang Ding, Longyue Wang, Xuebo Liu, Derek F. Wong, Dacheng Tao, Zhaopeng Tu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ACL/IJCNLP (1)1
2021 Improving Neural Machine Translation by Bidirectional Training
abstract
We present a simple and effective pretraining strategy -bidirectional training (BiT) for neural machine translation.Specifically, we bidirectionally update the model parameters at the early stage and then tune the model normally.To achieve bidirectional updating, we simply reconstruct the training samples from "src→tgt" to "src+tgt→tgt+src" without any complicated model modifications.Notably, our approach does not increase any parameters or training steps, requiring the parallel data merely.Experimental results show that BiT pushes the SOTA neural machine translation performance across 15 translation tasks on 8 language pairs (data sizes range from 160K to 38M) significantly higher.Encouragingly, our proposed model can complement existing data manipulation strategies, i.e. back translation, data distillation and data diversification.Extensive analyses show that our approach functions as a novel bilingual code-switcher, obtaining better bilingual alignment.
Liang Ding 0006, Dacheng Tao
EMNLP (1)1
2021 Towards Efficiently Diversifying Dialogue Generation Via Embedding Augmentation
abstract
Dialogue generation models face the challenge of producing generic and repetitive responses. Unlike previous augmentation methods that mostly focus on token manipulation and ignore the essential variety within a single sample using hard labels, we propose to promote the generation diversity of the neural dialogue models via soft embedding augmentation along with soft labels in this paper. Particularly, we select some key input tokens and fuse their embeddings together with embeddings from their semantic-neighbor tokens. The new embeddings serve as the input of the model to replace the original one. Besides, soft labels are used in loss calculation, resulting in multi-target supervision for a given input. Our experimental results on two datasets illustrate that our proposed method is capable of generating more diverse responses than raw models while remains a similar n-gram accuracy that ensures the quality of generated responses.
Yu Cao 0014, Liang Ding 0006, Zhiliang Tian
ICASSP2
2021 Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
Xuebo Liu 0002, Longyue Wang, Derek F. Wong, Liang Ding 0006, Lidia S. Chao, Zhaopeng Tu
ICLR4
2021 Understanding and Improving Lexical Choice in Non-Autoregressive Translation
Liang Ding 0006, Longyue Wang, Xuebo Liu 0002, Derek F. Wong, Dacheng Tao, Zhaopeng Tu
ICLR1
2020 Self-Attention with Cross-Lingual Position Representation
abstract
Position encoding (PE), an essential part of self-attention networks (SANs), is used to preserve the word order information for natural language processing tasks, generating fixed position indices for input sequences.However, in cross-lingual scenarios, e.g., machine translation, the PEs of source and target sentences are modeled independently.Due to word order divergences in different languages, modeling the cross-lingual positional relationships might help SANs tackle this problem.In this paper, we augment SANs with crosslingual position representations to model the bilingually aware latent structure for the input sentence.Specifically, we utilize bracketing transduction grammar (BTG)-based reordering information to encourage SANs to learn bilingual diagonal alignments.Experimental results on WMT'14 English⇒German, WAT'17 Japanese⇒English, and WMT'17 Chinese⇔English translation tasks demonstrate that our approach significantly and consistently improves translation quality over strong baselines.Extensive analyses confirm that the performance gains come from the cross-lingual information.
Liang Ding 0006, Longyue Wang, Dacheng Tao
ACL1
2020 Context-Aware Cross-Attention for Non-Autoregressive Translation
abstract
Non-autoregressive translation (NAT) significantly accelerates the inference process by predicting the entire target sequence.However, due to the lack of target dependency modelling in the decoder, the conditional generation process heavily depends on the cross-attention.In this paper, we reveal a localness perception problem in NAT cross-attention, for which it is difficult to adequately capture source context.To alleviate this problem, we propose to enhance signals of neighbour source tokens into conventional cross-attention.Experimental results on several representative datasets show that our approach can consistently improve translation quality over strong NAT baselines.Extensive analyses demonstrate that the enhanced cross-attention achieves better exploitation of source contexts by leveraging both local and global information.
Liang Ding 0006, Longyue Wang, Dacheng Tao, Zhaopeng Tu
COLING1