Jingang Wang

dblp:59/7807 · DBLP profile ↗
← Back
89ranked-venue papers
11as first author
77since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 3 first-author · 61 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 14 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Mathematical Reasoning to Code: Generalization of Process Reward Models in Test-Time Scaling
abstract
Recent advancements in improving the reasoning capabilities of Large Language Models have underscored the efficacy of Process Reward Models (PRMs) in addressing intermediate errors through structured feedback mechanisms. This study analyzes PRMs from multiple perspectives, including training methodologies, scalability, and generalization capabilities. We investigate the interplay between pre-training and reward model training FLOPs to assess their influence on PRM efficiency and accuracy in complex reasoning tasks. Our analysis reveals a pattern of diminishing returns in performance with increasing PRM scale, highlighting the importance of balancing model size and computational cost. Furthermore, the diversity of training datasets significantly impacts PRM performance, emphasizing the importance of diverse data to enhance both accuracy and efficiency. We further examine test-time scaling strategies, identifying Monte Carlo Tree Search as the most effective method when computational resources are abundant, while Best-of-N Sampling serves as a practical alternative under resource-limited conditions. Notably, our findings indicate that PRMs trained on mathematical datasets exhibit performance comparable to those tailored for code generation, suggesting robust cross-domain generalization. Employing a gradient-based metric, we observe that PRMs exhibit a preference for selecting responses with similar underlying patterns, further informing their optimization.
Zhengyu Chen 0001, Teng Xiao, Ruochen Zhou, Xuesheng Yang, Zhifang Sui, Jingang Wang
AAAI8
2026 Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
abstract
The low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of problem difficulty and fail to capture the alignment between model competence and problem difficulty in RL training, leading to suboptimal results. To address these challenges, we introduce Competence-Difficulty Alignment Sampling (CDAS). This approach allows for accurate and stable estimation of problem difficulties by aggregating historical performance discrepancies across problems. Subsequently, model competence is quantified to adaptively select problems whose difficulties align with the model's current competence using a fixed-point system. Extensive experiments in mathematical RL training show that CDAS consistently outperforms strong baselines, achieving the highest average accuracy of 45.89%. Furthermore, CDAS reduces the training step time overhead by 57.06% compared to the widely-used Dynamic Sampling strategy, verifying the efficiency of CDAS. Additional experiments on different tasks, model architectures, and model sizes demonstrate the generalization capability of CDAS.
Deyang Kong, Xiangyu Xi, Wei Wang 0225, Jingang Wang, Shikun Zhang, Wei Ye 0004
AAAI5
2026 Scaling and Transferability of Annealing Strategies in Large Language Model Training
abstract
Learning rate scheduling is crucial for training large language models, yet understanding the optimal annealing strategies across different model configurations remains challenging. In this work, we investigate the transferability of annealing dynamics in large language model training and refine a generalized predictive framework for optimizing annealing strategies under the Warmup-Steady-Decay (WSD) scheduler. Our improved framework incorporates training steps, maximum learning rate, and annealing behavior, enabling more efficient optimization of learning rate schedules. Our work provides a practical guidance for selecting optimal annealing strategies without exhaustive hyperparameter searches, demonstrating that smaller models can serve as reliable proxies for optimizing the training dynamics of larger models. We validate our findings on extensive experiments using both Dense and Mixture-of-Experts (MoE) models, demonstrating that optimal annealing ratios follow consistent patterns and can be transferred across different training configurations.
Zhengyu Chen 0001, Teng Xiao, Zheqi Lv, Jinluan Yang, Jingang Wang
AAAI7
2026 LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidance
abstract
Yuchun Fan, Bei Li, Peiguang Li, Yilin Wang, Yongyu Mu, Jian Yang, Xin Chen, Rongxiang Weng, Jingang Wang, Xunliang Cai, JingBo Zhu, Tong Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yuchun Fan, Peiguang Li, Yongyu Mu, Rongxiang Weng, Jingang Wang, Tong Xiao 0001
ACL (1)9
2026 MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
abstract
Junhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin, Xinyu Liu, Kechen Jiao, Xin Chen, Jingang Wang, Xunliang Cai, Tong Xiao, JingBo Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junhao Ruan, Abudukeyumu Abudula, Yongjing Yin, Kechen Jiao, Jingang Wang, Tong Xiao 0001
ACL (1)8
2026 BaseCal: Unsupervised Confidence Calibration via Base Model Signals
abstract
Hexiang Tan, Wanli Yang, Junwei Zhang, Xin Chen, Rui Tang, Du Su, Jingang Wang, Yuanzhuo Wang, Fei Sun, Xueqi Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hexiang Tan, Du Su, Jingang Wang, Yuanzhuo Wang, Fei Sun 0001, Xueqi Cheng 0001
ACL (1)7
2026 The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
abstract
Zihao Wei, Liang Pang, Jiahao Liu, Wenjie Shi, Jingcheng Deng, Shicheng Xu, Zenghao Duan, Jingang Wang, Fei Sun, Huawei Shen, Xueqi Cheng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zihao Wei, Liang Pang 0001, Wenjie Shi, Jingcheng Deng, Zenghao Duan, Jingang Wang, Fei Sun 0001, Huawei Shen, Xueqi Cheng 0001
ACL (1)8
2026 Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
abstract
Zhihao Xu, Rumei Li, Jiahuan Li, Rongxiang Weng, Jingang Wang, Xunliang Cai, Xiting Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahuan Li, Rongxiang Weng, Jingang Wang, Xiting Wang
ACL (1)5
2026 LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
abstract
Xuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng, Hongfei Yan, Jingang Wang, Xunliang Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuemiao Zhang, Can Ren, Chengying Tu, Rongxiang Weng, Hongfei Yan, Jingang Wang
ACL (1)6
2026 From haze degradation to haze evolution: an anisotropic-diffusion-inspired generative framework for image dehazing
Jingyuan Bai, Jingang Wang
Pattern Anal. Appl.2
2026 Cross-Domain detection of AI-Generated text: Integrating linguistic richness and lexical pair dispersion via deep learning
abstract
Cross-domain detection of AI-generated text is a crucial task for cybersecurity. In practical scenarios, after being trained on one or multiple known text generation sources (source domain), a detection model must be capable of effectively identifying text generated by unknown and unseen sources (target domain). Current approaches suffer from limited cross-domain generalization due to insufficient structural adaptation to domain discrepancies. To address this critical limitation, we propose RiDis ,a classification model that synergizes Linguistic Ri chness and Lexical Pair Dis persion for cross-domain AI-generated text detection. Through comprehensive statistical analysis, we establish Linguistic Richness and Lexical Pair Dispersion as discriminative indicators for distinguishing human-authored and machine-generated texts. Our architecture features two innovative components, a Semantic Coherence Extraction Module employing long-range receptive fields to capture linguistic richness through global semantic trend analysis, and a Contextual Dependency Extraction Module utilizing localized receptive fields to quantify lexical pair dispersion via fine-grained word association patterns. The framework further incorporates domain adaptation learning to enhance cross-domain detection robustness. Extensive evaluations demonstrate that our method achieves superior detection accuracy compared to state-of-the-art baselines across multiple domains, with experimental results showing significant performance improvements on cross-domain test scenarios.
Jingang Wang, Tong Xiao 0017, Cheng Zhang 0042, Peng Liu 0046
Pattern Recognit. Lett.1
2025 SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
abstract
As Large Language Models (LLMs) continue to advance in capability and influence, ensuring their security and preventing harmful outputs has become crucial. A promising approach to address these concerns involves training models to automatically generate adversarial prompts for red teaming. However, the evolving subtlety of vulnerabilities in LLMs challenges the effectiveness of current adversarial methods, which struggle to generate diverse, complex prompts and dynamically explore the weaknesses of these models. To tackle these challenges, we introduce the Self-Evolving Adversarial Safety (SEAS) optimization framework, which includes both a SEAS dataset and a SEAS pipeline. The SEAS dataset comprises complex adversarial prompts, while the SEAS pipeline operates through three stages: Initialization, Attack, and Adversarial Optimization. This framework generates a diverse range of adversarial prompts and dynamically explores the model's vulnerabilities to enhance its security. Our contributions include a novel adversarial framework, a comprehensive safety dataset, and empirical evidence demonstrating the effectiveness of SEAS.
Muxi Diao, Shiyang Liu, Guogang Liao, Jingang Wang, Weiran Xu
AAAI5
2025 Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies
abstract
Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, Jingang Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhengyu Chen 0001, Teng Xiao, Shiqi Chen 0002, Junxian He, Jingang Wang
ACL (1)8
2025 Sibyl: Empowering Empathetic Dialogue Generation in Large Language Models via Sensible and Visionary Commonsense Inference
abstract
Recently, there has been a heightened interest in building chatbots based on Large Language Models (LLMs) to emulate human-like qualities in multi-turn conversations. Despite having access to commonsense knowledge to better understand the psychological aspects and causality of dialogue context, even these powerful LLMs struggle to achieve the goals of empathy and emotional support. Current commonsense knowledge derived from dialogue contexts is inherently limited and often fails to adequately anticipate the future course of a dialogue. This lack of foresight can mislead LLMs and hinder their ability to provide effective support. In response to this challenge, we present an innovative framework named Sensible and Visionary Commonsense Knowledge (Sibyl). Designed to concentrate on the immediately succeeding dialogue, this paradigm equips LLMs with the capability to uncover the implicit requirements of the conversation, aiming to elicit more empathetic responses. Experimental results demonstrate that incorporating our paradigm for acquiring commonsense knowledge into LLMs comprehensively enhances the quality of their responses.
Lanrui Wang, Chenxu Yang, Zheng Lin 0001, Hongyin Tang, Yanan Cao 0001, Jingang Wang, Weiping Wang 0005
COLING8
2025 Jailbreak LLMs through Internal Stance Manipulation
abstract
To confront the ever-evolving safety risks of LLMs, automated jailbreak attacks have proven effective for proactively identifying security vulnerabilities at scale.Existing approaches, including GCG and AutoDAN, generate adversarial prompts for malicious requests that induce LLMs to respond following a fixed affirmative template.However, we observed that the reliance on the fixed output template is ineffective for certain malicious requests, leading to suboptimal jailbreak performance.In this work, we aim to develop a method that generalizes across all malicious requests.Our approach is inspired by the discovery of LLMs' intrinsic safety mechanisms: they tend to exhibit a similar refusal stance across diverse adversarial prompts, resulting in consistent rejections.We propose Stance Manipulation (SM), a novel automated jailbreak approach that generates adversarial prompts to suppress the refusal stance and induce affirmative responses.Our experiments across four mainstream open-source LLMs demonstrate the superiority of SM's performance.Under commonly used setting, SM achieves success rates over 77.1% across all models on Advbench.Specifically, for Llama-2-7b-chat, SM outperforms the best baseline by 25.4%.In further experiments with extended iterations, SM achieves over 92.2% attack success rate across all models.Our code is publicly available at https://github.com/Zed630/Stance- Manipulation
Shuangjie Fu, Du Su, Beining Huang, Fei Sun 0001, Jingang Wang, Wei Chen 0013, Huawei Shen, Xueqi Cheng 0001
EMNLP5
2025 TCPO: Thought-Centric Preference Optimization for Effective Embodied Decision-making
abstract
Kechen Jiao, Zhirui Fang, Jiahao Liu, Bei Li, Qifan Wang, Xinyu Liu, Junhao Ruan, Zhongjian Qiao, Yifan Zhu, Yaxin Xu, Jingang Wang, Xiu Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Kechen Jiao, Zhirui Fang, Junhao Ruan, Zhongjian Qiao, Yaxin Xu, Jingang Wang, Xiu Li 0001
EMNLP11
2025 IIET: Efficient Numerical Transformer via Implicit Iterative Euler Method
abstract
Xinyu Liu, Bei Li, Jiahao Liu, Junhao Ruan, Kechen Jiao, Hongyin Tang, Jingang Wang, Tong Xiao, JingBo Zhu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Junhao Ruan, Kechen Jiao, Hongyin Tang, Jingang Wang, Tong Xiao 0001
EMNLP7
2025 Too Consistent to Detect: A Study of Self-Consistent Errors in LLMs
abstract
Hexiang Tan, Fei Sun, Sha Liu, Du Su, Qi Cao, Xin Chen, Jingang Wang, Xunliang Cai, Yuanzhuo Wang, Huawei Shen, Xueqi Cheng. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Hexiang Tan, Fei Sun 0001, Du Su, Qi Cao 0005, Jingang Wang, Yuanzhuo Wang, Huawei Shen, Xueqi Cheng 0001
EMNLP7
2025 FIRE: Flexible Integration of Data Quality Ratings for Effective Pretraining
abstract
Selecting high-quality data can improve the pretraining efficiency of large language models (LLMs).Existing methods generally rely on heuristic techniques or single quality signals, limiting their ability to evaluate data quality comprehensively.In this work, we propose FIRE, a flexible and scalable framework for integrating multiple data quality raters, which allows for a comprehensive assessment of data quality across various dimensions.FIRE aligns multiple quality signals into a unified space, and integrates diverse data quality raters to provide a comprehensive quality signal for each data point.Further, we introduce a progressive data selection scheme based on FIRE that iteratively refines the selection of high-quality data points.Extensive experiments show that FIRE outperforms other data selection methods and significantly boosts pretrained model performance across a wide range of downstream tasks, while requiring less than 37.5% tokens needed by the Random baseline to reach the target performance.
Liangyu Xu, Xuemiao Zhang, Feiyu Duan, Rongxiang Weng, Jingang Wang
EMNLP6
2025 AgentRefine: Enhancing Agent Generalization through Refinement Tuning
abstract
Large Language Model (LLM) based agents have proved their ability to perform complex tasks like humans. However, there is still a large gap between open-sourced LLMs and commercial models like the GPT series. In this paper, we focus on improving the agent generalization capabilities of LLMs via instruction tuning. We first observe that the existing agent training corpus exhibits satisfactory results on held-in evaluation sets but fails to generalize to held-out sets. These agent-tuning works face severe formatting errors and are frequently stuck in the same mistake for a long while. We analyze that the poor generalization ability comes from overfitting to several manual agent environments and a lack of adaptation to new situations. They struggle with the wrong action steps and can not learn from the experience but just memorize existing observation-action relations. Inspired by the insight, we propose a novel AgentRefine framework for agent-tuning. The core idea is to enable the model to learn to correct its mistakes via observation in the trajectory. Specifically, we propose an agent synthesis framework to encompass a diverse array of environments and tasks and prompt a strong LLM to refine its error action according to the environment feedback. AgentRefine significantly outperforms state-of-the-art agent-tuning work in terms of generalization ability on diverse agent tasks. It also has better robustness facing perturbation and can generate diversified thought in inference. Our findings establish the correlation between agent generalization and self-refinement and provide a new paradigm for future research.
Dayuan Fu, Keqing He 0001, Yejie Wang, Wentao Hong, Zhuoma Gongque, Weihao Zeng 0003, Jingang Wang, Weiran Xu
ICLR8
2025 Earlier Tokens Contribute More: Learning Direct Preference Optimization From Temporal Decay Perspective
abstract
Direct Preference Optimization (DPO) has gained attention as an efficient alternative to reinforcement learning from human feedback (RLHF) for aligning large language models (LLMs) with human preferences. Despite its advantages, DPO suffers from a length bias, generating responses longer than those from the reference model. Existing solutions like SimPO and SamPO address this issue but uniformly treat the contribution of rewards across sequences, overlooking temporal dynamics. To this end, we propose an enhanced preference optimization method that incorporates a temporal decay factor controlled by a gamma parameter. This dynamic weighting mechanism adjusts the influence of each reward based on its position in the sequence, prioritizing earlier tokens that are more critical for alignment. By adaptively focusing on more relevant feedback, our approach mitigates overfitting to less pertinent data and remains responsive to evolving human preferences. Experimental results on several benchmarks show that our approach consistently outperforms vanilla DPO by 5.9-8.8 points on AlpacaEval 2 and 3.3-9.7 points on Arena-Hard across different model architectures and sizes. Furthermore, additional experiments on mathematical and reasoning benchmarks (MMLU, GSM8K, and MATH) confirm that our method enhances performance without compromising general capabilities. Our codebase would be available at \url{https://github.com/LotuSrc/D2PO}.
Ruichen Shao, Gangao Liu, ZhouXiang, Jingang Wang
ICLR6
2025 Dynamic Fisher-weighted Model Merging via Bayesian Optimization
abstract
Sanwoo Lee, Jiahao Liu, Qifan Wang, Jingang Wang, Xunliang Cai, Yunfang Wu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sanwoo Lee, Qifan Wang 0001, Jingang Wang, Yunfang Wu
NAACL (Long Papers)4
2025 NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables
abstract
Processing structured tabular data, particularly large and lengthy tables, constitutes a fundamental yet challenging task for large language models (LLMs). However, existing long-context benchmarks like Needle-in-a-Haystack primarily focus on unstructured text, neglecting the challenge of diverse structured tables. Meanwhile, previous tabular benchmarks mainly consider downstream tasks that require high-level reasoning abilities, and overlook models' underlying fine-grained perception of individual table cells, which is crucial for practical and robust LLM-based table applications. To address this gap, we introduce \textsc{NeedleInATable} (NIAT), a new long-context tabular benchmark that treats each table cell as a ``needle'' and requires models to extract the target cell based on cell locations or lookup questions. Our comprehensive evaluation of various LLMs and multimodal LLMs reveals a substantial performance gap between popular downstream tabular tasks and the simpler NIAT task, suggesting that they may rely on dataset-specific correlations or shortcuts to obtain better benchmark results but lack truly robust long-context understanding towards structured tables. Furthermore, we demonstrate that using synthesized NIAT training data can effectively improve performance on both NIAT task and downstream tabular tasks, which validates the importance of NIAT capability for LLMs' genuine table understanding ability. Our data, code and models will be released to facilitate future research.
Lanrui Wang, Mingyu Zheng, Hongyin Tang, Zheng Lin 0001, Yanan Cao 0001, Jingang Wang, Weiping Wang 0005
NeurIPS6
2025 ProFus: Progressive Radar-Vision Heterogeneous Modality Fusion for Maritime Target Detection
abstract
Maritime monitoring is crucial in both civilian and military applications, with shore-based radar and visual systems widely used due to their cost-effectiveness. However, single-sensor methods have notable limitations: radar systems, while offering wide detection coverage, suffer from high false alarm rates and lack detailed target information, whereas visual systems provide rich details but perform poorly in adverse weather conditions such as rain and fog. To address these issues, this paper proposes a progressive radar-vision fusion method for surface target detection. Due to the significant differences in data characteristics between radar and visual sensors, direct fusion is nearly infeasible. Instead, the proposed method adopts a stepwise fusion strategy, consisting of coordinate calibration, shallow feature fusion, and deep feature integration. Experimental results show that this approach achieves an mAP50of 86.7% and an mAP75of 54.5%, outperforming YOLOv10 by 1.0% and 1.5%, respectively. Moreover, the proposed method significantly surpasses existing state-of-the-art radar-vision fusion approaches, demonstrating its superior effectiveness in complex environments.
Jingang Wang, Shikai Wu, Peng Liu 0046
IEEE Geosci. Remote. Sens. Lett.1
2025 A generative image steganography method based on joint encoding of multi-object semantic information
Peng Liu 0046, Songbin Li, Jingang Wang
Pattern Anal. Appl.4
2025 Radar target tracking based on motion characteristic and distribution pattern matching
Jingang Wang, Songbin Li
Signal Process.1
2025 General Steganalysis of Generative Linguistic Steganography Based on Dynamic Segment-Level Lexical Association Extraction
abstract
In scenarios where steganographic texts from various steganographic domains generated by different generative steganography algorithms are mixed, most existing linguistic steganalysis methods lack corresponding network structures designed to account for the differences in steganographic texts from different domains, leading to the potential for further improvement in their general detection performance. To address the above issue, we propose a general generative linguistic steganalysis method based on the basic idea of dynamically extracting lexical association features of different steganographic domains at the segment level. We utilize dynamic-static text feature matrix to construct a word importance semantic encoding module to mine steganography-sensitive word features of different steganographic domains. Based on the obtained features, we propose a word correlation multi-scale perception module to focus on the segment-level lexical association changes caused by secret information embedding in different domains. Experimental results show that this method can improve the detection accuracy of existing mainstream linguistic steganalysis methods in various mixed steganography scenarios.
Songbin Li, Jingang Wang
IEEE Signal Process. Lett.3
2025 Heterogeneous Domain Remapping for Universal Detection of Generative Linguistic Steganography
abstract
Current researchers have proposed various steganalysis methods for detecting secret information within social media texts, which can achieve relatively optimal detection performance in specific steganographic domains. However, considering the practical application of social media, we can only obtain the text to be tested without prior knowledge of the steganographic domain it belongs to. Consequently, we are unable to prepare a supervised training dataset in advance. This places higher demands on steganalysis algorithms, necessitating their ability to generalize and detect any unknown steganography domain. To this end, we propose a universal detection method for generative linguistic steganography based on heterogeneous domain remapping. The core idea is to employ a neural structure composed of pre-trained embedding layers and capsule networks to extract steganography-sensitive correlation features. Subsequently, the concept of contrastive learning is utilized to remap the sensitive features from heterogeneous steganography domains into a unified domain. This process effectively extracts domain-invariant features, thereby enabling the detection of unknown steganographic domains. Experimental results demonstrate that the proposed method outperforms existing approaches by an average of over 2% across various steganography domains.
Tong Xiao 0017, Jingang Wang, Songbin Li
IEEE Signal Process. Lett.2
2025 Biresidual Compression Network With Conditional Diffusion Model for Hyperspectral Image Compression
abstract
Hyperspectral image (HSI) compression presents the challenge of preserving both spectral and spatial fidelity while achieving high compression rates. Current compression methods frequently depend on band-by-band compression or simplistic joint modeling, which complicates the balance between spectral consistency and perceptual quality. To address this issue, a compression driven generation framework (BRC-CDM) is proposed, which decouples the extraction of compressed representations from the reconstruction of high-quality images. We introduce reference band information through channel-level concatenation to guide the spectral residual compression network in collaboratively extracting residual information in both spatial and spectral domains. During the prediction phase, a spectral gaussian grid compensation structure is further integrated to enhance the accuracy of predictions for the target bands. Ultimately, by compressing the residual information based on the differences between the predicted bands and the true bands, an efficient representation of the residual information in the compressed domain is achieved. The predicted image is then fused with the compressed residual to obtain a more expressive and potentially compressed representation. During the decoding phase, the diffusion process of the conditional diffusion model (CDM) utilizes compressed representations as potential conditions to guide the model in progressively reconstructing images over multiple time steps. Experimental findings demonstrate that the bi-residual compression network achieves superior peak signal-to-noise ratio (PSNR) across nine datasets, with an improvement of approximately 1 dB over SOTA methods. Moreover, BRC-CDM attains a PSNR that surpasses that of the majority of current methodologies, while also delivering enhanced spectral fidelity and perceptual quality. To foster reproducibility and further development, we release the full implementation of BRC-CDM at https://github.com/Nicle-L/BRC-CDM.
Lili Zhang 0005, Jingang Wang, Lele Qu
IEEE Trans. Geosci. Remote. Sens.3
2025 Lightweight single-head attention mechanism-driven diffusion model for hyperspectral image classification
Qizhi Fang, Jingang Wang
J. Supercomput.3
2024 What Makes Quantization for Large Language Model Hard? An Empirical Study from the Lens of Perturbation
abstract
Quantization has emerged as a promising technique for improving the memory and computational efficiency of large language models (LLMs). Though the trade-off between performance and efficiency is well-known, there is still much to be learned about the relationship between quantization and LLM performance. To shed light on this relationship, we propose a new perspective on quantization, viewing it as perturbations added to the weights and activations of LLMs. We call this approach ``the lens of perturbation". Using this lens, we conduct experiments with various artificial perturbations to explore their impact on LLM performance. Our findings reveal several connections between the properties of perturbations and LLM performance, providing insights into the failure cases of uniform quantization and suggesting potential solutions to improve the robustness of LLM quantization. To demonstrate the significance of our findings, we implement a simple non-uniform quantization approach based on our insights. Our experiments show that this approach achieves minimal performance degradation on both 4-bit weight quantization and 8-bit quantization for weights and activations. These results validate the correctness of our approach and highlight its potential to improve the efficiency of LLMs without sacrificing performance.
Zhuocheng Gong, Jingang Wang, Dongyan Zhao 0001, Rui Yan 0001
AAAI3
2024 CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video Grounding
abstract
This paper studies the spatio-temporal video grounding task, which aims to localize a spatio-temporal tube in an untrimmed video based on the given text description of an event. Existing one-stage approaches suffer from insufficient space-time interaction in two aspects: i) less precise prediction of event temporal boundaries, and ii) inconsistency in object prediction for the same event across adjacent frames. To address these issues, we propose a framework of Comprehensive Space-Time entAnglement (CoSTA) to densely entangle space-time multi-modal features for spatio-temporal localization. Specifically, we propose a space-time collaborative encoder to extract comprehensive video features and leverage Transformer to perform spatio-temporal multi-modal understanding. Our entangled decoder couples temporal boundary prediction and spatial localization via an entangled query, boasting an enhanced ability to capture object-event relationships. We conduct extensive experiments on the challenging benchmarks of HC-STVG and VidSTG, where CoSTA outperforms existing state-of-the-art methods, demonstrating its effectiveness for this task.
Yaoyuan Liang, Yansong Tang, Zhao Yang 0002, Ziran Li, Jingang Wang, Wenbo Ding 0001, Shao-Lun Huang
AAAI6
2024 MCL-NER: Cross-Lingual Named Entity Recognition via Multi-View Contrastive Learning
abstract
Cross-lingual named entity recognition (CrossNER) faces challenges stemming from uneven performance due to the scarcity of multilingual corpora, especially for non-English data. While prior efforts mainly focus on data-driven transfer methods, a significant aspect that has not been fully explored is aligning both semantic and token-level representations across diverse languages. In this paper, we propose Multi-view Contrastive Learning for Cross-lingual Named Entity Recognition (MCL-NER). Specifically, we reframe the CrossNER task into a problem of recognizing relationships between pairs of tokens. This approach taps into the inherent contextual nuances of token-to-token connections within entities, allowing us to align representations across different languages. A multi-view contrastive learning framework is introduced to encompass semantic contrasts between source, codeswitched, and target sentences, as well as contrasts among token-to-token relations. By enforcing agreement within both semantic and relational spaces, we minimize the gap between source sentences and their counterparts of both codeswitched and target sentences. This alignment extends to the relationships between diverse tokens, enhancing the projection of entities across languages. We further augment CrossNER by combining self-training with labeled source data and unlabeled target data. Our experiments on the XTREME benchmark, spanning 40 languages, demonstrate the superiority of MCL-NER over prior data-driven and model-based approaches. It achieves a substantial increase of nearly +2.0 F1 scores across a broad spectrum and establishes itself as the new state-of-the-art performer.
Ying Mo, Jian Yang 0030, Qifan Wang 0001, Jingang Wang, Zhoujun Li 0001
AAAI6
2024 DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning
abstract
Yejie Wang, Keqing He, Guanting Dong, Pei Wang, Weihao Zeng, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang, Xunliang Cai. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yejie Wang, Keqing He 0001, Guanting Dong 0001, Weihao Zeng 0003, Muxi Diao, Weiran Xu, Jingang Wang, Mengdi Zhang 0002
ACL (1)8
2024 Beyond the Known: Investigating LLMs Performance on Out-of-Domain Intent Detection
abstract
Out-of-domain (OOD) intent detection aims to examine whether the user’s query falls outside the predefined domain of the system, which is crucial for the proper functioning of task-oriented dialogue (TOD) systems. Previous methods address it by fine-tuning discriminative models. Recently, some studies have been exploring the application of large language models (LLMs) represented by ChatGPT to various downstream tasks, but it is still unclear for their ability on OOD detection task.This paper conducts a comprehensive evaluation of LLMs under various experimental settings, and then outline the strengths and weaknesses of LLMs. We find that LLMs exhibit strong zero-shot and few-shot capabilities, but is still at a disadvantage compared to models fine-tuned with full resource. More deeply, through a series of additional analysis experiments, we discuss and summarize the challenges faced by LLMs and provide guidance for future work including injecting domain knowledge, strengthening knowledge transfer from IND(In-domain) to OOD, and understanding long instructions.
Keqing He 0001, Yejie Wang, Xiaoshuai Song, Yutao Mou, Jingang Wang, Yunsen Xian, Weiran Xu
LREC/COLING6
2024 Task-agnostic Distillation of Encoder-Decoder Language Models
abstract
Finetuning pretrained language models (LMs) have enabled appealing performance on a diverse array of tasks. The intriguing task-agnostic property has driven a shifted focus from task-specific to task-agnostic distillation of LMs. While task-agnostic, compute-efficient, performance-preserved LMs can be yielded by task-agnostic distillation, previous studies mainly sit in distillation of either encoder-only LMs (e.g., BERT) or decoder-only ones (e.g., GPT) yet largely neglect that distillation of encoder-decoder LMs (e.g., T5) can posit very distinguished behaviors. Frustratingly, we discover that existing task-agnostic distillation methods can fail to handle the distillation of encoder-decoder LMs. To the demand, we explore a few paths and uncover a path named as MiniEnD that successfully tackles the distillation of encoder-decoder LMs in a task-agnostic fashion. We examine MiniEnD on language understanding and abstractive summarization. The results showcase that MiniEnD is generally effective and is competitive compared to other alternatives. We further scale MiniEnD up to distillation of 3B encoder-decoder language models with interpolated distillation. The results imply the opportunities and challenges in distilling large language models (e.g., LLaMA).
Chen Zhang 0020, Yang Yang 0129, Qiuchi Li, Jingang Wang, Dawei Song 0001
LREC/COLING4
2024 Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-Context Models
abstract
Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model's effective memorization length.However, through thorough investigations, we find limitations for currently existing evaluations on model's memorization capability.We provide an extensive survey for limitations in this work and propose a new method called forgetting curve to measure the memorization capability of long-context models.We show that forgetting curve has the advantage of being robust to the tested corpus and the experimental settings, of not relying on prompts and can be applied to any model size.We apply our forgetting curve to a large variety of models involving both transformer and RNN/SSM based architectures.Our measurement provides empirical evidence for the effectiveness of transformer extension techniques while raises questions for the effective length of RNN/SSM based models.We also examine the difference between our measurement and existing benchmarks as well as popular metrics for various models.Our code and results can be found at https://github.com/1azybug/ForgettingCurve.
Runsong Zhao, Pengcheng Huang 0004, Chunyang Xiao, Jingang Wang, Tong Xiao 0001
EMNLP6
2024 How Do Your Code LLMs perform? Empowering Code Instruction Tuning with Really Good Data
abstract
Yejie Wang, Keqing He, Dayuan Fu, Zhuoma GongQue, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong, Muxi Diao, Jingang Wang, Mengdi Zhang, Xunliang Cai, Weiran Xu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yejie Wang, Keqing He 0001, Dayuan Fu, Zhuoma Gongque, Heyang Xu, Yanxu Chen, Zhexu Wang, Yujia Fu, Guanting Dong 0001, Muxi Diao, Jingang Wang, Mengdi Zhang 0002, Weiran Xu
EMNLP11
2024 Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
abstract
The scaling of large language models (LLMs) is a critical research area for the efficiency and effectiveness of model training and deployment.Our work investigates the transferability and discrepancies of scaling laws between Dense Models and Mixture of Experts (MoE) models.Through a combination of theoretical analysis and extensive experiments, including consistent loss scaling, optimal batch size and learning rate scaling, and resource allocation strategies scaling, our findings reveal that the power-law scaling framework also applies to MoE Models, indicating that the fundamental principles governing the scaling behavior of these models are preserved, even though the architecture differs.Additionally, MoE Models demonstrate superior generalization, resulting in lower testing losses with the same training compute budget compared to Dense Models.These findings indicate the scaling consistency and transfer generalization capabilities of MoE Models, providing new insights for optimizing MoE Model training and deployment strategies.
Zhengyu Chen 0001, Keqing He 0001, Min Zhang 0068, Jingang Wang
EMNLP6
2024 Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning
abstract
Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution.'' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30.95 and 44.27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3.8B DeepNet by an average of 2.9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5.7 accuracy points on the LM Harness Evaluation.
Rui Wang 0028, Qingyan Guo, Junliang Guo, Xu Tan 0003, Tong Xiao 0001, Jingang Wang
NeurIPS10
2024 Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression Generation
abstract
The Multi-modal Large Language Model (MLLM) based Referring Expression Generation (REG) task has gained increasing popularity, which aims to generate an unambiguous text description that applies to exactly one object or region in the image by leveraging foundation models. We empirically found that there exists a potential trade-off between the detailedness and the correctness of the descriptions for the referring objects. On the one hand, generating sentences with more details is usually required in order to provide more precise object descriptions. On the other hand, complicated sentences could easily increase the probability of hallucinations. To address this issue, we propose a training-free framework, named ``unleash-then-eliminate'', which first elicits the latent information in the intermediate layers, and then adopts a cycle-consistency-based decoding method to alleviate the production of hallucinations. Furthermore, to reduce the computational load of cycle-consistency-based decoding, we devise a Probing-based Importance Estimation method to statistically estimate the importance weights of intermediate layers within a subset. These importance weights are then incorporated into the decoding process over the entire dataset, intervening in the next token prediction from intermediate layers. Extensive experiments conducted on the RefCOCOg and PHD benchmarks show that our proposed framework could outperform existing methods on both semantic and hallucination-related metrics. Code will be made available in https://github.com/Glupayy/unleash-eliminate.
Yaoyuan Liang, Zhuojun Cai, Guanbo Huang, Ziran Li, Jingang Wang, Shao-Lun Huang
NeurIPS9
2024 FIRP: Faster LLM Inference via Future Intermediate Representation Prediction
Pengfei Wu 0006, Zhuocheng Gong, Qifan Wang 0001, Jinpeng Li 0003, Jingang Wang, Dongyan Zhao 0001
NLPCC (3)6
2024 Two-Stage Collaborative Sea Clutter Suppression Based on Subregion Gazing and Reconstruction
abstract
Radar target detection is an important technical means for effective monitoring of sea surface targets. Under strong sea clutter, the echo characteristics of some targets with small radar cross-sections (RCS) may be submerged, affecting the performance of the detector, such as CFAR (Constant False Alarm Rate). The importance of sea clutter suppression as one of the preprocessing steps for detection is self-evident. In this letter, we propose a two-stage collaborative sea clutter suppression algorithm based on subregion gazing and reconstruction, which combines data-driven neural models and vector factorization theory to first magnify the subregions where targets may exist in the echo, and then use vector factorization methods to suppress clutter and noise within the echo reception window. To verify the effectiveness of the proposed algorithm, we conducted long-term observations of several channel buoys in the southern waters of China, forming an effective dataset. Experimental results on this dataset demonstrate that, by employing the sea clutter suppression method proposed in this paper as a data preprocessing technique, the F1-score of the detection algorithm can average over 56.81%.
Jingang Wang, Songbin Li
IEEE Geosci. Remote. Sens. Lett.1
2024 SANet: A Compressed Speech Encoder and Steganography Algorithm Independent Steganalysis Deep Neural Network
abstract
Most of the existing steganalysis methods for low-bit-rate compressed speech are specifically designed for a particular speech encoder or category of steganography methods, limiting their generalization capability. These methods require pre-selection of codewords affected by the specific steganographic process as input to the steganalysis models. In order to overcome this limitation and enhance the practicality of steganalysis algorithms, we propose a compressedSpeech encoder and steganographyAlgorithm independent steganalysisNetwork, namedSANet. Irrespective of the specific steganography algorithm used, modifications to the codewords will impact the sequential correlation characteristics of uncompressed domain (time domain) speech. Additionally, the compressed speech streams from different coders are unified in the uncompressed domain format. Therefore, this article introduces an intermediate representation based on the uncompressed domain and develops a neural network that utilizes collaborative correlation features to extract steganography-sensitive characteristics from this representation. Experimental results demonstrate that our proposed method achieves state-of-the-art detection performance for various steganography algorithms under different speech encoders.
Songbin Li, Jingang Wang, Peng Liu 0046
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multi-Task Multi-Attention Transformer for Generative Named Entity Recognition
abstract
Most previous sequential labeling models are task-specific, while recent years have witnessed the rise of generative models due to the advantage of unifying all named entity recognition (NER) tasks into the encoder-decoder framework. Although achieving promising performance, our pilot studies demonstrate that existing generative models are ineffective at detecting entity boundaries and estimating entity types. In this paper, we propose a multi-task Transformer, which incorporates an entity boundary detection task into the named entity recognition task. More concretely, we achieve entity boundary detection by classifying the relations between tokens within the sentence. To improve the accuracy of entity-type mapping during decoding, we adopt an external knowledge base to calculate the prior entity-type distributions and then incorporate the information into the model via the self- and cross-attention mechanisms. We perform experiments on extensive NER benchmarks, including flat, nested, and discontinuous NER datasets involving long entities. It substantially increases nearly$+0.3 \sim +1.5\;{F_1}$scores across a broad spectrum or performs closely to the best generative NER model. Experimental results show that our approach improves the performance of the generative NER model considerably.
Ying Mo, Hongyin Tang, Qifan Wang 0001, Zenglin Xu, Jingang Wang, Xiaojun Quan, Wei Wu 0014, Zhoujun Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Maritime Radar Target Detection Model Self-Evolution Based on Semisupervised Learning
abstract
Radar target detection in sea clutter aims to effectively discern the presence of maritime targets within the current radar echo. With the advancement of deep learning technology, an ever-growing number of researchers are turning to neural networks as the foundation for constructing detection models. These sophisticated neural models demonstrate promising performance in public datasets. On this basis, we propose an innovative self-evolution framework for radar target detection models using semisupervised learning (SesuL). The proposed approach aims to enhance the performance of the detection model under various radar conditions. Notably, this research presents the first-ever attempt within the literature to introduce such an approach for pulse-compression radar. To bolster the performance of the proposed model, novel techniques are introduced for sample selection, sample augmentation, and model optimization. Experimental findings provide compelling evidence supporting the superiority of the proposed method in terms of detection performance and robustness under unknown conditions, surpassing existing techniques. In light of practical deployment considerations, future efforts should be directed toward investigating the fusion of radar and other sensors, such as visible light, to enhance the detection performance.
Jingang Wang, Songbin Li
IEEE Trans. Geosci. Remote. Sens.1
2023 RankCSE: Unsupervised Sentence Representations Learning via Learning to Rank
abstract
Jiduan Liu, Jiahao Liu, Qifan Wang, Jingang Wang, Wei Wu, Yunsen Xian, Dongyan Zhao, Kai Chen, Rui Yan. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Jiduan Liu, Qifan Wang 0001, Jingang Wang, Wei Wu 0014, Yunsen Xian, Dongyan Zhao 0001, Rui Yan 0001
ACL (1)4
2023 Decoupling Pseudo Label Disambiguation and Representation Learning for Generalized Intent Discovery
abstract
Yutao Mou, Xiaoshuai Song, Keqing He, Chen Zeng, Pei Wang, Jingang Wang, Yunsen Xian, Weiran Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yutao Mou, Xiaoshuai Song, Keqing He 0001, Jingang Wang, Yunsen Xian, Weiran Xu
ACL (1)6
2023 MUSTIE: Multimodal Structural Transformer for Web Information Extraction
abstract
Qifan Wang, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Qifan Wang 0001, Jingang Wang, Xiaojun Quan, Fuli Feng, Zenglin Xu, Shaoliang Nie, Sinong Wang, Madian Khabsa, Hamed Firooz, Dongfang Liu
ACL (1)2
2023 FutureTOD: Teaching Future Knowledge to Pre-trained Language Model for Task-Oriented Dialogue
abstract
Weihao Zeng, Keqing He, Yejie Wang, Chen Zeng, Jingang Wang, Yunsen Xian, Weiran Xu. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Weihao Zeng 0003, Keqing He 0001, Yejie Wang, Jingang Wang, Yunsen Xian, Weiran Xu
ACL (1)5
2023 Seen to Unseen: Exploring Compositional Generalization of Multi-Attribute Controllable Dialogue Generation
abstract
Existing controllable dialogue generation work focuses on the single-attribute control and lacks generalization capability to out-of-distribution multiple attribute combinations.In this paper, we explore the compositional generalization for multi-attribute controllable dialogue generation where a model can learn from seen attribute values and generalize to unseen combinations.We propose a prompt-based disentangled controllable dialogue generation model, DCG.It learns attribute concept composition by generating attribute-oriented prompt vectors and uses a disentanglement loss to disentangle different attributes for better generalization.Besides, we design a unified reference-free evaluation framework for multiple attributes with different levels of granularities.Experiment results on two benchmarks prove the effectiveness of our method and the evaluation metric.* The first two authors contribute equally.Weiran Xu is the corresponding author.
Weihao Zeng 0003, Keqing He 0001, Ruotong Geng, Jingang Wang, Wei Wu 0014, Weiran Xu
ACL (1)5
2023 Lifting the Curse of Capacity Gap in Distilling Language Models
abstract
Chen Zhang, Yang Yang, Jiahao Liu, Jingang Wang, Yunsen Xian, Benyou Wang, Dawei Song. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Chen Zhang 0020, Yang Yang 0129, Jingang Wang, Yunsen Xian, Benyou Wang, Dawei Song 0001
ACL (1)4
2023 Large Language Models Meet Open-World Intent Discovery and Recognition: An Evaluation of ChatGPT
abstract
Xiaoshuai Song, Keqing He, Pei Wang, Guanting Dong, Yutao Mou, Jingang Wang, Yunsen Xian, Xunliang Cai, Weiran Xu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Xiaoshuai Song, Keqing He 0001, Guanting Dong 0001, Yutao Mou, Jingang Wang, Yunsen Xian, Weiran Xu
EMNLP6
2023 APrompt: Attention Prompt Tuning for Efficient Adaptation of Pre-trained Language Models
abstract
Qifan Wang, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Qifan Wang 0001, Yuning Mao, Jingang Wang, Hanchao Yu, Shaoliang Nie, Sinong Wang, Fuli Feng, Lifu Huang, Xiaojun Quan, Zenglin Xu, Dongfang Liu
EMNLP3
2023 Multi-Task Transformer with Relation-Attention and Type-Attention for Named Entity Recognition
abstract
Named entity recognition (NER) is an important research problem in natural language processing. There are three types of NER tasks, including flat, nested and discontinuous entity recognition. Most previous sequential labeling models are task-specific, while recent years have witnessed the rising of generative models due to the advantage of unifying all NER tasks into the seq2seq model framework. Although achieving promising performance, our pilot studies demonstrate that existing generative models are ineffective at detecting entity boundaries and estimating entity types. This paper proposes a multi-task Transformer, which incorporates an entity boundary detection task into the named entity recognition task. More concretely, we achieve entity boundary detection by classifying the relations between tokens within the sentence. To improve the accuracy of entity-type mapping during decoding, we adopt an external knowledge base to calculate the prior entity-type distributions and then incorporate the information into the model via the self and cross-attention mechanisms. We perform experiments on an extensive set of NER benchmarks, including two flat, three nested, and three discontinuous NER datasets. Experimental results show that our approach considerably improves the generative NER model’s performance.
Ying Mo, Hongyin Tang, Qifan Wang 0001, Zenglin Xu, Jingang Wang, Wei Wu 0014, Zhoujun Li 0001
ICASSP6
2023 Sea Surface Object Detection Based on Background Dynamic Perception and Cross-Layer Semantic Interaction
abstract
Sea surface object detection plays an important role in the coastal defense monitoring system. Existing target detection methods mostly lack the adaptive perception of background changes. In addition, these methods fail to further integrate and interact with the multi-layer features extracted from the deep backbone network. To address these two issues, we first propose a Background Dynamic Perception module, which uses environmental information as an auxiliary. We train the detector to dynamically capture the background changes through a multi-task learning framework. Moreover, we propose a Cross-Layer Semantic Interaction module, which can achieve cross-layer interaction and reduce information loss. Based on the above modules, we propose a sea surface object detection network. To verify the performance, we collected real sea surface data and built a sea surface object dataset. Experimental results demonstrate that our method achieves 74.4% AP on the dataset, outperforming the latest methods.
Songbin Li, Xiangzhi Yang, Jingang Wang
ICME3
2023 LUNA: Language as Continuing Anchors for Referring Expression Comprehension
abstract
Referring expression comprehension aims to localize a natural language description in an image. Using location priors to help reduce inaccuracies in cross-modal alignments is the state of the art for CNN-based methods tackling this problem. Recent Transformer-based models cast aside this idea, making the case for steering away from hand-designed components. In this work, we propose LUNA, which uses language as continuing anchors to guide box prediction in a Transformer decoder, and thus show that language-guided location priors can be effectively exploited in a Transformer-based architecture. Our method first initializes an anchor box from the input expression via a small "proto-decoder,'' and then uses this anchor and its refined successors as location guidance in a modified Transformer decoder. At each decoder layer, the anchor box is first used as a query for gathering multi-modal context, and then updated based on the gathered context (producing the next, refined anchor). In the end, a lightweight assessment pathway evaluates the quality of all produced anchors, yielding the final prediction in a dynamic way. This approach allows box decoding to be conditioned on learned anchors, which facilitates accurate grounding, as we shown in the experiments. Our method outperforms existing state-of-the-art methods on the datasets of ReferIt Game, RefCOCO/+/g, and Flickr30K Entities.
Yaoyuan Liang, Zhao Yang 0002, Yansong Tang, Jiashuo Fan, Ziran Li, Jingang Wang, Philip Torr 0001, Shao-Lun Huang
ACM Multimedia6
2023 Maritime Radar Target Detection in Sea Clutter Based on CNN With Dual-Perspective Attention
abstract
Radar-based maritime target detection plays an important role in ocean monitoring. Considering the practical application, pulse-compression radar is widely used in terms of civilian offshore surface target detection. The existence of sea clutter will greatly interfere the detection performance of pulse-compression radar. This leads to the low detection performance of traditional algorithms like constant false alarm rate (CFAR). Deep learning methods have made strides in many fields recently, such as natural language processing and speech recognition. Inspired by this idea, we propose a maritime radar target detection method in sea clutter based on convolution neural network (CNN) and dual-perspective attention (DPA). The proposed method first encodes the radar echo in high-dimensional space and then extracts the correlation features from the global and local perspectives through the attention mechanism. We deployed the X-band pulse-compression radar on the coast of Hainan, China, and collected a lot of measured data. Experimental results demonstrate that the detection performance of our method outperforms the traditional CFAR methods and the latest deep learning-based methods. In the measured dataset, our proposed method can reach a detection probability of 93.59% under a false alarm rate (FAR) of$1e-3$, reaching the practical application level.
Jingang Wang, Songbin Li
IEEE Geosci. Remote. Sens. Lett.1
2023 Solve the Puzzle of Instance Segmentation in Videos: A Weakly Supervised Framework With Spatio-Temporal Collaboration
abstract
Instance segmentation in videos, which aims to segment and track multiple objects in video frames, has garnered a flurry of research attention in recent years. In this paper, we present a novel weakly supervised framework with Spatio-Temporal Collaboration for instance Segmentation in videos, namely STC-Seg. Concretely, STC-Seg demonstrates four contributions. First, we leverage the complementary representations from unsupervised depth estimation and optical flow to produce effective pseudo-labels for training deep networks and predicting high-quality instance masks. Second, to enhance the mask generation, we devise a puzzle loss, which enables end-to-end training using box-level annotations. Third, our tracking module jointly utilizes bounding-box diagonal points with spatio-temporal discrepancy to model movements, which largely improves the robustness to different object appearances. Finally, our framework is flexible and enables image-level instance segmentation methods to operate the video-level task. We conduct an extensive set of experiments on the KITTI MOTS and YT-VIS datasets. Experimental results demonstrate that our method achieves strong performance and even outperforms fully supervised TrackR-CNN and MaskTrack R-CNN. We believe that STC-Seg can be a valuable addition to the community, as it reflects the tip of an iceberg about the innovative opportunities in the weakly supervised paradigm for instance segmentation in videos.
Liqi Yan, Qifan Wang 0001, Siqi Ma 0005, Jingang Wang, Changbin Yu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Detection of Generative Linguistic Steganography Based on Explicit and Latent Text Word Relation Mining Using Deep Learning
abstract
Covert communication channels can be easily constructed using text steganography based on social media. Offenders can easily utilize these channels to engage in various criminal activities, which brings great challenges in maintaining the security of cyberspace. Among the text information hiding methods, generative linguistic steganography poses the biggest threat to network security because it does not need the original carrier and has high embedding efficiency. The existing generative linguistic steganalysis methods fail to deeply mine text word relation, hence the detection performance is relatively unsatisfactory. In this article, we prove that there is explicit and latent steganography-sensitive text word relation. Based on this, we propose a generative linguistic steganalysis method based onExplicit andLatent text word relationMining, namedELM. First, we employ a distributed readin module to convert words into real number vectors. Then, MRA (Mining Relation by Attentions) is proposed to mine the explicit and latent text word relation. Finally, global adaptive classification module is presented to exploit the mined relation feature to predict whether secret information is embedded in the current text segment. Experimental results demonstrate that the detection performance of ELM is better than the existing generative linguistic steganalysis methods.
Songbin Li, Jingang Wang, Peng Liu 0046
IEEE Trans. Dependable Secur. Comput.2
2022 Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification
abstract
Shengding Hu, Ning Ding, Huadong Wang, Zhiyuan Liu, Jingang Wang, Juanzi Li, Wei Wu, Maosong Sun. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Shengding Hu, Ning Ding 0002, Zhiyuan Liu 0001, Jingang Wang, Juan-Zi Li, Wei Wu 0014, Maosong Sun 0001
ACL (1)5
2022 Unified Knowledge Prompt Pre-training for Customer Service Dialogues
abstract
Dialogue bots have been widely applied in customer service scenarios to provide timely and user-friendly experience. These bots must classify the appropriate domain of a dialogue, understand the intent of users, and generate proper responses. Existing dialogue pre-training models are designed only for several dialogue tasks and ignore weakly-supervised expert knowledge in customer service dialogues. In this paper, we propose a novel unified knowledge prompt pre-training framework, UFA (Unified Model F or All Tasks), for customer service dialogues. We formulate all the tasks of customer service dialogues as a unified text-to-text generation task and introduce a knowledge-driven prompt strategy to jointly learn from a mixture of distinct dialogue tasks. We pre-train UFA on a large-scale Chinese customer service corpus collected from practical scenarios and get significant improvements on both natural language understanding (NLU) and natural language generation (NLG) benchmarks.
Keqing He 0001, Jingang Wang, Chaobo Sun, Wei Wu 0014
CIKM2
2022 CLOWER: A Pre-trained Language Model with Contrastive Learning over Word and Character Representations
abstract
Pre-trained Language Models (PLMs) have achieved remarkable performance gains across numerous downstream tasks in natural language understanding. Various Chinese PLMs have been successively proposed for learning better Chinese language representation. However, most current models use Chinese characters as inputs and are not able to encode semantic information contained in Chinese words. While recent pre-trained models incorporate both words and characters simultaneously, they usually suffer from deficient semantic interactions and fail to capture the semantic relation between words and characters. To address the above issues, we propose a simple yet effective PLM CLOWER, which adopts the Contrastive Learning Over Word and charactER representations. In particular, CLOWER implicitly encodes the coarse-grained information (i.e., words) into the fine-grained representations (i.e., characters) through contrastive learning on multi-grained information. CLOWER is of great value in realistic scenarios since it can be easily incorporated into any existing fine-grained based PLMs without modifying the production pipelines. Extensive experiments conducted on a range of downstream tasks demonstrate the superior performance of CLOWER over several state-of-the-art baselines.
Borun Chen, Hongyin Tang, Jiahao Bu, Kai Zhang 0038, Jingang Wang, Qifan Wang 0001, Hai-Tao Zheng 0002, Wei Wu 0014, Liqian Yu
COLING5
2022 Generalized Intent Discovery: Learning from Open World Dialogue System
abstract
Traditional intent classification models are based on a pre-defined intent set and only recognize limited in-domain (IND) intent classes. But users may input out-of-domain (OOD) queries in a practical dialogue system. Such OOD queries can provide directions for future improvement. In this paper, we define a new task, Generalized Intent Discovery (GID), which aims to extend an IND intent classifier to an open-world intent set including IND and OOD intents. We hope to simultaneously classify a set of labeled IND intent classes while discovering and recognizing new unlabeled OOD types incrementally. We construct three public datasets for different application scenarios and propose two kinds of frameworks, pipeline-based and end-to-end for future work. Further, we conduct exhaustive experiments and qualitative analysis to comprehend key challenges and provide new guidance for future GID research.
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Yi Huang 0017, Junlan Feng, Weiran Xu
COLING5
2022 Structural Bias for Aspect Sentiment Triplet Extraction
abstract
Structural bias has recently been exploited for aspect sentiment triplet extraction (ASTE) and led to improved performance. On the other hand, it is recognized that explicitly incorporating structural bias would have a negative impact on efficiency, whereas pretrained language models (PLMs) can already capture implicit structures. Thus, a natural question arises: Is structural bias still a necessity in the context of PLMs? To answer the question, we propose to address the efficiency issues by using an adapter to integrate structural bias in the PLM and using a cheap-to-compute relative position structure in place of the syntactic dependency structure. Benchmarking evaluation is conducted on the SemEval datasets. The results show that our proposed structural adapter is beneficial to PLMs and achieves state-of-the-art performance over a range of strong baselines, yet with a light parameter demand and low latency. Meanwhile, we give rise to the concern that the current evaluation default with data of small scale is under-confident. Consequently, we release a large-scale dataset for ASTE. The results on the new dataset hint that the structural adapter is confidently effective and efficient to a large scale. Overall, we draw the conclusion that structural bias shall still be a necessity even with PLMs.
Chen Zhang 0020, Fang Ma, Jingang Wang, Wei Wu 0014, Dawei Song 0001
COLING4
2022 One-Shot Detection of Malicious TLS Traffic
abstract
Network security protocols (such as Transport Layer Security, TLS) are increasingly used by hackers to evade detection. The detection of encrypted malicious traffic is becoming a critical task for cyber security. To accomplish this task, researchers have proposed several machine learning or deep learning based methods. However, existing methods require a huge amount of malicious flows as annotated samples, which is difficult or even impossible to obtain for emerging attacks. In this paper, we propose a novel method to detect malicious TLS traffic with only one sample, i.e., one annotated malicious flow. Our observation is that there are a large number of benign HTTPS (HTTP+TLS) flows in the network, and these flows are similar but with subtle differences in behavior. If a model can efficiently distinguish different HTTPS flows by traffic behavior, it should know how to extract inherent characteristics of TLS traffic. Based on the observation, we visit Alexa top websites and train ResNet models in conjunction with statistical traffic data. The obtained conjunction models are used as the feature extractor for malicious TLS flows. To achieve one-shot classification, we further train a feed-forward neural network to transfer the extracted features to a classification decision boundary of the support vector machine (SVM). We test our approach with a thorough set of experiments, and the experimental results show that our approach significantly outperforms the baseline methods.
Gaofeng He, Qianfeng Wei, Jingang Wang, Haiting Zhu, Bingfeng Xu
CSCWD3
2022 VIRT: Improving Representation-based Text Matching via Virtual Interaction
abstract
Text matching is a fundamental research problem in natural language understanding.Interaction-based approaches treat the text pair as a single sequence and encode it through cross encoders, while representation-based models encode the text pair independently with siamese or dual encoders.Interactionbased models require dense computations and thus are impractical in real-world applications.Representation-based models have become the mainstream paradigm for efficient text matching.However, these models suffer from severe performance degradation due to the lack of interactions between the pair of texts.To remedy this, we propose a Virtual InteRacTion mechanism (VIRT) for improving representation-based text matching while maintaining its efficiency.In particular, we introduce an interactive knowledge distillation module that is only applied during training.It enables deep interaction between texts by effectively transferring knowledge from the interaction-based model.A light interaction strategy is designed to fully leverage the learned interactive knowledge.Experimental results on six text matching benchmarks demonstrate the superior performance of our method over several state-of-the-art representationbased models.We further show that VIRT can be integrated into existing methods as plugins to lift their performances.
Yang Yang 0129, Hongyin Tang, Qifan Wang 0001, Jingang Wang, Tong Xu 0001, Wei Wu 0014, Enhong Chen
EMNLP6
2022 XPrompt: Exploring the Extreme of Prompt Tuning
abstract
Prompt tuning learns soft prompts to condition the frozen Pre-trained Language Models (PLMs) for performing downstream tasks in a parameter-efficient manner.While prompt tuning has gradually reached the performance level of fine-tuning as the model scale increases, there is still a large performance gap between prompt tuning and fine-tuning for models of moderate and small scales (typically less than 11B parameters).In this paper, we empirically show that the trained prompt tokens can have a negative impact on a downstream task and thus degrade its performance.To bridge the gap, we propose a novel PROMPT tuning model with an eXtremely small scale (XPROMPT) under the regime of lottery tickets hypothesis.Specifically, XPROMPT eliminates the negative prompt tokens at different granularity levels through a hierarchical structured pruning, yielding a more parameter-efficient prompt yet with a competitive performance.Comprehensive experiments are carried out on the SuperGLUE tasks, and the results indicate that XPROMPT is able to close the performance gap at smaller model scales. 1 Recently, Prompt-Tuning (Lester et al., 2021; Liu et al., 2021b) has been proposed to address this issue by prepending a soft prompt to the input and only updating the parameters of prompt tokens during tuning.Prompt-Tuning provides a parameter-efficient alternative to fine-tuning, since the scale of the soft prompt is tens of thousand smaller.It is also conceptually simpler and more flexible than other parameter-efficient tuning methods (such as Adapters), that require intrusive modifications to transformer layers (Houlsby et al., 2019;Guo et al., 2021).Using fewer tunable parameters, prompt tuning achieves competitive performance to fine-tuning with the increase of the model scale.However, there is still a large performance gap between prompt tuning and fine-tuning for models of smaller scales (as shown in Figure 1).This paper aims to fill the gap, from the perspective of the lottery tickets hypothesis (LTH) (Frankle and Carbin, 2019).We are motivated by an observation that, on a specific task, not all prompt tokens contribute equally to the task performance, while certain prompt tokens may even bring a negative
Fang Ma, Chen Zhang 0020, Jingang Wang, Qifan Wang 0001, Wei Wu 0014, Xiaojun Quan, Dawei Song 0001
EMNLP4
2022 Watch the Neighbors: A Unified K-Nearest Neighbor Contrastive Learning Framework for OOD Intent Discovery
abstract
Discovering out-of-domain (OOD) intent is important for developing new skills in taskoriented dialogue systems.The key challenges lie in how to transfer prior in-domain (IND) knowledge to OOD clustering, as well as jointly learn OOD representations and cluster assignments.Previous methods suffer from indomain overfitting problem, and there is a natural gap between representation learning and clustering objectives.In this paper, we propose a unified K-nearest neighbor contrastive learning framework to discover OOD intents.Specifically, for IND pre-training stage, we propose a KCL objective to learn inter-class discriminative features, while maintaining intraclass diversity, which alleviates the in-domain overfitting problem.For OOD clustering stage, we propose a KCC method to form compact clusters by mining true hard negative samples, which bridges the gap between clustering and representation learning.Extensive experiments on three benchmark datasets show that our method achieves substantial improvements over the state-of-the-art methods. 1
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Weiran Xu
EMNLP5
2022 UniNL: Aligning Representation Learning with Scoring Function for OOD Detection via Unified Neighborhood Learning
abstract
Detecting out-of-domain (OOD) intents from user queries is essential for avoiding wrong operations in task-oriented dialogue systems.The key challenge is how to distinguish indomain (IND) and OOD intents.Previous methods ignore the alignment between representation learning and scoring function, limiting the OOD detection performance.In this paper, we propose a unified neighborhood learning framework (UniNL) to detect OOD intents.Specifically, we design a K-nearest neighbor contrastive learning (KNCL) objective for representation learning and introduce a KNNbased scoring function for OOD detection.We aim to align representation learning with scoring function.Experiments and analysis on two benchmark datasets show the effectiveness of our method.1
Yutao Mou, Keqing He 0001, Yanan Wu 0002, Jingang Wang, Wei Wu 0014, Weiran Xu
EMNLP5
2022 Making Pretrained Language Models Good Long-tailed Learners
abstract
Prompt-tuning has shown appealing performance in few-shot classification by virtue of its capability in effectively exploiting pre-trained knowledge.This motivates us to check the hypothesis that prompt-tuning is also a promising choice for long-tailed classification, since the tail classes are intuitively few-shot ones.To achieve this aim, we conduct empirical studies to examine the hypothesis.The results demonstrate that prompt-tuning makes pretrained language models at least good long-tailed learners.For intuitions on why prompt-tuning can achieve good performance in long-tailed classification, we carry out in-depth analyses by progressively bridging the gap between prompttuning and commonly used finetuning.The summary is that the classifier structure and parameterization form the key to making good long-tailed learners, in comparison with the less important input structure.Finally, we verify the applicability of our finding to few-shot classification. 1
Chen Zhang 0020, Jingang Wang, Wei Wu 0014, Dawei Song 0001
EMNLP3
2022 Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language Models
abstract
Despite the great success of pre-trained language models (PLMs) in a large set of natural language processing (NLP) tasks, there has been a growing concern about their security in real-world applications. Backdoor attack, which poisons a small number of training samples by inserting backdoor triggers, is a typical threat to security. Trained on the poisoned dataset, a victim model would perform normally on benign samples but predict the attacker-chosen label on samples containing pre-defined triggers. The vulnerability of PLMs under backdoor attacks has been proved with increasing evidence in the literature. In this paper, we present several simple yet effective training strategies that could effectively defend against such attacks. To the best of our knowledge, this is the first work to explore the possibility of backdoor-free adaptation for PLMs. Our motivation is based on the observation that, when trained on the poisoned dataset, the PLM's adaptation follows a strict order of two stages: (1) a moderate-fitting stage, where the model mainly learns the major features corresponding to the original task instead of subsidiary features of backdoor triggers, and (2) an overfitting stage, where both features are learned adequately. Therefore, if we could properly restrict the PLM's adaptation to the moderate-fitting stage, the model would neglect the backdoor triggers but still achieve satisfying performance on the original task. To this end, we design three methods to defend against backdoor attacks by reducing the model capacity, training epochs, and learning rate, respectively. Experimental results demonstrate the effectiveness of our methods in defending against several representative NLP backdoor attacks. We also perform visualization-based analysis to attain a deeper understanding of how the model learns different features, and explore the effect of the poisoning ratio. Finally, we explore whether our methods could defend against backdoor attacks for the pre-trained CV model. The codes are publicly available at https://github.com/thunlp/Moderate-fitting.
Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Yangdong Deng, Zhiyuan Liu 0001, Jingang Wang, Wei Wu 0014, Maosong Sun 0001, Ming Gu 0001
NeurIPS9
2022 Personalized Abstractive Opinion Tagging
abstract
An opinion tag is a sequence of words on a specific aspect of a product or service. Opinion tags reflect key characteristics of product reviews and help users quickly understand their content in e-commerce portals. The task of abstractive opinion tagging has previously been proposed to automatically generate a ranked list of opinion tags for a given review. However, current models for opinion tagging are not personalized, even though personalization is an essential ingredient of engaging user interactions, especially in e-commerce. In this paper, we focus on the task of personalized abstractive opinion tagging. There are two main challenges when developing models for the end-to-end generation of personalized opinion tags: sparseness of reviews and difficulty to integrate multi-type signals, i.e., explicit review signals and implicit behavioral signals. To address these challenges, we propose an end-to-end model, named POT, that consists of three main components: (1) a review-based explicit preference tracker component based on a hierarchical heterogeneous review graph to track user preferences from reviews; (2)a behavior-based implicit preference tracker component using a heterogeneous behavior graph to track the user preferences from implicit behaviors; and (3) a personalized rank-aware tagging component to generate a ranked sequence of personalized opinion tags. In our experiments, we evaluate POT on a real-world dataset collected from e-commerce platforms and the results demonstrate that it significantly outperforms strong baselines.
Mengxue Zhao, Yang Yang 0129, Jingang Wang, Wei Wu 0014, Pengjie Ren, Maarten de Rijke, Zhaochun Ren
SIGIR4
2022 General Frame-Wise Steganalysis of Compressed Speech Based on Dual-Domain Representation and Intra-Frame Correlation Leaching
abstract
Frame-wise steganalysis is of significance for active steganography defense. By frame-wise detection, we can accurately find the embedding position of secret information and destroy the covert channel further. However, there is currently no research specifically aiming at frame-wise steganalysis of low-bit-rate compressed speech. Besides, most of the existing steganalysis methods are specifically designed for a specific category of steganography methods. They are difficult to apply to practical scenarios where the steganography algorithms are uncertain. In this paper, a general frame-wise steganalysis method for low-bit-rate compressed speech is proposed. To extract rich feature from a speech frame, we propose a dual-domain representation, which conducts feature extraction both in the compressed domain and the decoded time domain. In addition, we propose an efficient steganalysis network named Stegaformer to leach the intra-frame correlation from the obtained representation to enable steganalysis. In Stegaformer, an adaptive local correlation enhancement module is introduced to effectively models the local characteristics, which compensates for the drawback of traditional Transformer-based models. Experimental results show that our method performs better than the existing steganalysis methods in detecting multiple steganography methods for a speech frame.
Songbin Li, Jingang Wang, Peng Liu 0046
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Improving Document Representations by Generating Pseudo Query Embeddings for Dense Retrieval
abstract
Hongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Fuzheng Zhang, Wei Wu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hongyin Tang, Xingwu Sun, Beihong Jin, Jingang Wang, Wei Wu 0014
ACL/IJCNLP (1)4
2021 ASAP: A Chinese Review Dataset Towards Aspect Category Sentiment Analysis and Rating Prediction
abstract
Jiahao Bu, Lei Ren, Shuang Zheng, Yang Yang, Jingang Wang, Fuzheng Zhang, Wei Wu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Jiahao Bu, Shuang Zheng 0007, Yang Yang 0129, Jingang Wang, Wei Wu 0014
NAACL-HLT5
2021 Detection of Multiple Steganography Methods in Compressed Speech Based on Code Element Embedding, Bi-LSTM and CNN With Attention Mechanisms
abstract
Steganographic algorithms in low-bit-rate compressed speech bring convenience to realize covert communication, meanwhile result in safety issues. The existing steganalysis methods are normally designed for one specific category of steganographic methods, thus lacking generalization capability. In this paper, we propose a general steganalysis method based on code element (CE) embedding, Bi-LSTM and CNN with attention mechanisms. Firstly, CEs in each frame are converted to a multi-hot vector. And each multi-hot vector will be mapped into a fixed-length embedding vector to get a more compact representation by utilizing dictionaries. Then, Bi-LSTM and CNN are applied to extract the contextual information and the local characteristics respectively of these embedding vectors. In addition, the attention mechanisms are introduced in different layers of the network to assign different weights to the output feature within each layer. Finally, the prediction results can be generated by the fully connected layer. Experimental results show that our method performs better than the existing steganalysis methods for detecting multiple steganography methods in the low-bit-rate compressed speech streams.
Songbin Li, Jingang Wang, Peng Liu 0046, Miao Wei, Qiandong Yan
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Query-aware Tip Generation for Vertical Search
abstract
As a concise form of user reviews, tips have unique advantages to explain the search results, assist users' decision making, and further improve user experience in vertical search scenarios. Existing work on tip generation does not take query into consideration, which limits the impact of tips in search scenarios. To address this issue, this paper proposes a query-aware tip generation framework, integrating query information into encoding and subsequent decoding processes. Two specific adaptations of Transformer and Recurrent Neural Network (RNN) are proposed. For Transformer, the query impact is incorporated into the self-attention computation of both the encoder and the decoder. As for RNN, the query-aware encoder adopts a selective network to distill query-relevant information from the review, while the query-aware decoder integrates the query information into the attention computation during decoding. The framework consistently outperforms the competing methods on both public and real-world industrial datasets. Last but not least, online deployment experiments on Dianping demonstrate the advantage of the proposed framework for tip generation as well as its online business values.
Yang Yang 0129, Junmei Hao, Canjia Li, Jingang Wang, Peixu Hou, Zhongyuan Wang 0006
CIKM5
2020 A Hybrid Discriminative Mixture Model for Cumulative Citation Recommendation
abstract
This paper explores Cumulative Citation Recommendation (CCR) for Knowledge Base Acceleration (KBA). The CCR task aims to detect potential citations of a set of target entities with priorities from a volume of temporally-ordered stream corpus. Previous approaches for CCR that build an individual relevance model for each entity fail to deal with unseen entities without annotation. A compromised solution is to build a global entity-unspecific model for all entities without respect to the relationship information among entities, which cannot guarantee achieving a satisfactory result for each entity. Moreover, most previous methods can not adequately exploit prior knowledge embedded in entities or documents due to considering all kinds of features indifferently. In this paper, we propose a novel entity and document class-dependent discriminative mixture model by introducing one intermediate layer to model the correlation between entity-document pairs and hybrid latent entity-document classes. The model can better adjust to different types of entities and documents, and achieve better performance when dealing with a broad range of entity and document classes. An extensive set of experiments has been conducted on two offical datasets, and the experimental results demonstrate that the proposed model can achieve the state-of-the-art performance.
Lerong Ma, Lejian Liao, Jingang Wang
IEEE Trans. Knowl. Data Eng.4
2019 Earlier Attention? Aspect-Aware LSTM for Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis (ABSA) aims to predict fine-grained sentiments of comments with respect to given aspect terms or categories. In previous ABSA methods, the importance of aspect has been realized and verified. Most existing LSTM-based models take aspect into account via the attention mechanism, where the attention weights are calculated after the context is modeled in the form of contextual vectors. However, aspect-related information may be already discarded and aspect-irrelevant information may be retained in classic LSTM cells in the context modeling process, which can be improved to generate more effective context representations. This paper proposes a novel variant of LSTM, termed as aspect-aware LSTM (AA-LSTM), which incorporates aspect information into LSTM cells in the context modeling stage before the attention mechanism. Therefore, our AA-LSTM can dynamically produce aspect-aware contextual representations. We experiment with several representative LSTM-based models by replacing the classic LSTM cells with the AA-LSTM cells. Experimental results on SemEval-2014 Datasets demonstrate the effectiveness of AA-LSTM.
Lejian Liao, Dandan Song 0005, Jingang Wang, Zhongyuan Wang 0006, Heyan Huang
IJCAI4
2018 A Multi-Task Learning Approach for Improving Product Title Compression with User Search Log Data
abstract
It is a challenging and practical research problem to obtain effective compression of lengthy product titles for E-commerce. This is particularly important as more and more users browse mobile E-commerce apps and more merchants make the original product titles redundant and lengthy for Search Engine Optimization. Traditional text summarization approaches often require a large amount of preprocessing costs and do not capture the important issue of conversion rate in E-commerce. This paper proposes a novel multi-task learning approach for improving product title compression with user search log data. In particular, a pointer network-based sequence-to-sequence approach is utilized for title compression with an attentive mechanism as an extractive method and an attentive encoder-decoder approach is utilized for generating user search queries. The encoding parameters (i.e., semantic embedding of original titles) are shared among the two tasks and the attention distributions are jointly optimized. An extensive set of experiments with both human annotated data and online deployment demonstrate the advantage of the proposed research for both compression qualities and online business values.
Jingang Wang, Long Qiu, Sheng Li 0017, Jun Lang 0001, Luo Si, Man Lan
AAAI1
2018 An Adversarial Joint Learning Model for Low-Resource Language Semantic Textual Similarity
Man Lan, Yuanbin Wu, Jingang Wang, Long Qiu, Sheng Li 0017, Jun Lang 0001, Luo Si
ECIR4
2017 PSVM: a preference-enhanced SVM model using preference data for classification
Lerong Ma, Lejian Liao, Jingang Wang
Sci. China Inf. Sci.4
2016 Cold Start Cumulative Citation Recommendation for Knowledge Base Acceleration
Jingang Wang, Jingtian Jiang, Lejian Liao, Chin-Yew Lin
ECIR1
2016 Supervised Local Contexts Aggregation for Effective Session Search
Jingang Wang, Tao Wu 0019, Pengjie Ren, Zhumin Chen, Luo Si
ECIR2
2015 LDTM: A Latent Document Type Model for Cumulative Citation Recommendation
abstract
This paper studies Cumulative Citation Recommendation (CCR) -given an entity in Knowledge Bases, how to effectively detect its potential citations from volume text streams.Most previous approaches treated all kinds of features indifferently to build a global relevance model, in which the prior knowledge embedded in documents cannot be exploited adequately.To address this problem, we propose a latent document type discriminative model by introducing a latent layer to capture the correlations between documents and their underlying types.The model can better adjust to different types of documents and yield flexible performance when dealing with a broad range of document types.An extensive set of experiments has been conducted on TREC-KBA-2013 dataset, and the results demonstrate that this model can yield a significant performance gain in recommendation quality as compared to the state-of-the-art.
Jingang Wang, Lejian Liao, Luo Si, Chin-Yew Lin
EMNLP1
2015 An Entity Class-Dependent Discriminative Mixture Model for Cumulative Citation Recommendation
abstract
This paper studies Cumulative Citation Recommendation (CCR) for Knowledge Base Acceleration (KBA). The CCR task aims to detect potential citations of a set of target entities with priorities from a volume of temporally-ordered stream corpus. Previous approaches for CCR that build an individual relevance model for each entity fail to handle unseen entities without annotation. A baseline solution is to build a global entity-unspecific model for all entities regardless of the relationship information among entities, which cannot guarantee to achieve satisfactory result for each entity. In this paper, we propose a novel entity class-dependent discriminative mixture model by introducing a latent entity class layer to model the correlations between entities and latent entity classes. The model can better adjust to different types of entities and achieve better performance when dealing with a broad range of entities. An extensive set of experiments has been conducted on TREC-KBA-2013 dataset, and the experimental results demonstrate that the proposed model can achieve the state-of-the-art performance.
Jingang Wang, Qifan Wang 0001, Luo Si, Lejian Liao, Chin-Yew Lin
SIGIR1
2015 Resorting Relevance Evidences to Cumulative Citation Recommendation for Knowledge Base Acceleration
Jingang Wang, Lejian Liao, Lerong Ma, Chin-Yew Lin, Yong Rui
WAIM1
2013 A Self-healing Framework for QoS-Aware Web Service Composition via Case-Based Reasoning
Guoqiang Li 0003, Lejian Liao, Jingang Wang, Fuzhen Sun, Guangcheng Liang
APWeb4