EDBT 2026 Demo / reviewers in the wild / expert
Yequan Wang
dblp:188/9082
· DBLP profile ↗
38ranked-venue papers
4as first author
34since 2021 · last 2026
0000-0001-7530-6125ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 2 first-author · 28 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Harmonizing the Past, Present, and Future: A Null-Space Constrained Region-Specific Method for Continual Learning in LLMsabstractJinhui Chen, Shizhu He, Xingchang Yang, Huanxuan Liao, Yequan Wang, Xiangwen Liao, Wenhao Teng, Kang Liu, Jun Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shizhu He, Xingchang Yang, Huanxuan Liao, Yequan Wang, Xiangwen Liao, Wenhao Teng, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 5 |
| 2026 | If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMsabstractSiqi Fan, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu, Peng Han, Shuo Shang, Aixin Sun, Yequan Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Siqi Fan 0001, Xiusheng Huang, Yiqun Yao, Xuezhi Fang, Kang Liu 0001, Peng Han 0005, Shuo Shang, Aixin Sun, Yequan Wang |
ACL (1) | 9 |
| 2026 | Theory-optimal Quantization Based on FlatnessabstractXiusheng Huang, Zhe Li, Xuanwu Yin, Lu Wang, Yequan Wang, Dong Li, Emad Barsoum, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiusheng Huang, Xuanwu Yin, Lu Wang 0029, Yequan Wang, Dong Li 0025, Emad Barsoum, Kang Liu 0001 |
ACL (1) | 5 |
| 2026 | Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMsabstractWhile recent self-training approaches have reduced reliance on human-labeled data for aligning LLMs, they still face critical limitations: (i) sensitivity to synthetic data quality, leading to instability and bias amplification in iterative training; (ii) ineffective optimization due to a diminishing gap between positive and negative responses over successive training iterations.In this paper, we propose Team-based self-Play with dual Adaptive Weighting (TPAW), a novel self-play algorithm designed to improve alignment in a fully self-supervised setting.TPAW adopts a team-based framework in which the current policy model both collaborates with and competes against historical checkpoints, promoting more stable and efficient optimization.To further enhance learning, we design two adaptive weighting mechanisms: (i) a response reweighting scheme that adjusts the importance of target responses, and (ii) a player weighting strategy that dynamically modulates each team member's contribution during training.Initialized from a SFT model, TPAW iteratively refines alignment without requiring additional human supervision.Experimental results demonstrate that TPAW consistently outperforms existing baselines across various base models and LLM benchmarks. Yigeng Zhou, Zesheng Shi, Yequan Wang |
ACL (1) | 4 |
| 2026 | Spectral Disentanglement: Rank-Aware Task Adaptation for Rehearsal-free Continual Learning in LLMsabstractHuanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao, Kang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Huanxuan Liao, Shizhu He, Yupu Hao, Yequan Wang, Wenhao Teng, Xiangwen Liao, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 4 |
| 2026 | Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under ConflictsabstractRetrieval-Augmented Generation (RAG) has become a standard paradigm for grounding Large Language Models (LLMs) with external knowledge.However, RAG performance often degrades substantially when faced with noisy, outdated, or conflicting retrieved information.In this work, we empirically demonstrate that Prior-Guided Reasoning-a strategy that explicitly elicits the model's parametric knowledge as prior information to guide reasoning on retrieved documents-effectively mitigates the impact of external conflicts.Building on this, we propose BrPr (Bernoulligated reinforcement learning for Prior-Guided reasoning), a framework that achieves robust performance across varying degrees of external inconsistency.Furthermore, by employing a Bernoulli-gated dropout mechanism during training, BrPr distills the prior-driven reasoning capability into the model parameters, enabling efficient latent reasoning without explicit prior generation.The experimental results demonstrate that BrPr consistently exhibits superior robustness to external conflicts and noise. Xiaowei Yuan, Ziyang Huang 0005, Zhao Yang 0004, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 4 |
| 2025 | Knowledge Editing with Dynamic Knowledge Graphs for Multi-Hop Question AnsweringabstractMulti-hop question answering (MHQA) poses a significant challenge for large language models (LLMs) due to the extensive knowledge demands involved. Knowledge editing, which aims to precisely modify the LLMs to incorporate specific knowledge without negatively impacting other unrelated knowledge, offers a potential solution for addressing MHQA challenges with LLMs. However, current solutions struggle to effectively resolve issues of knowledge conflicts. Most parameter-preserving editing methods are hindered by inaccurate retrieval and overlook secondary editing issues, which can introduce noise into the reasoning process of LLMs. In this paper, we introduce KEDKG, a novel knowledge editing method that leverages a dynamic knowledge graph for MHQA, designed to ensure the reliability of answers. KEDKG involves two primary steps: dynamic knowledge graph construction and knowledge graph augmented generation. Initially, KEDKG autonomously constructs a dynamic knowledge graph to store revised information while resolving potential knowledge conflicts. Subsequently, it employs a fine-grained retrieval strategy coupled with an entity and relation detector to enhance the accuracy of graph retrieval for LLM generation. Experimental results on benchmarks show that KEDKG surpasses previous state-of-the-art models, delivering more accurate and reliable answers in environments with dynamic information. Yigeng Zhou, Jing Li 0034, Yequan Wang, Xuebo Liu 0002, Daojing He, Fangming Liu, Min Zhang 0005 |
AAAI | 4 |
| 2025 | Multi-Modality Expansion and Retention for LLMs through Parameter Merging and DecouplingabstractJunlin Li, Guodong Du, Jing Li, Sim Kuan Goh, Wenya Wang, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guodong Du 0002, Jing Li 0034, Sim Kuan Goh, Wenya Wang 0001, Yequan Wang, Fangming Liu, Ho-Kin Tang, Saleh Alharbi, Daojing He, Min Zhang 0005 |
ACL (1) | 6 |
| 2025 | Exploiting Contextual Knowledge in LLMs through V-usable Information based Layer EnhancementabstractXiaowei Yuan, Zhao Yang, Ziyang Huang, Yequan Wang, Siqi Fan, Yiming Ju, Jun Zhao, Kang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xiaowei Yuan, Zhao Yang 0004, Ziyang Huang 0005, Yequan Wang, Siqi Fan 0001, Yiming Ju, Jun Zhao 0001, Kang Liu 0001 |
ACL (1) | 4 |
| 2025 | ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5abstractAutomatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children’s speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes. Jiaming Zhou 0001, Shiwan Zhao, Jiabei He 0001, Haoqin Sun, Hui Wang 0075, Aobo Kong, Xi Yang 0023, Yequan Wang, Yonghua Lin |
ACL (1) | 11 |
| 2025 | Capability Localization: Capabilities Can be Localized rather than Individual KnowledgeabstractLarge scale language models have achieved superior performance in tasks related to natural language processing, however, it is still unclear how model parameters affect performance improvement. Previous studies assumed that individual knowledge is stored in local parameters, and the storage form of individual knowledge is dispersed parameters, parameter layers, or parameter chains, which are not unified. We found through fidelity and reliability evaluation experiments that individual knowledge cannot be localized. Afterwards, we constructed a dataset for decoupling experiments and discovered the potential for localizing data commonalities. To further reveal this phenomenon, this paper proposes a **C**ommonality **N**euron **L**ocalization (**CNL**) method, which successfully locates commonality neurons and achieves a neuron overlap rate of 96.42% on the GSM8K dataset. Finally, we have demonstrated through cross data experiments that commonality neurons are a collection of capability neurons that possess the capability to enhance performance. Our code is available at https://github.com/nlpkeg/Capability-Neuron-Localization. Xiusheng Huang, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
ICLR | 3 |
| 2025 | Few-Shot Learner Generalizes Across AI-Generated Image DetectionabstractCurrent fake image detectors trained on large synthetic image datasets perform satisfactorily on limited studied generative models. However, these detectors suffer a notable performance decline over unseen models. Besides, collecting adequate training data from online generative models is often expensive or infeasible. To overcome these issues, we propose Few-Shot Detector (FSD), a novel AI-generated image detector which learns a specialized metric space for effectively distinguishing unseen fake images using very few samples. Experiments show that FSD achieves state-of-the-art performance by $+11.6\%$ average accuracy on the GenImage dataset with only $10$ additional samples. More importantly, our method is better capable of capturing the intra-category commonality in unseen images without further training. Our code is available at https://github.com/teheperinko541/Few-Shot-AIGI-Detector. Shiyu Wu, Yequan Wang |
ICML | 4 |
| 2025 | Not All Layers of LLMs Are Necessary During InferenceabstractDue to the large number of parameters, the inference phase of Large Language Models (LLMs) is resource-intensive. However, not all requests posed to LLMs are equally difficult to handle. Through analysis, we show that for some tasks, LLMs can achieve results comparable to the final output at some intermediate layers. That is, not all layers of LLMs are necessary during inference. If we can predict at which layer the inferred results match the final results (produced by evaluating all layers), we could significantly reduce the inference cost. To this end, we propose a simple yet effective algorithm named AdaInfer to adaptively terminate the inference process for an input instance. AdaInfer relies on easily obtainable statistical features and classic classifiers like SVM. Experiments on well-known LLMs like the Llama2 series and OPT, show that AdaInfer can achieve an average of 17.8% pruning ratio, and up to 43% on sentiment tasks, with nearly no performance drop (<1%). Because AdaInfer does not alter LLM parameters, the LLMs incorporated with AdaInfer maintain generalizability across tasks. Siqi Fan 0001, Xin Jiang 0005, Xiang Li 0001, Xuying Meng, Peng Han 0005, Shuo Shang, Aixin Sun, Yequan Wang |
IJCAI | 8 |
| 2025 | SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged SeniorsabstractWhile voice technologies increasingly serve aging populations, current systems exhibit significant performance gaps due to inadequate training data capturing elderly-specific vocal characteristics like presbyphonia and dialectal variations. The limited data available on super-aged individuals in existing elderly speech datasets, coupled with overly simple recording styles and annotation dimensions, exacerbates this issue. To address the critical scarcity of speech data from individuals aged 75 and above, we introduce SeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This dataset contains 55.53 hours of speech from 101 natural conversations involving 202 participants, ensuring a strategic balance across gender, region, and age. Through detailed annotation across multiple dimensions, it can support a wide range of speech tasks. We perform extensive experiments on speaker verification, speaker diarization, speech recognition, and speech editing tasks, offering crucial insights for the development of speech technologies targeting this age group. Code is available at https://github.com/flageval-baai/SeniorTalk and data at https://huggingface.co/datasets/evan0617/seniortalk. Yang Chen 0034, Hui Wang 0075, Jiabei He 0001, Jiaming Zhou 0001, Xi Yang 0023, Yequan Wang, Yonghua Lin |
NeurIPS | 8 |
| 2024 | Spectral-Based Graph Neural Networks for Complementary Item RecommendationabstractModeling complementary relationships greatly helps recommender systems to accurately and promptly recommend the subsequent items when one item is purchased. Unlike traditional similar relationships, items with complementary relationships may be purchased successively (such as iPhone and Airpods Pro), and they not only share relevance but also exhibit dissimilarity. Since the two attributes are opposites, modeling complementary relationships is challenging. Previous attempts to exploit these relationships have either ignored or oversimplified the dissimilarity attribute, resulting in ineffective modeling and an inability to balance the two attributes. Since Graph Neural Networks (GNNs) can capture the relevance and dissimilarity between nodes in the spectral domain, we can leverage spectral-based GNNs to effectively understand and model complementary relationships. In this study, we present a novel approach called Spectral-based Complementary Graph Neural Networks (SComGNN) that utilizes the spectral properties of complementary item graphs. We make the first observation that complementary relationships consist of low-frequency and mid-frequency components, corresponding to the relevance and dissimilarity attributes, respectively. Based on this spectral observation, we design spectral graph convolutional networks with low-pass and mid-pass filters to capture the low-frequency and mid-frequency components. Additionally, we propose a two-stage attention mechanism to adaptively integrate and balance the two attributes. Experimental results on four e-commerce datasets demonstrate the effectiveness of our model, with SComGNN significantly outperforming existing baseline models. Haitong Luo, Xuying Meng, Suhang Wang, Hanyun Cao, Weiyao Zhang, Yequan Wang, Yujun Zhang 0001 |
AAAI | 6 |
| 2024 | BiPFT: Binary Pre-trained Foundation Transformer with Low-Rank Estimation of Binarization Residual PolynomialsabstractPretrained foundation models offer substantial benefits for a wide range of downstream tasks, which can be one of the most potential techniques to access artificial general intelligence. However, scaling up foundation transformers for maximal task-agnostic knowledge has brought about computational challenges, especially on resource-limited devices such as mobiles. This work proposes the first Binary Pretrained Foundation Transformer (BiPFT) for natural language understanding (NLU) tasks, which remarkably saves 56 times operations and 28 times memory. In contrast to previous task-specific binary transformers, BiPFT exhibits a substantial enhancement in the learning capabilities of binary neural networks (BNNs), promoting BNNs into the era of pre-training. Benefiting from extensive pretraining data, we further propose a data-driven binarization method. Specifically, we first analyze the binarization error in self-attention operations and derive the polynomials of binarization error. To simulate full-precision self-attention, we define binarization error as binarization residual polynomials, and then introduce low-rank estimators to model these polynomials. Extensive experiments validate the effectiveness of BiPFTs, surpassing task-specific baseline by 15.4% average performance on the GLUE benchmark. BiPFT also demonstrates improved robustness to hyperparameter changes, improved optimization efficiency, and reduced reliance on downstream distillation, which consequently generalize on various NLU tasks and simplify the downstream pipeline of BNNs. Our code and pretrained models are publicly available at https://github.com/Xingrun-Xing/BiPFT. Xingrun Xing, Xianlin Zeng, Yequan Wang |
AAAI | 5 |
| 2024 | Multimodal Reasoning with Multimodal Knowledge GraphabstractMultimodal reasoning with large language models (LLMs) often suffers from hallucinations and the presence of deficient or outdated knowledge within LLMs.Some approaches have sought to mitigate these issues by employing textual knowledge graphs, but their singular modality of knowledge limits comprehensive cross-modal understanding.In this paper, we propose the Multimodal Reasoning with Multimodal Knowledge Graph (MR-MKG) method, which leverages multimodal knowledge graphs (MMKGs) to learn rich and semantic knowledge across modalities, significantly enhancing the multimodal reasoning capabilities of LLMs.In particular, a relation graph attention network is utilized for encoding MMKGs and a cross-modal alignment module is designed for optimizing image-text alignment.A MMKGgrounded dataset is constructed to equip LLMs with initial expertise in multimodal reasoning through pretraining.Remarkably, MR-MKG achieves superior performance while training on only a small fraction of parameters, approximately 2.25% of the LLM's parameter size.Experimental results on multimodal question answering and multimodal analogy reasoning tasks demonstrate that our MR-MKG method outperforms previous state-of-the-art models. Junlin Lee, Yequan Wang, Jing Li 0034, Min Zhang 0005 |
ACL (1) | 2 |
| 2024 | Commonsense Knowledge Editing Based on Free-Text in LLMsabstractKnowledge editing technology is crucial for maintaining the accuracy and timeliness of large language models (LLMs) .However, the setting of this task overlooks a significant portion of commonsense knowledge based on freetext in the real world, characterized by broad knowledge scope, long content and non instantiation.The editing objects of previous methods (e.g., MEMIT) were single token or entity, which were not suitable for commonsense knowledge in free-text form.To address the aforementioned challenges, we conducted experiments from two perspectives: knowledge localization and knowledge editing.Firstly, we introduced Knowledge Localization for Free-Text(KLFT) method, revealing the challenges associated with the distribution of commonsense knowledge in MLP and Attention layers, as well as in decentralized distribution.Next, we propose a Dynamics-aware Editing Method(DEM), which utilizes a Dynamicsaware Module to locate the parameter positions corresponding to commonsense knowledge, and uses Knowledge Editing Module to update knowledge.The DEM method fully explores the potential of the MLP and Attention layers, and successfully edits commonsense knowledge based on free-text.The experimental results indicate that the DEM can achieve excellent editing performance.The code and dataset file in URL 1 . Xiusheng Huang, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 2 |
| 2024 | Improving Zero-shot LLM Re-Ranker with Risk MinimizationabstractIn the Retrieval-Augmented Generation (RAG) system, advanced Large Language Models (LLMs) have emerged as effective Query Likelihood Models (QLMs) in an unsupervised way, which re-rank documents based on the probability of generating the query given the content of a document.However, directly prompting LLMs to approximate QLMs inherently is biased, where the estimated distribution might diverge from the actual document-specific distribution.In this study, we introduce a novel framework, UR 3 , which leverages Bayesian decision theory to both quantify and mitigate this estimation bias.Specifically, UR 3 reformulates the problem as maximizing the probability of document generation, thereby harmonizing the optimization of query and document generation probabilities under a unified risk minimization objective.Our empirical results indicate that UR 3 significantly enhances re-ranking, particularly in improving the Top-1 accuracy.It benefits the QA tasks by achieving higher accuracy with fewer input documents. Xiaowei Yuan, Zhao Yang 0004, Yequan Wang, Jun Zhao 0001, Kang Liu 0001 |
EMNLP | 3 |
| 2024 | Masked Structural Growth for 2x Faster Language Model Pre-trainingabstractAccelerating large language model pre-training is a critical issue in present research. In this paper, we focus on speeding up pre-training by progressively growing from a small Transformer structure to a large one. There are two main research problems associated with progressive growth: determining the optimal growth schedule, and designing efficient growth operators. In terms of growth schedule, the impact of each single dimension on a schedule’s efficiency is underexplored by existing work. Regarding the growth operators, existing methods rely on the initialization of new weights to inherit knowledge, and achieve only non-strict function preservation, limiting further improvements on training dynamics. To address these issues, we propose Masked Structural Growth (MSG), including (i) growth schedules involving all possible dimensions and (ii) strictly function-preserving growth operators that is independent of the initialization of new weights. Experiments show that MSG is significantly faster than related work: we achieve up to 2.2x speedup in pre-training different types of language models while maintaining comparable or better downstream performances. Code is publicly available at https://github.com/cofe-ai/MSG. Yiqun Yao, Zheng Zhang 0006, Jing Li 0034, Yequan Wang |
ICLR | 4 |
| 2024 | SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking MechanismsabstractTowards energy-efficient artificial intelligence similar to the human brain, the bio-inspired spiking neural networks (SNNs) have advantages of biological plausibility, event-driven sparsity, and binary activation. Recently, large-scale language models exhibit promising generalization capability, making it a valuable issue to explore more general spike-driven models. However, the binary spikes in existing SNNs fail to encode adequate semantic information, placing technological challenges for generalization. This work proposes the first fully spiking mechanism for general language tasks, including both discriminative and generative ones. Different from previous spikes with 0,1 levels, we propose a more general spike formulation with bi-directional, elastic amplitude, and elastic frequency encoding, while still maintaining the addition nature of SNNs. In a single time step, the spike is enhanced by direction and amplitude information; in spike frequency, a strategy to control spike firing rate is well designed. We plug this elastic bi-spiking mechanism in language modeling, named SpikeLM. It is the first time to handle general language tasks with fully spike-driven models, which achieve much higher accuracy than previously possible. SpikeLM also greatly bridges the performance gap between SNNs and ANNs in language modeling. Our code is available at https://github.com/Xingrun-Xing/SpikeLM. Xingrun Xing, Zheng Zhang 0006, Ziyi Ni, Shitao Xiao, Yiming Ju, Siqi Fan 0001, Yequan Wang |
ICML | 7 |
| 2024 | Reasons and Solutions for the Decline in Model Performance after EditingabstractKnowledge editing technology has received widespread attention for low-cost updates of incorrect or outdated knowledge in large-scale language models. However, recent research has found that edited models often exhibit varying degrees of performance degradation. The reasons behind this phenomenon and potential solutions have not yet been provided. In order to investigate the reasons for the performance decline of the edited model and optimize the editing method, this work explores the underlying reasons from both data and model perspectives. Specifically, 1) from a data perspective, to clarify the impact of data on the performance of editing models, this paper first constructs a **M**ulti-**Q**uestion **D**ataset (**MQD**) to evaluate the impact of different types of editing data on model performance. The performance of the editing model is mainly affected by the diversity of editing targets and sequence length, as determined through experiments. 2) From a model perspective, this article explores the factors that affect the performance of editing models. The results indicate a strong correlation between the L1-norm of the editing model layer and the editing accuracy, and clarify that this is an important factor leading to the bottleneck of editing performance. Finally, in order to improve the performance of the editing model, this paper further proposes a **D**ump **for** **S**equence (**D4S**) method, which successfully overcomes the previous editing bottleneck by reducing the L1-norm of the editing layer, allowing users to perform multiple effective edits and minimizing model damage. Our code is available at https://github.com/nlpkeg/D4S. Xiusheng Huang, Yequan Wang, Kang Liu 0001 |
NeurIPS | 3 |
| 2024 | Contrastive Language-knowledge Graph Pre-trainingabstractRecent years have witnessed a surge of academic interest in knowledge-enhanced pre-trained language models (PLMs) that incorporate factual knowledge to enhance knowledge-driven applications. Nevertheless, existing studies primarily focus on shallow, static, and separately pre-trained entity embeddings, with few delving into the potential of deep contextualized knowledge representation for knowledge incorporation. Consequently, the performance gains of such models remain limited. In this article, we introduce a simple yet effective knowledge-enhanced model, College ( Co ntrastive L anguage-Know le dge G raph Pr e -training), which leverages contrastive learning to incorporate factual knowledge into PLMs. This approach maintains the knowledge in its original graph structure to provide the most available information and circumvents the issue of heterogeneous embedding fusion. Experimental results demonstrate that our approach achieves more effective results on several knowledge-intensive tasks compared to previous state-of-the-art methods. Our code and trained models are available at https://github.com/Stacy027/COLLEGE . Xiaowei Yuan, Kang Liu 0001, Yequan Wang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2023 | Knowledgeable Parameter Efficient Tuning Network for Commonsense Question AnsweringabstractCommonsense question answering is important for making decisions about everyday matters.Although existing commonsense question answering works based on fully fine-tuned PLMs have achieved promising results, they suffer from prohibitive computation costs as well as poor interpretability.Some works improve the PLMs by incorporating knowledge to provide certain evidence, via elaborately designed GNN modules which require expertise.In this paper, we propose a simple knowledgeable parameter efficient tuning network to couple PLMs with external knowledge for commonsense question answering.Specifically, we design a trainable parameter-sharing adapter attached to a parameter-freezing PLM to incorporate knowledge at a small cost.The adapter is equipped with both entity-and query-related knowledge via two auxiliary knowledge-related tasks (i.e., span masking and relation discrimination).To make the adapter focus on the relevant knowledge, we design gating and attention mechanisms to respectively filter and fuse the query information from the PLM.Extensive experiments on two benchmark datasets show that KPE is parameter-efficient and can effectively incorporate knowledge for improving commonsense question answering. Ziwang Zhao, Linmei Hu, Yingxia Shao, Yequan Wang |
ACL (1) | 5 |
| 2023 | Finding Generalization Measures by Contrasting Signal and NoiseabstractGeneralization is one of the most fundamental challenges in deep learning, aiming to predict model performances on unseen data. Empirically, such predictions usually rely on a validation set, while recent works showed that an unlabeled validation set also works. Without validation sets, it is extremely difficult to obtain non-vacuous generalization bounds, which leads to a weaker task of finding generalization measures that monotonically relate to generalization error. In this paper, we propose a new generalization measure REF Complexity (RElative Fitting degree between signal and noise), motivated by the intuition that a given model-algorithm pair may generalize well if it fits signal (e.g., true labels) fast while fitting noise (e.g., random labels) slowly. Empirically, REF Complexity monotonically relates to test accuracy in real-world datasets without accessing additional validation sets, achieving -0.988 correlation on CIFAR-10 and -0.960 correlation on CIFAR-100. We further theoretically verify the utility of REF Complexity under three different cases, including convex and smooth regimes with stochastic gradient descent, smooth regimes (not necessarily convex) with stochastic gradient Langevin dynamics, and linear regimes with gradient descent. The code is available at https://github.com/962086838/REF-complexity. Jiaye Teng, Bohang Zhang, Haowei He, Yequan Wang, Yang Yuan 0010 |
ICML | 5 |
| 2023 | Multi-Layer Collaborative Bandit for Multivariate Time Series Anomaly DetectionabstractMultivariate Time Series Anomaly Detection (MTSAD) detects abnormal indicators from Multivariate Time Series (MTS), and provides the rank of the multiple abnormal indicators to meet the expert's detection interest in current environment, which underpin the security and stability of intelligent cyber-physical systems. However, popular integration-based methods, which are pre-defined, fall short in locating the exact abnormal indicator, nor can they perceive the environmental dynamic and evolve accordingly. Let alone meeting the expert's interest. As a result, the expert's workload is exaggerated. These issues motivate us to propose a novel multi-layer collaborative bandit framework MULA for MTSAD. MULA decomposes MTS and pairs individual time series with a bandit arm, which locates the abnormal indicator directly. Then, MULA sorts the indicators by abnormal scores computed based on the expert's feedback, which facilitates the experts. Besides, to address the adaptability issue, we devise a dual signal to comprehensively monitor environmental changes, and design a multi-layer collaborative mechanism for MULA to adapt to the dynamic environment. Theoretical analysis and experiments on public datasets demonstrate the superiority of MULA compared to the state-of-art. Weiyao Zhang, Xuying Meng, Jinyang Li 0009, Yequan Wang, Yujun Zhang 0001 |
IWQoS | 4 |
| 2023 | EmpMFF: A Multi-factor Sequence Fusion Framework for Empathetic Response GenerationabstractEmpathy is one of the fundamental abilities of dialog systems. In order to build more intelligent dialogue systems, it’s important to learn how to demonstrate empathy toward others. Existing studies focus on identifying and leveraging the user’s coarse emotion to generate empathetic responses. However, human emotion and dialog act (e.g., intent) evolve as the talk goes along in an empathetic dialogue. This leads to the generated responses with very different intents from the human responses. As a result, empathy failure is ultimately caused. Therefore, using fine-grained emotion and intent sequential data on conversational emotions and dialog act is crucial for empathetic response generation. On the other hand, existing empathy models overvalue the empathy of responses while ignoring contextual relevance, which results in repetitive model-generated responses. To address these issues, we propose a Multi-Factor sequence Fusion framework (EmpMFF) based on conditional variational autoencoder. To generate empathetic responses, the proposed EmpMFF encodes a combination of contextual, emotion, and intent information into a continuous latent variable, which is then fed into the decoder. Experiments on the EmpatheticDialogues benchmark dataset demonstrate that EmpMFF exhibits exceptional performance in both automatic and human evaluations. Xiaobing Pang, Yequan Wang, Siqi Fan 0001, Lisi Chen 0001, Shuo Shang, Peng Han 0005 |
WWW | 2 |
| 2023 | EDU-Capsule: aspect-based sentiment analysis at clause level
Ting Lin, Aixin Sun, Yequan Wang |
Knowl. Inf. Syst. | 3 |
| 2023 | UaMC: user-augmented conversation recommendation via multi-modal graph learning and context mining
Siqi Fan 0001, Yequan Wang, Xiaobing Pang, Lisi Chen 0001, Peng Han 0005, Shuo Shang |
World Wide Web (WWW) | 2 |
| 2022 | CofeNet: Context and Former-Label Enhanced Net for Complicated Quotation ExtractionabstractQuotation extraction aims to extract quotations from written text. There are three components in a quotation: source refers to the holder of the quotation, cue is the trigger word(s), and content is the main body. Existing solutions for quotation extraction mainly utilize rule-based approaches and sequence labeling models. While rule-based approaches often lead to low recalls, sequence labeling models cannot well handle quotations with complicated structures. In this paper, we propose the Context and Former-Label Enhanced Net () for quotation extraction. is able to extract complicated quotations with components of variable lengths and complicated structures. On two public datasets (and ) and one proprietary dataset (), we show that our achieves state-of-the-art performance on complicated quotation extraction. Yequan Wang, Xiang Li 0001, Aixin Sun, Xuying Meng, Huaming Liao, Jiafeng Guo |
COLING | 1 |
| 2022 | Interactive Information Extraction by Semantic Information GraphabstractInformation extraction (IE) mainly focuses on three highly correlated subtasks, i.e., entity extraction, relation extraction and event extraction. Recently, there are studies using Abstract Meaning Representation (AMR) to utilize the intrinsic correlations among these three subtasks. AMR based models are capable of building the relationship of arguments. However, they are hard to deal with relations. In addition, the noises of AMR (i.e., tags unrelated to IE tasks, nodes with unconcerned conception, and edge types with complicated hierarchical structures) disturb the decoding processing of IE. As a result, the decoding processing limited by the AMR cannot be worked effectively. To overcome the shortages, we propose an Interactive Information Extraction (InterIE) model based on a novel Semantic Information Graph (SIG). SIG can guide our InterIE model to tackle the three subtasks jointly. Furthermore, the well-designed SIG without noise is capable of enriching entity and event trigger representation, and capturing the edge connection between the information types. Experimental results show that our InterIE achieves state-of-the-art performance on all IE subtasks on the benchmark dataset (i.e., ACE05-E+ and ACE05-E). More importantly, the proposed model is not sensitive to the decoding order, which goes beyond the limitations of AMR based methods. Siqi Fan 0001, Yequan Wang, Jing Li 0034, Zheng Zhang 0006, Shuo Shang, Peng Han 0005 |
IJCAI | 2 |
| 2022 | Packet Representation Learning for Traffic ClassificationabstractWith the surging development of information technology, to provide a high quality of network services, there are increasing demands and challenges for network analysis. As all data on the Internet are encapsulated and transferred by network packets, packets are widely used for various network traffic analysis tasks, from application identification to intrusion detection. Considering the choice of features and how to represent them can greatly affect the performance of downstream tasks, it is critical to learn high-quality packet representations. In addition, existing packet-level works ignore packet representations but focus on trying to get good performance with independent analysis of different classification tasks. In the real world, although a packet may have different class labels for different tasks, the packet representation learned from one task can also help understand its complex packet patterns in other tasks, while existing works omit to leverage them. Xuying Meng, Yequan Wang, Runxin Ma, Haitong Luo, Yujun Zhang 0001 |
KDD | 2 |
| 2022 | Aspect-Based Sentiment Analysis Through EDU-Level Attentions
Ting Lin, Aixin Sun, Yequan Wang |
PAKDD (1) | 3 |
| 2021 | Interactive Anomaly Detection in Dynamic Communication NetworksabstractNetwork flows are the basic components of the Internet. Considering the serious consequences of abnormal flows, it is crucial to provide timely anomaly detection in dynamic communication networks. To obtain accurate anomaly detection results in dynamic networks, supervision from experts is highly demanded. However, to obtain high-quality ground truth of abnormal flows, we suffer from two major problems: (1)limited labor resources: experts with the latest domain knowledge are much fewer than the large number of flows; and (2)dynamic environment: considering the new abnormal patterns (i.e., new attacks) and continuously changing network structures, it requires timely supervision to adaptively update the parameters. To tackle these problems, we propose HADDN, a novel bandit framework for periodic-updated anomaly detection in dynamic communication networks. We formulate the task as a bandit problem, where by interactions, supervision is offered by human experts to provide the ground truth to a fraction of flows. We construct semi-parametric expected rewards to optimize the estimation of flows’ abnormality in limited interactions. Also, we utilize feature-based clusters and structural correlations to make connections between historical flows and new flows to improve both efficiency and accuracy of abnormality estimation. What’s more, we provide two implementations for the semi-parametric expected reward of the proposed HADDN with theoretical proof. Experimental evaluations on public datasets demonstrate the substantial improvement of our proposed approaches compared to state-of-art anomaly detection methods. Xuying Meng, Yequan Wang, Suhang Wang, Di Yao 0001, Yujun Zhang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2019 | Path Travel Time Estimation using Attribute-related Hybrid Trajectories NetworkabstractEstimation of path travel time provides great value to applications like bus line designs and route plannings. Existing approaches are mainly based on single-source trajectory datasets that are usually large in size to ensure a satisfactory performance. This leads to two limitations: 1) Large-scale data may not always be attainable, e.g. city-scale public bus data is usually small compared to taxi data due to relative fewer bus trips in a day. 2) Considering only single-source trajectory data neglects the potential estimation-improving insights of external data, e.g. trajectory dataset of other vehicle sources obtained from the same geographical region. A challenge is how to effectively utilize such other trajectory sources. Moreover, existing work does not attend the important attributes of a trajectory including vehicle ID, day of week, rainfall level etc., which are important for estimating the path travel time. Motivated by these and the recent successes of neural network models, we propose Attribute-related Hybrid Trajectories Network~(AtHy-TNet), a neural model that effectively utilizes the attribute correlations, as well as the spatial and temporal relationships across hybrid trajectory data. We apply this to a novel problem of estimating path travel time of a type of vehicles using a hybrid trajectory dataset that includes trajectories from other vehicle types. We demonstrate in our experiments the benefits of considering hybrid data for travel time estimation, and show that AtHy-TNet significantly outperforms state-of-the-art methods on real-world trajectory datasets. Xi Lin 0007, Yequan Wang, Xiaokui Xiao, Zengxiang Li, Sourav S. Bhowmick |
CIKM | 2 |
| 2019 | Aspect-level Sentiment Analysis using AS-CapsulesabstractAspect-level sentiment analysis aims to provide complete and detailed view of sentiment analysis from different aspects. Existing solutions usually adopt a two-staged approach: first detecting aspect category in a document, then categorizing the polarity of opinion expressions for detected aspect(s). Inevitably, such methods lead to error accumulation. Moreover, aspect detection and aspect-level sentiment classification are highly correlated with each other. The key issue here is how to perform aspect detection and aspect-level sentiment classification jointly, and effectively. In this paper, we propose the aspect-level sentiment capsules model (AS-Capsules), which is capable of performing aspect detection and sentiment classification simultaneously, in a joint manner. AS-Capsules utilizes the correlation between aspect and sentiment through shared components including capsule embedding, shared encoders, and shared attentions. AS-Capsules is also capable of communicating with different capsules through a shared Recurrent Neural Network (RNN). More importantly, AS-Capsules model does not require any linguistic knowledge as additional input. Instead, through the attention mechanism, this model is able to attend aspect related words and sentiment words corresponding to different aspect(s). Experiments show that the AS-Capsules model achieves state-of-the-art performances on a benchmark dataset for aspect-level sentiment analysis. Yequan Wang, Aixin Sun, Minlie Huang, Xiaoyan Zhu 0001 |
WWW | 1 |
| 2018 | Sentiment Analysis by CapsulesabstractIn this paper, we propose RNN-Capsule, a capsule model based on Recurrent Neural Network (RNN) for sentiment analysis. For a given problem, one capsule is built for each sentiment category e.g., 'positive' and 'negative'. Each capsule has an attribute, a state, and three modules: representation module, probability module, and reconstruction module. The attribute of a capsule is the assigned sentiment category. Given an instance encoded in hidden vectors by a typical RNN, the representation module builds capsule representation by the attention mechanism. Based on capsule representation, the probability module computes the capsule's state probability. A capsule's state is active if its state probability is the largest among all capsules for the given instance, and inactive otherwise. On two benchmark datasets (i.e., Movie Review and Stanford Sentiment Treebank) and one proprietary dataset (i.e., Hospital Feedback), we show that RNN-Capsule achieves state-of-the-art performance on sentiment classification. More importantly, without using any linguistic knowledge, RNN-Capsule is capable of outputting words with sentiment tendencies reflecting capsules' attributes. The words well reflect the domain specificity of the dataset. Yequan Wang, Aixin Sun, Jialong Han, Ying Liu 0004, Xiaoyan Zhu 0001 |
WWW | 1 |
| 2016 | Attention-based LSTM for Aspect-level Sentiment ClassificationabstractAspect-level sentiment classification is a finegrained task in sentiment analysis.Since it provides more complete and in-depth results, aspect-level sentiment analysis has received much attention these years.In this paper, we reveal that the sentiment polarity of a sentence is not only determined by the content but is also highly related to the concerned aspect.For instance, "The appetizers are ok, but the service is slow.",for aspect taste, the polarity is positive while for service, the polarity is negative.Therefore, it is worthwhile to explore the connection between an aspect and the content of a sentence.To this end, we propose an Attention-based Long Short-Term Memory Network for aspect-level sentiment classification.The attention mechanism can concentrate on different parts of a sentence when different aspects are taken as input.We experiment on the SemEval 2014 dataset and results show that our model achieves state-ofthe-art performance on aspect-level sentiment classification. Yequan Wang, Minlie Huang, Xiaoyan Zhu 0001, Li Zhao 0007 |
EMNLP | 1 |