VLDB 2026 Research / reviewers in the wild / expert
Bin Ji 0002
dblp:119/1943-2
· DBLP profile ↗
33ranked-venue papers
9as first author
31since 2021 · last 2026
0000-0002-5508-5051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 18 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Leveraging Image as Compressed Visual Prompt and Hierarchical Visual Knowledge for Effective Image Utilization in MLLMsabstractMultimodal Large Language Models (MLLMs) integrate text and images for complex reasoning tasks, but efficiently utilizing image remains a challenge due to redundancy and noise. Traditional methods take the entire image features as visual prompt into the MLLMs, leading to excessive visual tokens that disrupt textual information expression. Thus, recent studies treat image features as visual knowledge, storing them in the feed-forward network for retrieval when needed. These methods, completely removing images from the input, may hinder the activation of image-related knowledge. Besides, current visual knowledge focuses on fine-grained details but overlooks the hierarchical process of visual perception. As described in feature integration theory, global structure is first processed before details are integrated. Ignoring this process may lead to a fragmented visual understanding, making it difficult to capture high-level semantic relationships. To overcome these issues, we propose a novel image utilization mechanism in MLLMs. We leverage a compression-based attention mechanism to generate the compressed visual prompt, which not only mitigates the interference of excessively long visual prompts but also preserves crucial visual information necessary for activating knowledge in the MLLM. Furthermore, we extract hierarchical visual features as visual knowledge using wavelet transforms, allowing the model to capture both global structures and fine-grained details. Experiments show that our method achieves state-of-the-art performance. Shezheng Song, Kangcheng Ding, Shan Zhao 0002, Shasha Li 0001, Xiaopeng Li 0006, Chengyu Wang 0008, Qian Wan 0007, Bin Ji 0002, Jie Yu 0008 |
AAAI | 8 |
| 2026 | SPAR: Step-wise Path Dispatching and Asymmetric Re-routing for Efficient MoE Inference
Qingxiao Zhang, Xiaopeng Li 0006, Jinzhu Kong, Xiaodong Liu 0004, Bin Ji 0002, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008 |
ICIC (26) | 5 |
| 2026 | EMSEdit: Efficient Multi-Step Meta-Learning-based Model EditingabstractLarge Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted parameter modifications, with meta-learning-based model editing (MLME) demonstrating strong effectiveness and efficiency. However, we find that MLME struggles in low-data regimes and incurs high training costs due to the use of KL divergence. To address these issues, we propose $\textbf{E}$fficient $\textbf{M}$ulti-$\textbf{S}$tep $\textbf{Edit (EMSEdit)}$, which leverages multi-step backpropagation (MSBP) to effectively capture gradient-activation mapping patterns within editing samples, performs multi-step edits per sample to enhance editing performance under limited data, and introduces norm-based regularization to preserve unedited knowledge while improving training efficiency. Experiments on two datasets and three LLMs show that EMSEdit consistently outperforms state-of-the-art methods in both sequential and batch editing. Moreover, MSBP can be seamlessly integrated into existing approaches to yield additional performance gains. Further experiments on a multi-hop reasoning editing task demonstrate EMSEdit's robustness in handling complex edits, while ablation studies validate the contribution of each design component. Our code is available at https://github.com/xpq-tech/emsedit. Xiaopeng Li 0006, Shasha Li 0001, Xi Wang 0018, Shezheng Song, Bin Ji 0002, Shangwen Wang, Jun Ma 0015, Xiaodong Liu 0004, Mina Liu, Jie Yu 0008 |
WWW | 5 |
| 2026 | Emp: enhance memory in data pruning
Jinying Xiao, Ping Li 0034, Jie Nie, Bin Ji 0002, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Qingbo Wu 0003, Jie Yu 0008 |
Data Min. Knowl. Discov. | 4 |
| 2026 | SEAttack: A self-evolving jailbreak attack to induce toxic responses for non-toxic queries in large language models
Huijun Liu 0003, Shasha Li 0001, Bin Ji 0002, Xiaohu Du, Xiaopeng Li 0006, Jun Ma 0015, Jie Yu 0008 |
Inf. Process. Manag. | 3 |
| 2026 | Evaluating Large Language Models on Named Entity RecognitionabstractLarge language models (LLMs) are popping up all over the place, and they have been gaining prominence due to their exceptional abilities in conducting various tasks. Although extensive LLM evaluation has been explored on natural language understanding tasks like text classification and sentiment analysis, evaluating LLMs on named entity recognition (NER) still remains under-explored. To fill this gap, we evaluate twenty-eight representative LLMs on thirteen datasets across five domains, whose parameters range from 3 billion to 175 billion, from four perspectives, that is, supervised fine-tuning (SFT), parameter scales, hallucinations, and prompt designs. We propose an LLM-based NER framework (LLM-NER) for the evaluation, which consists of a Recognition phase and a Check phase. Specifically, the Check guides LLMs to examine the correctness of recognized entities, which is designed to mitigate hallucinations in the NER scenario. Qualitative and quantitative evaluation analyses demonstrate that in the NER scenario: 1) SFT empowers LLMs to understand and follow human instructions; 2) LLMs' ability generally improves as their parameter scales consistently increase; 3) hallucinations exist in all evaluated LLMs, and guiding LLMs to check their outputs is a feasible way to alleviate hallucinations; and 4) all evaluated LLMs are sensitive to prompt designs. Based on the analyses, we highlight a number of promising directions for future study. Moreover, our evaluation shows high consistency with two LLM evaluation leaderboards, which evaluate LLMs on other tasks, demonstrating the rationality of our evaluation design. Bin Ji 0002, Huijun Liu 0003, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, See-Kiong Ng |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Towards Verifiable Text Generation with Generative AgentabstractText generation with citations makes it easy to verify the factuality of Large Language Models’ (LLMs) generations. Existing one-step generation studies expose distinct shortages in answer refinement and in-context demonstration matching. In light of these challenges, we propose R2-MGA, a Retrieval and Reflection Memory-augmented Generative Agent. Specifically, it first retrieves the memory bank to obtain the best-matched memory snippet, then reflects the retrieved snippet as a reasoning rationale, next combines the snippet and the rationale as the best-matched in-context demonstration. Additionally, it is capable of in-depth answer refinement with two specifically designed modules. We evaluate R2-MGA across five LLMs on the ALCE benchmark. The results reveal R2-MGA’ exceptional capabilities in text generation with citations. In particular, compared to the selected baselines, it delivers up to +58.8% and +154.7% relative performance gains on answer correctness and citation quality, respectively. Extensive analyses strongly support the motivations of R2-MGA. Bin Ji 0002, Huijun Liu 0003, Mingzhe Du, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Jie Yu 0008, See-Kiong Ng |
AAAI | 1 |
| 2025 | SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringabstractThe general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability. Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Huijun Liu 0003, Bin Ji 0002, Xi Wang 0018, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004 |
AAAI | 5 |
| 2025 | Cross-Modal Reasoning-Based Unsupervised Multi-modal Entity Linking
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008 |
DASFAA (3) | 4 |
| 2025 | Multi-modal Entity Linking Model Based on Knowledge Distillation
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008 |
ICIC (24) | 4 |
| 2025 | Model Editing for LLMs4Code: How Far are we?abstractLarge Language Models for Code (LLMs4Code) have been found to exhibit outstanding performance in the software engineering domain, especially the remarkable performance in coding tasks. However, even the most advanced LLMs4Code can inevitably contain incorrect or outdated code knowledge. Due to the high cost of training LLMs4Code, it is impractical to re-train the models for fixing these problematic code knowledge. Model editing is a new technical field for effectively and efficiently correcting erroneous knowledge in LLMs, where various model editing techniques and benchmarks have been proposed recently. Despite that, a comprehensive study that thoroughly compares and analyzes the performance of the state-of-the-art model editing techniques for adapting the knowledge within LLMs4Code across various code-related tasks is notably absent. To bridge this gap, we perform the first systematic study on applying state-of-the-art model editing approaches to repair the inaccuracy of LLMs4Code. To that end, we introduce a benchmark named CLMEEval, which consists of two datasets, i.e., CoNaLa-Edit (CNLE) with 21K+ code generation samples and CodeSearchNet-Edit (CSNE) with 16K+ code summarization samples. With the help of CLMEEval, we evaluate six advanced model editing techniques on three LLMs4Code: CodeLlama (7B), CodeQwen1.5 (7B), and Stable-Code (3B). Our findings include that the external memorization-based GRACE approach achieves the best knowledge editing effectiveness and specificity (the editing does not influence untargeted knowledge), while generalization (whether the editing can generalize to other semantically-identical inputs) is a universal challenge for existing techniques. Furthermore, building on in-depth case analysis, we introduce an enhanced version of GRACE called A-GRACE, which incorporates contrastive learning to better capture the semantics of the inputs. Results demonstrate that A-GRACE notably enhances generalization while maintaining similar levels of effectiveness and specificity compared to the vanilla GRACE. Xiaopeng Li 0006, Shangwen Wang, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004, Bin Ji 0002 |
ICSE | 8 |
| 2025 | LSAQ: Layer-Specific Adaptive Quantization for Large Language Model DeploymentabstractAs Large Language Models (LLMs) demonstrate exceptional performance across various domains, deploying LLMs on edge devices has emerged as a new trend. Quantization techniques, which reduce the size and memory requirements of LLMs, are effective for deploying LLMs on resource-limited edge devices. However, existing one-size-fits-all quantization methods often fail to dynamically adjust the memory requirements of LLMs, limiting their applications to practical edge devices with various computation resources. To tackle this issue, we propose Layer-Specific Adaptive Quantization (LSAQ), a system for adaptive quantization and dynamic deployment of LLMs based on layer importance. Specifically, LSAQ evaluates the importance of LLMs’ neural layers by constructing top-k token sets from the inputs and outputs of each layer and calculating their Jaccard similarity. Based on layer importance, our system adaptively adjusts quantization strategies in real time according to the computation resource of edge devices, which applies higher quantization precision to layers with higher importance, and vice versa. Experimental results show that LSAQ consistently outperforms the selected quantization baselines in terms of perplexity and zero-shot tasks. Additionally, it can devise appropriate quantization schemes for different usage scenarios to facilitate the deployment of LLMs. Binrui Zeng, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Xiaopeng Li 0006, Shangwen Wang, Xinran Hong, Yongtao Tang |
IJCNN | 2 |
| 2025 | Identifying Knowledge Editing Types in Large Language ModelsabstractWarning: This paper contains examples of toxic text. Knowledge editing has emerged as an efficient technique for updating the knowledge of large language models (LLMs), attracting increasing attention in recent years. However, there is a lack of effective measures to prevent the malicious misuse of this technique, which could lead to harmful edits in LLMs. These malicious modifications could cause LLMs to generate toxic content, misleading users into inappropriate actions. In front of this risk, we introduce a new task, Knowledge Editing Type Identification (KETI), aimed at identifying different types of edits in LLMs, thereby providing timely alerts to users when encountering illicit edits. As part of this task, we propose KETIBench, which includes five types of harmful edits covering the most popular toxic types, as well as one benign factual edit. We develop five classical classification models and three BERT-based models as baseline identifiers for both open-source and closed-source LLMs. Our experimental results, across 92 trials involving four models and three knowledge editing methods, demonstrate that all eight baseline identifiers achieve decent identification performance, highlighting the feasibility of identifying malicious edits in LLMs. Additional analyses reveal that the performance of the identifiers is independent of the reliability of the knowledge editing methods and exhibits cross-domain generalization, enabling the identification of edits from unknown sources. All data and code are available in https://github.com/xpq-tech/KETI. Xiaopeng Li 0006, Shasha Li 0001, Shangwen Wang, Shezheng Song, Bin Ji 0002, Huijun Liu 0003, Jun Ma 0015, Jie Yu 0008 |
KDD (2) | 5 |
| 2025 | Rethinking Residual Distribution in Locate-then-Edit Model EditingabstractModel editing enables targeted updates to the knowledge of large language models (LLMs) with minimal retraining. Among existing approaches, locate-then-edit methods constitute a prominent paradigm: they first identify critical layers, then compute residuals at the final critical layer based on the target edit, and finally apply least-squares-based multi-layer updates via $\textbf{residual distribution}$. While empirically effective, we identify a counterintuitive failure mode: residual distribution, a core mechanism in these methods, introduces weight shift errors that undermine editing precision. Through theoretical and empirical analysis, we show that such errors increase with the distribution distance, batch size, and edit sequence length, ultimately leading to inaccurate or suboptimal edits. To address this, we propose the $\textbf{B}$oundary $\textbf{L}$ayer $\textbf{U}$pdat$\textbf{E (BLUE)}$ strategy to enhance locate-then-edit methods. Sequential batch editing experiments on three LLMs and two datasets demonstrate that BLUE not only delivers an average performance improvement of 35.59\%, significantly advancing the state of the art in model editing, but also enhances the preservation of LLMs' general capabilities. Our code is available at https://github.com/xpq-tech/BLUE. Xiaopeng Li 0006, Shangwen Wang, Shasha Li 0001, Shezheng Song, Bin Ji 0002, Ma Jun, Jie Yu 0008 |
NeurIPS | 5 |
| 2025 | A Survey of AI Inference Technologies for On-Device SystemsabstractIn recent years, artificial intelligence(AI) technologies represented by foundation models have experienced rapid development. Concurrently, On-device AI inference has become the primary approach for intelligent technology applications, offering advantages such as low latency, high security, and personalization. However, due to the limited resources of on-device systems, on-device AI inference faces new challenges, including improving computational efficiency, optimizing task parallelism, and model optimization. This survey addresses these challenges from a software and algorithmic perspective, focusing on three key areas: Operator Computation: Explores methods to accelerate matrix multiplication and convolution, as well as techniques like operator fusion and vectorized computation. Task Inference: Analyzes heterogeneous and distributed computing, memory allocation, and energy-efficient tuning to improve the parallel execution and energy efficiency of inference tasks. AI Models: Covers model compression, lookup table quantization, and model architecture design to reduce computational complexity and storage requirements. By analyzing these areas, the survey aims to improve inference speed, reduce resource dependency, and provide insights into the future trends of on-device AI technology. Wenzhu Wang, Ke Li 0026, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Qingbo Wu 0003 |
IEEE Internet Things J. | 3 |
| 2025 | Win-Win Cooperation: Bundling Sequence and Span Models for Named Entity RecognitionabstractFor Named Entity Recognition (NER), sequence labeling-based and span-based paradigms are quite different. Previous studies have demonstrated the clear complementary advantages of the two paradigms, but few models have tried to incorporate them into a single NER model as far as we know. In our previous work, we proposed a paradigm called Bundling Learning (BL) to explore the above issue, which bundles the two NER paradigms, enabling NER models to jointly tune their parameters by weighted summing each paradigm's training loss. However, three critical issues remain unresolved: When does BL work? Why does BL work? Can BL enhance existing state-of-the-art NER models? To address the first two issues, we design three NER models: a sequence labeling-based model – SeqNER, a span-based NER model – SpanNER, and BL-NER which bundles SeqNER and SpanNER. We draw two conclusions regarding the two issues based on the experimental results on eleven NER datasets. To investigate the third issue, we apply BL to five existing state-of-the-art NER models, including three sequence labeling-based and two span-based models. Experimental results indicate consistent NER performance gains, suggesting a feasible way to construct new state-of-the-art NER systems by applying BL to the current state-of-the-art systems. Moreover, investigation results show that BL reduces both entity boundary and type prediction errors. In addition, we compare two commonly used label tagging methods and three types of span semantic representations. Bin Ji 0002, Huijun Liu 0003, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | From Static to Dynamic: Knowledge Metabolism for Large Language ModelsabstractThe immense parameter space of Large Language Models (LLMs) endows them with superior knowledge retention capabilities, allowing them to excel in a variety of natural language processing tasks. However, it also instigates difficulties in consistently tuning LMs to incorporate the most recent knowledge, which may further lead LMs to produce inaccurate and fabricated content. To alleviate this issue, we propose a knowledge metabolism framework for LLMs. This framework proactively sustains the credibility of knowledge through an auxiliary external memory component and directly delivers pertinent knowledge for LM inference, thereby suppressing hallucinations caused by obsolete internal knowledge during the LM inference process. Benchmark experiments demonstrate DynaMind's effectiveness in overcoming this challenge. The code and demo of DynaMind are available at: https://github.com/Elfsong/DynaMind. Mingzhe Du, Anh Tuan Luu, Bin Ji 0002, See-Kiong Ng |
AAAI | 3 |
| 2024 | Chain-of-Thought Improves Text Generation with Citations in Large Language ModelsabstractPrevious studies disclose that Large Language Models (LLMs) suffer from hallucinations when generating texts, bringing a novel and challenging research topic to the public, which centers on enabling LLMs to generate texts with citations. Existing work exposes two limitations when using LLMs to generate answers to questions with provided documents: unsatisfactory answer correctness and poor citation quality. To tackle the above issues, we investigate using Chain-of-Thought (CoT) to elicit LLMs’ ability to synthesize correct answers from multiple documents, as well as properly cite these documents. Moreover, we propose a Citation Insurance Mechanism, which enables LLMs to detect and cite those missing citations. We conduct experiments on the ALCE benchmark with six open-source LLMs. Experimental results demonstrate that: (1) the CoT prompting strategy significantly improves the quality of text generation with citations; (2) the Citation Insurance Mechanism delivers impressive gains in citation quality at a low cost; (3) our best approach performs comparably as previous best ChatGPT-based baselines. Extensive analyses further validate the effectiveness of the proposed approach. Bin Ji 0002, Huijun Liu 0003, Mingzhe Du, See-Kiong Ng |
AAAI | 1 |
| 2024 | Offline Textual Adversarial Attacks against Large Language ModelsabstractThis work centers on textual adversarial attacks against large language models (LLMs) and proposes a new reproducible benchmark for future study. Unlike pre-trained language models (PLMs) which can output predicted class probabilities as feedback to instruct the generation of adversarial examples, LLMs cannot accurately provide such feedback due to their generative nature, making existing attack modes unsuitable. To address this issue, we propose Offline-Attack, an offline method tailored for LLMs that contains a novel Transformer-based Adversarial Machine Translation (AMT) framework. AMT is trained on one self-constructed large-scale adversarial dataset and used to translate original texts to adversarial examples. To mitigate training bias, we induce LLMs to generate stable prediction confidence and incorporate it into AMT training process. The evaluation, spanning four text classification datasets against LLaMA-2-13b-chat, showcases Offline-Attack’s robust performance, particularly achieving 44.3% attack success rate on average. Moreover, Offline-Attack exhibits promising attack ability to other LLMs like Vicuna-33b and ChatGPT. Our study paves the way for future study by presenting strong and reproducible baselines for textual adversarial attacks against LLMs. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Miaomiao Li 0001, Xi Wang 0018 |
IJCNN | 2 |
| 2024 | Mercury: A Code Efficiency Benchmark for Code Large Language ModelsabstractAmidst the recent strides in evaluating Large Language Models for Code (Code LLMs), existing benchmarks have mainly focused on the functional correctness of generated code, neglecting the importance of their computational efficiency. To fill the gap, we present Mercury, the first code efficiency benchmark for Code LLMs. It comprises 1,889 Python tasks, each accompanied by adequate solutions that serve as real-world efficiency baselines, enabling a comprehensive analysis of the runtime distribution. Based on the distribution, we introduce a new metric Beyond, which computes a runtime-percentile-weighted Pass score to reflect functional correctness and code efficiency simultaneously. On Mercury, leading Code LLMs can achieve 65% on Pass, while less than 50% on Beyond. Given that an ideal Beyond score would be aligned with the Pass score, it indicates that while Code LLMs exhibit impressive capabilities in generating functionally correct code, there remains a notable gap in their efficiency. Finally, our empirical experiments reveal that Direct Preference Optimization (DPO) serves as a robust baseline for enhancing code efficiency compared with Supervised Fine Tuning (SFT), which paves a promising avenue for future exploration of efficient code generation. Our code and data are available on GitHub: https://github.com/Elfsong/Mercury. Mingzhe Du, Anh Tuan Luu, Bin Ji 0002, Qian Liu 0033, See-Kiong Ng |
NeurIPS | 3 |
| 2024 | Span-based joint entity and relation extraction augmented with sequence tagging mechanism
Bin Ji 0002, Shasha Li 0001, Hao Xu 0015, Jie Yu 0008, Jun Ma 0015, Huijun Liu 0003 |
Sci. China Inf. Sci. | 1 |
| 2024 | A More Context-Aware Approach for Textual Adversarial Attacks Using Probability Difference-Guided Beam SearchabstractTextual adversarial attacks expose the vulnerabilities of text classifiers and can be used to improve their robustness. Previous context-aware attack models suffer from several limitations. They generally rely on out-of-date substitutes, solely consider the gold label probability, and use the greedy search when generating adversarial examples, often limiting the attack efficiency. To tackle these issues, we proposeMC-PDBS, aMoreContext-aware textual adversarial attack model usingProbabilityDifference-guidedBeamSearch. MC-PDBS generates substitutes using the newest perturbed text sequences in each attack iteration, enabling the generation of more context-aware adversarial examples. The probability difference is an overall consideration of the probabilities of all class labels, which is more effective than the gold label probability in guiding the selection of attack paths. In addition, the beam search enables MC-PDBS to search attack paths from multiple search channels, thereby avoiding the limited search space problem. Extensive experiments and human evaluation demonstrate that MC-PDBS outperforms previous best models in a series of evaluation metrics, particularly bringing up to a +19.5% attack success rate. Extensive analyses further confirm the effectiveness of MC-PDBS. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Zibo Yi, Mengxue Du, Miaomiao Li 0001, Jie Liu 0002, Zeyao Mo |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | QAE: A Hard-Label Textual Attack Considering the Comprehensive Quality of Adversarial Examples
Miaomiao Li 0001, Jie Yu 0008, Jun Ma 0015, Shasha Li 0001, Huijun Liu 0003, Mengxue Du, Bin Ji 0002 |
NLPCC (2) | 7 |
| 2023 | Learning Well-Separated and Representative Prototypes for Few-Shot Event Detection
Shasha Li 0001, Bin Ji 0002, Ting Wang 0009 |
NLPCC (2) | 3 |
| 2023 | Dynamic Multi-View Fusion Mechanism for Chinese Relation ExtractionabstractAbstract Recently, many studies incorporate external knowledge into character-level feature based models to improve the performance of Chinese relation extraction. However, these methods tend to ignore the internal information of the Chinese character and cannot filter out the noisy information of external knowledge. To address these issues, we propose a mixture-of-view-experts framework (MoVE) to dynamically learn multi-view features for Chinese relation extraction. With both the internal and external knowledge of Chinese characters, our framework can better capture the semantic information of Chinese characters. To demonstrate the effectiveness of the proposed framework, we conduct extensive experiments on three real-world datasets in distinct domains. Experimental results show consistent and significant superiority and robustness of our proposed framework. Our code and dataset will be released at: https://gitee.com/tmg-nudt/multi-view-of-expert-for-chinese-relation-extraction Bin Ji 0002, Shasha Li 0001, Jun Ma 0015, Long Peng 0002, Jie Yu 0008 |
PAKDD (1) | 2 |
| 2022 | Few-shot Named Entity Recognition with Entity-level Prototypical Network Enhanced by Dispersedly Distributed PrototypesabstractFew-shot named entity recognition (NER) enables us to build a NER system for a new domain using very few labeled examples. However, existing prototypical networks for this task suffer from roughly estimated label dependency and closely distributed prototypes, thus often causing misclassifications. To address the above issues, we propose EP-Net, an Entity-level Prototypical Network enhanced by dispersedly distributed prototypes. EP-Net builds entity-level prototypes and considers text spans to be candidate entities, so it no longer requires the label dependency. In addition, EP-Net trains the prototypes from scratch to distribute them dispersedly and aligns spans to prototypes in the embedding space using a space projection. Experimental results on two evaluation tasks and the Few-NERD settings demonstrate that EP-Net consistently outperforms the previous strong models in terms of overall performance. Extensive analyses further validate the effectiveness of EP-Net. Bin Ji 0002, Shasha Li 0001, Shaoduo Gan, Jie Yu 0008, Jun Ma 0015, Huijun Liu 0003 |
COLING | 1 |
| 2022 | Topic-Grained Text Representation-Based Model for Document Retrieval
Mengxue Du, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Bin Ji 0002, Huijun Liu 0003, Wuhang Lin, Zibo Yi |
ICANN (3) | 5 |
| 2022 | Textual adversarial attacks by exchanging text-self wordsabstractAdversarial attacks expose the vulnerability of deep neural networks. Compared to image adversarial attacks, textual adversarial attacks are more challenging due to the discrete nature of texts. Recent synonym-based methods achieve the current state-of-the-art results. However, these methods introduce new words against the original text, leading to that humans easily perceive the difference between the adversarial example and the original text. Motivated by the fact that humans are usually unaware of chaotic word order in some cases, we propose exchange-attack (EA), a concise and effective word-level textual adversarial attack model. Specifically, the EA model generates adversarial examples by exchanging words of the original text itself according to the contributions that these words make regarding classification results. Intuitively, the smaller the distance between the two exchanged words, the more difficult the chaotic word order to be perceived by humans. We thus take the word distance into consideration when generating the chaotic word orders. Extensive experiments on several text classification data sets show that the EA model consistently outperforms the selected baselines in terms of averaged after-attack accuracy, modification rate, query number, and semantic similarity. And human evaluation results reveal that humans difficultly perceive the adversarial examples generated by the EA model. In addition, quantitative and qualitative analyses further validate the effectiveness of the EA model, including that the generated adversarial examples are grammatically correct and semantically preserved. Huijun Liu 0003, Jie Yu 0008, Jun Ma 0015, Shasha Li 0001, Bin Ji 0002, Zibo Yi, Miaomiao Li 0001, Long Peng 0002, Xiaodong Liu 0004 |
Int. J. Intell. Syst. | 5 |
| 2022 | A novel bundling learning paradigm for named entity recognition
Bin Ji 0002, Yalong Xie, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Yun Ji, Huijun Liu 0003 |
Knowl. Based Syst. | 1 |
| 2021 | A Unified Summarization Model with Semantic Guide and Keyword Coverage Mechanism
Wuhang Lin, Jianling Li, Zibo Yi, Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015 |
ICANN (5) | 4 |
| 2021 | Span Representation Generation Method in Entity-Relation Joint Extraction
Yongtao Tang, Jie Yu 0008, Shasha Li 0001, Bin Ji 0002, Yusong Tan, Qingbo Wu 0003 |
ICIC (2) | 4 |
| 2020 | Span-based Joint Entity and Relation Extraction with Attention-based Span-specific and Contextual Semantic RepresentationsabstractSpan-based joint extraction models have shown their efficiency on entity recognition and relation extraction.These models regard text spans as candidate entities and span tuples as candidate relation tuples.Span semantic representations are shared in both entity recognition and relation extraction, while existing models cannot well capture semantics of these candidate entities and relations.To address these problems, we introduce a span-based joint extraction framework with attention-based semantic representations.Specially, attentions are utilized to calculate semantic representations, including span-specific and contextual ones.We further investigate effects of four attention variants in generating contextual semantic representations.Experiments show that our model outperforms previous systems and achieves state-of-the-art results on ACE2005, CoNLL2004 and ADE. Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003 |
COLING | 1 |
| 2020 | Research on Chinese medical named entity recognition based on collaborative cooperation of multiple neural network models
Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Jintao Tang, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003, Yun Ji |
J. Biomed. Informatics | 1 |