VLDB 2026 Research / reviewers in the wild / expert
Jun Ma 0015
dblp:91/4845-15
· DBLP profile ↗
51ranked-venue papers
2as first author
40since 2021 · last 2026
0000-0003-2258-0854ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 15 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SPAR: Step-wise Path Dispatching and Asymmetric Re-routing for Efficient MoE Inference
Qingxiao Zhang, Xiaopeng Li 0006, Jinzhu Kong, Xiaodong Liu 0004, Bin Ji 0002, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008 |
ICIC (26) | 7 |
| 2026 | EMSEdit: Efficient Multi-Step Meta-Learning-based Model EditingabstractLarge Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted parameter modifications, with meta-learning-based model editing (MLME) demonstrating strong effectiveness and efficiency. However, we find that MLME struggles in low-data regimes and incurs high training costs due to the use of KL divergence. To address these issues, we propose $\textbf{E}$fficient $\textbf{M}$ulti-$\textbf{S}$tep $\textbf{Edit (EMSEdit)}$, which leverages multi-step backpropagation (MSBP) to effectively capture gradient-activation mapping patterns within editing samples, performs multi-step edits per sample to enhance editing performance under limited data, and introduces norm-based regularization to preserve unedited knowledge while improving training efficiency. Experiments on two datasets and three LLMs show that EMSEdit consistently outperforms state-of-the-art methods in both sequential and batch editing. Moreover, MSBP can be seamlessly integrated into existing approaches to yield additional performance gains. Further experiments on a multi-hop reasoning editing task demonstrate EMSEdit's robustness in handling complex edits, while ablation studies validate the contribution of each design component. Our code is available at https://github.com/xpq-tech/emsedit. Xiaopeng Li 0006, Shasha Li 0001, Xi Wang 0018, Shezheng Song, Bin Ji 0002, Shangwen Wang, Jun Ma 0015, Xiaodong Liu 0004, Mina Liu, Jie Yu 0008 |
WWW | 7 |
| 2026 | Emp: enhance memory in data pruning
Jinying Xiao, Ping Li 0034, Jie Nie, Bin Ji 0002, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Qingbo Wu 0003, Jie Yu 0008 |
Data Min. Knowl. Discov. | 7 |
| 2026 | SEAttack: A self-evolving jailbreak attack to induce toxic responses for non-toxic queries in large language models
Huijun Liu 0003, Shasha Li 0001, Bin Ji 0002, Xiaohu Du, Xiaopeng Li 0006, Jun Ma 0015, Jie Yu 0008 |
Inf. Process. Manag. | 6 |
| 2026 | Evaluating Large Language Models on Named Entity RecognitionabstractLarge language models (LLMs) are popping up all over the place, and they have been gaining prominence due to their exceptional abilities in conducting various tasks. Although extensive LLM evaluation has been explored on natural language understanding tasks like text classification and sentiment analysis, evaluating LLMs on named entity recognition (NER) still remains under-explored. To fill this gap, we evaluate twenty-eight representative LLMs on thirteen datasets across five domains, whose parameters range from 3 billion to 175 billion, from four perspectives, that is, supervised fine-tuning (SFT), parameter scales, hallucinations, and prompt designs. We propose an LLM-based NER framework (LLM-NER) for the evaluation, which consists of a Recognition phase and a Check phase. Specifically, the Check guides LLMs to examine the correctness of recognized entities, which is designed to mitigate hallucinations in the NER scenario. Qualitative and quantitative evaluation analyses demonstrate that in the NER scenario: 1) SFT empowers LLMs to understand and follow human instructions; 2) LLMs' ability generally improves as their parameter scales consistently increase; 3) hallucinations exist in all evaluated LLMs, and guiding LLMs to check their outputs is a feasible way to alleviate hallucinations; and 4) all evaluated LLMs are sensitive to prompt designs. Based on the analyses, we highlight a number of promising directions for future study. Moreover, our evaluation shows high consistency with two LLM evaluation leaderboards, which evaluate LLMs on other tasks, demonstrating the rationality of our evaluation design. Bin Ji 0002, Huijun Liu 0003, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, See-Kiong Ng |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2026 | Fault Localization from the Semantic Code Search PerspectiveabstractThe software development process is characterized by an iterative cycle of continuous functionality implementation and debugging, essential for the enhancement of software quality and adaptability to changing requirements. This process incorporates two isolatedly studied tasks: Code Search (CS), which retrieves reference code from a code corpus to aid in code implementation, and Fault Localization (FL), which identifies code entities responsible for bugs within the software project to boost software debugging. The basic observation of this study is that these two tasks exhibit similarities since they both address search problems. Notably, CS techniques have demonstrated greater effectiveness than FL ones, possibly because of the precise semantic details of the required code offered by natural language queries, which are not readily accessible to FL methods. Drawing inspiration from this, we hypothesize that a fault localizer could achieve greater proficiency if semantic information about the buggy methods were made available. Based on this idea, we propose \(\texttt{CosFL}\) , an FL approach that decomposes the FL task into two steps: query generation , which describes the functionality of the problematic code in natural language, and fault retrieval , which uses CS to find program elements semantically related to the query, allowing for finishing the FL task from a CS perspective. Specifically, to depict the buggy functionalities and generate high-quality queries, \(\texttt{CosFL}\) extensively harnesses the code analysis, semantic comprehension, text generation, and decision-making capabilities of LLMs. Moreover, to enhance the accuracy of CS, \(\texttt{CosFL}\) captures varying levels of context information and employs a multi-granularity CS strategy, which facilitates a more precise identification of buggy methods from a holistic view. The evaluation on 835 real bugs from 23 Java projects shows that \(\texttt{CosFL}\) successfully localizes 324 bugs within Top-1, which significantly outperforms the state-of-the-art approaches by 26.6%–57.3%. The ablation study and sensitivity analysis further validate the importance of different components and the robustness of \(\texttt{CosFL}\) across different backend models. Yihao Qin, Shangwen Wang, Yan Lei 0005, Zhuo Zhang 0007, Bo Lin 0011, Xin Peng 0010, Jun Ma 0015, Liqian Chen, Xiaoguang Mao |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2025 | Towards Verifiable Text Generation with Generative AgentabstractText generation with citations makes it easy to verify the factuality of Large Language Models’ (LLMs) generations. Existing one-step generation studies expose distinct shortages in answer refinement and in-context demonstration matching. In light of these challenges, we propose R2-MGA, a Retrieval and Reflection Memory-augmented Generative Agent. Specifically, it first retrieves the memory bank to obtain the best-matched memory snippet, then reflects the retrieved snippet as a reasoning rationale, next combines the snippet and the rationale as the best-matched in-context demonstration. Additionally, it is capable of in-depth answer refinement with two specifically designed modules. We evaluate R2-MGA across five LLMs on the ALCE benchmark. The results reveal R2-MGA’ exceptional capabilities in text generation with citations. In particular, compared to the selected baselines, it delivers up to +58.8% and +154.7% relative performance gains on answer correctness and citation quality, respectively. Extensive analyses strongly support the motivations of R2-MGA. Bin Ji 0002, Huijun Liu 0003, Mingzhe Du, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Jie Yu 0008, See-Kiong Ng |
AAAI | 6 |
| 2025 | SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringabstractThe general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability. Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Huijun Liu 0003, Bin Ji 0002, Xi Wang 0018, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004 |
AAAI | 7 |
| 2025 | Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined BehaviorsabstractTo provide flexibility and low-level interaction capabilities, the “unsafe” tag in Rust is essential, but undermines memory safety and introduces Undefined Behaviors (UBs) that reduce safety. Eliminating UBs requires a deep understanding of Rust’s safety rules and strong typing. Traditional methods require depth analysis of code, which is laborious and depends on knowledge design. The powerful semantic understanding capabilities of LLM offer new opportunities to solve this problem. Although existing large model debugging frameworks excel in semantic tasks, limited by fixed processes and lack adaptive and dynamic adjustment capabilities. Inspired by the dual process theory of decision-making (“Fast and Slow Thinking”), we present a LLM-based framework called RustBrain that automatically and flexibly minimizes UBs in Rust projects. Fast thinking extracts features to generate solutions, while slow thinking decomposes, verifies, and generalizes them abstractly. To apply verification and generalization results to solution generation, enabling dynamic adjustments and precise outputs, RustBrain integrates two thinking through a feedback mechanism. Experimental results on Miri dataset show a 94.3% pass rate and 80.4% execution rate, improving flexibility and Rust projects safety. Renshuang Jiang, Pan Dong, Zhenling Duan, Xiaoxiang Fang, Jun Ma 0015, Shuai Zhao 0004, Zhe Jiang 0004 |
DAC | 7 |
| 2025 | Cross-Modal Reasoning-Based Unsupervised Multi-modal Entity Linking
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008 |
DASFAA (3) | 3 |
| 2025 | Multi-modal Entity Linking Model Based on Knowledge Distillation
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008 |
ICIC (24) | 3 |
| 2025 | Model Editing for LLMs4Code: How Far are we?abstractLarge Language Models for Code (LLMs4Code) have been found to exhibit outstanding performance in the software engineering domain, especially the remarkable performance in coding tasks. However, even the most advanced LLMs4Code can inevitably contain incorrect or outdated code knowledge. Due to the high cost of training LLMs4Code, it is impractical to re-train the models for fixing these problematic code knowledge. Model editing is a new technical field for effectively and efficiently correcting erroneous knowledge in LLMs, where various model editing techniques and benchmarks have been proposed recently. Despite that, a comprehensive study that thoroughly compares and analyzes the performance of the state-of-the-art model editing techniques for adapting the knowledge within LLMs4Code across various code-related tasks is notably absent. To bridge this gap, we perform the first systematic study on applying state-of-the-art model editing approaches to repair the inaccuracy of LLMs4Code. To that end, we introduce a benchmark named CLMEEval, which consists of two datasets, i.e., CoNaLa-Edit (CNLE) with 21K+ code generation samples and CodeSearchNet-Edit (CSNE) with 16K+ code summarization samples. With the help of CLMEEval, we evaluate six advanced model editing techniques on three LLMs4Code: CodeLlama (7B), CodeQwen1.5 (7B), and Stable-Code (3B). Our findings include that the external memorization-based GRACE approach achieves the best knowledge editing effectiveness and specificity (the editing does not influence untargeted knowledge), while generalization (whether the editing can generalize to other semantically-identical inputs) is a universal challenge for existing techniques. Furthermore, building on in-depth case analysis, we introduce an enhanced version of GRACE called A-GRACE, which incorporates contrastive learning to better capture the semantics of the inputs. Results demonstrate that A-GRACE notably enhances generalization while maintaining similar levels of effectiveness and specificity compared to the vanilla GRACE. Xiaopeng Li 0006, Shangwen Wang, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004, Bin Ji 0002 |
ICSE | 4 |
| 2025 | LSAQ: Layer-Specific Adaptive Quantization for Large Language Model DeploymentabstractAs Large Language Models (LLMs) demonstrate exceptional performance across various domains, deploying LLMs on edge devices has emerged as a new trend. Quantization techniques, which reduce the size and memory requirements of LLMs, are effective for deploying LLMs on resource-limited edge devices. However, existing one-size-fits-all quantization methods often fail to dynamically adjust the memory requirements of LLMs, limiting their applications to practical edge devices with various computation resources. To tackle this issue, we propose Layer-Specific Adaptive Quantization (LSAQ), a system for adaptive quantization and dynamic deployment of LLMs based on layer importance. Specifically, LSAQ evaluates the importance of LLMs’ neural layers by constructing top-k token sets from the inputs and outputs of each layer and calculating their Jaccard similarity. Based on layer importance, our system adaptively adjusts quantization strategies in real time according to the computation resource of edge devices, which applies higher quantization precision to layers with higher importance, and vice versa. Experimental results show that LSAQ consistently outperforms the selected quantization baselines in terms of perplexity and zero-shot tasks. Additionally, it can devise appropriate quantization schemes for different usage scenarios to facilitate the deployment of LLMs. Binrui Zeng, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Xiaopeng Li 0006, Shangwen Wang, Xinran Hong, Yongtao Tang |
IJCNN | 6 |
| 2025 | Identifying Knowledge Editing Types in Large Language ModelsabstractWarning: This paper contains examples of toxic text. Knowledge editing has emerged as an efficient technique for updating the knowledge of large language models (LLMs), attracting increasing attention in recent years. However, there is a lack of effective measures to prevent the malicious misuse of this technique, which could lead to harmful edits in LLMs. These malicious modifications could cause LLMs to generate toxic content, misleading users into inappropriate actions. In front of this risk, we introduce a new task, Knowledge Editing Type Identification (KETI), aimed at identifying different types of edits in LLMs, thereby providing timely alerts to users when encountering illicit edits. As part of this task, we propose KETIBench, which includes five types of harmful edits covering the most popular toxic types, as well as one benign factual edit. We develop five classical classification models and three BERT-based models as baseline identifiers for both open-source and closed-source LLMs. Our experimental results, across 92 trials involving four models and three knowledge editing methods, demonstrate that all eight baseline identifiers achieve decent identification performance, highlighting the feasibility of identifying malicious edits in LLMs. Additional analyses reveal that the performance of the identifiers is independent of the reliability of the knowledge editing methods and exhibits cross-domain generalization, enabling the identification of edits from unknown sources. All data and code are available in https://github.com/xpq-tech/KETI. Xiaopeng Li 0006, Shasha Li 0001, Shangwen Wang, Shezheng Song, Bin Ji 0002, Huijun Liu 0003, Jun Ma 0015, Jie Yu 0008 |
KDD (2) | 7 |
| 2025 | SNCD: A fast and scalable distributed near-miss code clone detector for big code based on partial index
Rulin Xie, Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan |
Future Gener. Comput. Syst. | 7 |
| 2025 | Win-Win Cooperation: Bundling Sequence and Span Models for Named Entity RecognitionabstractFor Named Entity Recognition (NER), sequence labeling-based and span-based paradigms are quite different. Previous studies have demonstrated the clear complementary advantages of the two paradigms, but few models have tried to incorporate them into a single NER model as far as we know. In our previous work, we proposed a paradigm called Bundling Learning (BL) to explore the above issue, which bundles the two NER paradigms, enabling NER models to jointly tune their parameters by weighted summing each paradigm's training loss. However, three critical issues remain unresolved: When does BL work? Why does BL work? Can BL enhance existing state-of-the-art NER models? To address the first two issues, we design three NER models: a sequence labeling-based model – SeqNER, a span-based NER model – SpanNER, and BL-NER which bundles SeqNER and SpanNER. We draw two conclusions regarding the two issues based on the experimental results on eleven NER datasets. To investigate the third issue, we apply BL to five existing state-of-the-art NER models, including three sequence labeling-based and two span-based models. Experimental results indicate consistent NER performance gains, suggesting a feasible way to construct new state-of-the-art NER systems by applying BL to the current state-of-the-art systems. Moreover, investigation results show that BL reduces both entity boundary and type prediction errors. In addition, we compare two commonly used label tagging methods and three types of span semantic representations. Bin Ji 0002, Huijun Liu 0003, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | How to Bridge the Gap Between Modalities: Survey on Multimodal Large Language ModelabstractWe explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions and answering image-based questions, bridging the gap towards real-world human-computer interactions and hinting at a potential pathway to artificial general intelligence. However, MLLMs still face challenges in addressing the semantic gap in multimodal data, which may lead to erroneous outputs, posing potential risks to society. Selecting the appropriate modality alignment method is crucial, as improper methods might require more parameters without significant performance improvements. This paper aims to explore modality alignment methods for LLMs and their current capabilities. Implementing effective modality alignment can help LLMs address environmental issues and enhance accessibility. The study surveys existing modality alignment methods for MLLMs, categorizing them into four groups: (1) Multimodal Converter, which transforms data into a format that LLMs can understand; (2) Multimodal Perceiver, which improves how LLMs percieve different types of data; (3) Tool Learning, which leverages external tools to convert data into a common format, usually text; and (4) Data-Driven Method, which teaches LLMs to understand specific data types within datasets. Shezheng Song, Xiaopeng Li 0006, Shasha Li 0001, Shan Zhao 0002, Jie Yu 0008, Jun Ma 0015, Xiaoguang Mao, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | PMET: Precise Model Editing in a TransformerabstractModel editing techniques modify a minor proportion of knowledge in Large Language Models (LLMs) at a relatively low cost, which have demonstrated notable success. Existing methods assume Transformer Layer (TL) hidden states are values of key-value memories of the Feed-Forward Network (FFN). They usually optimize the TL hidden states to memorize target knowledge and use it to update the weights of the FFN in LLMs. However, the information flow of TL hidden states comes from three parts: Multi-Head Self-Attention (MHSA), FFN, and residual connections. Existing methods neglect the fact that the TL hidden states contains information not specifically required for FFN. Consequently, the performance of model editing decreases. To achieve more precise model editing, we analyze hidden states of MHSA and FFN, finding that MHSA encodes certain general knowledge extraction patterns. This implies that MHSA weights do not require updating when new knowledge is introduced. Based on above findings, we introduce PMET, which simultaneously optimizes Transformer Component (TC, namely MHSA and FFN) hidden states, while only using the optimized TC hidden states of FFN to precisely update FFN weights. Our experiments demonstrate that PMET exhibits state-of-the-art performance on both the \textsc{counterfact} and zsRE datasets. Our ablation experiments substantiate the effectiveness of our enhancements, further reinforcing the finding that the MHSA encodes certain general knowledge extraction patterns and indicating its storage of a small amount of factual knowledge. Our code is available at \url{https://github.com/xpq-tech/PMET}. Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Jun Ma 0015, Jie Yu 0008 |
AAAI | 5 |
| 2024 | Offline Textual Adversarial Attacks against Large Language ModelsabstractThis work centers on textual adversarial attacks against large language models (LLMs) and proposes a new reproducible benchmark for future study. Unlike pre-trained language models (PLMs) which can output predicted class probabilities as feedback to instruct the generation of adversarial examples, LLMs cannot accurately provide such feedback due to their generative nature, making existing attack modes unsuitable. To address this issue, we propose Offline-Attack, an offline method tailored for LLMs that contains a novel Transformer-based Adversarial Machine Translation (AMT) framework. AMT is trained on one self-constructed large-scale adversarial dataset and used to translate original texts to adversarial examples. To mitigate training bias, we induce LLMs to generate stable prediction confidence and incorporate it into AMT training process. The evaluation, spanning four text classification datasets against LLaMA-2-13b-chat, showcases Offline-Attack’s robust performance, particularly achieving 44.3% attack success rate on average. Moreover, Offline-Attack exhibits promising attack ability to other LLMs like Vicuna-33b and ChatGPT. Our study paves the way for future study by presenting strong and reproducible baselines for textual adversarial attacks against LLMs. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Miaomiao Li 0001, Xi Wang 0018 |
IJCNN | 5 |
| 2024 | A Knowledge Graph Based Technology of Operating System Software Repository Evolution AnalysisabstractThe operating system software repository is a collection of software package resources that are used to build operating system distributions. It also serves as a platform for users to install and update the system software. Owing to the extensive array of software packages within the software repository, intricate dependency relationships among these packages, and the asynchronous nature of updates for different software packages, the evolution and upgrade of both the software repository and operating system version are challenging to predict. This presents significant hidden risks for version updates and ecological compatibility governance. To tackle this issue, the paper designs a knowledge graph for the operating system software repository. Additionally, it suggests a method for identifying changes and ensuring coherence within the software repository by utilizing the knowledge graph. We select typical open source operating system software repositories for evolutionary analysis. Results demonstrate that the proposed method, as compared to traditional dependency detection methods such as Boolean expressions, examines package dependencies and overall self-consistency of the repository from a macro perspective of the operating system. This approach effectively identifies and guides the improvement of deep inconsistent relationships. It offers a new tool and perspective for constructing, managing, and maintaining operating system software repositories. Jun Ma 0015, Xiaoling Li 0002, Xinran Hong, Jie Yu 0008, Shasha Li 0001 |
ISPA | 1 |
| 2024 | How to Pet a Two-Headed Snake? Solving Cross-Repository Compatibility Issues with HeraabstractMany programming languages and operating system communities maintain software repositories to build their own ecosystems. The repositories often provide management tools to help users using the packages. The tools are often, if not all the times, well-designed to handle intra-repository dependencies without considering inter-repository dependencies. The users, however, often need packages from different repositories, and thus may suffer from compatibility issues. We refer to these issues as Cross-repository Compatibility (CC) issues. Existing works typically focus on a single software repository and are insufficient to detect CC issues. Zhouyang Jia, Shanshan Li 0001, Ying Wang 0038, Jun Ma 0015, Xiaoling Li 0002, Xiangke Liao |
ASE | 5 |
| 2024 | DIM: Dynamic Integration of Multimodal Entity Linking with Large Language Model
Shezheng Song, Shasha Li 0001, Jie Yu 0008, Shan Zhao 0002, Xiaopeng Li 0006, Jun Ma 0015, Xiaodong Liu 0004, Xiaoguang Mao |
PRCV (5) | 6 |
| 2024 | FEAttack: A Fast and Efficient Hard-Label Textual Attack Framework
Miaomiao Li 0001, Jun Ma 0015, Jie Yu 0008, Shasha Li 0001, Huijun Liu 0003, Xi Wang 0018 |
WASA (2) | 2 |
| 2024 | Span-based joint entity and relation extraction augmented with sequence tagging mechanism
Bin Ji 0002, Shasha Li 0001, Hao Xu 0015, Jie Yu 0008, Jun Ma 0015, Huijun Liu 0003 |
Sci. China Inf. Sci. | 5 |
| 2024 | A More Context-Aware Approach for Textual Adversarial Attacks Using Probability Difference-Guided Beam SearchabstractTextual adversarial attacks expose the vulnerabilities of text classifiers and can be used to improve their robustness. Previous context-aware attack models suffer from several limitations. They generally rely on out-of-date substitutes, solely consider the gold label probability, and use the greedy search when generating adversarial examples, often limiting the attack efficiency. To tackle these issues, we proposeMC-PDBS, aMoreContext-aware textual adversarial attack model usingProbabilityDifference-guidedBeamSearch. MC-PDBS generates substitutes using the newest perturbed text sequences in each attack iteration, enabling the generation of more context-aware adversarial examples. The probability difference is an overall consideration of the probabilities of all class labels, which is more effective than the gold label probability in guiding the selection of attack paths. In addition, the beam search enables MC-PDBS to search attack paths from multiple search channels, thereby avoiding the limited search space problem. Extensive experiments and human evaluation demonstrate that MC-PDBS outperforms previous best models in a series of evaluation metrics, particularly bringing up to a +19.5% attack success rate. Extensive analyses further confirm the effectiveness of MC-PDBS. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Zibo Yi, Mengxue Du, Miaomiao Li 0001, Jie Liu 0002, Zeyao Mo |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | WMWatcher: Preventing Workload-Related Misconfigurations in Production EnvironmentabstractAmong the misconfigurations with increasing preva-lence and severity in recent years, workload-related misconfigu-rations, i.e. misconfigurations under certain workloads with valid configuration values, account for a significant portion. Since the runtime constraints of configuration parameters are influenced by workloads, piror researches could not handle workload-related misconfigurations at present. To solve the situation mentioned above, we conducted an empirical study on how configuration variables interact with other program variables, and summarized five handling type of the interactions happen in branch statements. Based on the study, we proposed WMWatcher to help system admins to prevent workload-related misconfigurations in production environment. WMWatcher infers the runtime constraints of configuration parameters under certain workload by instrumenting probes in source code and monitoring the corresponding status. The experiments on seven open-source software systems proved that WMWatcher could automatically instrument proper probes while bringing only 2.33% extra runtime overhead at most. And the case study demonstrates the effectiveness of WMWatcher in preventing workload-related misconfigurations in real-world scenarios. Shulin Zhou, Zhijie Jiang, Shanshan Li 0001, Xiaodong Liu 0004, Zhouyang Jia, Yuanliang Zhang, Jun Ma 0015, Haibo Mi |
APSEC | 7 |
| 2023 | QAE: A Hard-Label Textual Attack Considering the Comprehensive Quality of Adversarial Examples
Miaomiao Li 0001, Jie Yu 0008, Jun Ma 0015, Shasha Li 0001, Huijun Liu 0003, Mengxue Du, Bin Ji 0002 |
NLPCC (2) | 3 |
| 2023 | Dynamic Multi-View Fusion Mechanism for Chinese Relation ExtractionabstractAbstract Recently, many studies incorporate external knowledge into character-level feature based models to improve the performance of Chinese relation extraction. However, these methods tend to ignore the internal information of the Chinese character and cannot filter out the noisy information of external knowledge. To address these issues, we propose a mixture-of-view-experts framework (MoVE) to dynamically learn multi-view features for Chinese relation extraction. With both the internal and external knowledge of Chinese characters, our framework can better capture the semantic information of Chinese characters. To demonstrate the effectiveness of the proposed framework, we conduct extensive experiments on three real-world datasets in distinct domains. Experimental results show consistent and significant superiority and robustness of our proposed framework. Our code and dataset will be released at: https://gitee.com/tmg-nudt/multi-view-of-expert-for-chinese-relation-extraction Bin Ji 0002, Shasha Li 0001, Jun Ma 0015, Long Peng 0002, Jie Yu 0008 |
PAKDD (1) | 4 |
| 2023 | MulCS: Towards a Unified Deep Representation for Multilingual Code SearchabstractCode search aims to search for relevant code snippets through queries, which has become an essential requirement to assist programmers in software development. With the availability of large and rapidly growing source code repositories covering various languages, multilingual code search can leverage more training data to learn complementary information across languages. Contrastive learning can naturally understand the similarity between functionally equivalent code across different languages by narrowing the distance between objects with the same function while keeping dissimilar objects further apart. Some works exist addressing monolingual code search problems with contrastive learning, however, they mainly exploit every specific programming language’s textual semantics or syntactic structures for code representation. Due to the high diversity of different languages in terms of syntax, format, and structure, these methods limit the performance of contrastive learning in multilingual training. To bridge this gap, we propose a unified semantic graph representation approach toward multilingual code search called MulCS. Specifically, we first design a general semantic graph construction strategy across different languages by Intermediate Representation (IR). Furthermore, we introduce the contrastive learning module integrated into a gated graph neural network (GGNN) to enhance query-multilingual code matching. The extensive experiments on three representative languages illustrate that our method outperforms state-of-the-art models by 10.7% to 77.5% in terms of MRR on average. Yingwei Ma, Yue Yu 0001, Shanshan Li 0001, Zhouyang Jia, Jun Ma 0015, Rulin Xu, Wei Dong 0006, Xiangke Liao |
SANER | 5 |
| 2023 | When Database Meets New Storage Devices: Understanding and Exposing Performance Mismatches via ConfigurationsabstractNVMe SSD hugely boosts the I/O speed, with up to GB/s throughput and microsecond-level latency. Unfortunately, DBMS users can often find their high-performanced storage devices tend to deliver less-than-expected or even worse performance when compared to their traditional peers. While many works focus on proposing new DBMS designs to fully exploit NVMe SSDs, few systematically study the symptoms, root causes and possible detection methods of such performance mismatches on existing databases. In this paper, we start with an empirical study where we systematically expose and analyze the performance mismatches on six popular databases via controlled configuration tuning. From the study, we find that all six databases can suffer from performance mismatches. Moreover, we conclude that the root causes can be categorized as databases' unawareness of new storage devices characteristics in I/O size, I/O parallelism and I/O sequentiality. We report 17 mismatches to developers and 15 are confirmed. Additionally, we realize testing all configuration knobs yields low efficiency. Therefore, we propose a fast performance mismatch detection framework and evaluation shows that our framework brings two orders of magnitude speedup than baseline without sacrificing effectiveness. Haochen He, Erci Xu, Shanshan Li 0001, Zhouyang Jia, Si Zheng 0003, Yue Yu 0001, Jun Ma 0015, Xiangke Liao |
Proc. VLDB Endow. | 7 |
| 2022 | Few-shot Named Entity Recognition with Entity-level Prototypical Network Enhanced by Dispersedly Distributed PrototypesabstractFew-shot named entity recognition (NER) enables us to build a NER system for a new domain using very few labeled examples. However, existing prototypical networks for this task suffer from roughly estimated label dependency and closely distributed prototypes, thus often causing misclassifications. To address the above issues, we propose EP-Net, an Entity-level Prototypical Network enhanced by dispersedly distributed prototypes. EP-Net builds entity-level prototypes and considers text spans to be candidate entities, so it no longer requires the label dependency. In addition, EP-Net trains the prototypes from scratch to distribute them dispersedly and aligns spans to prototypes in the embedding space using a space projection. Experimental results on two evaluation tasks and the Few-NERD settings demonstrate that EP-Net consistently outperforms the previous strong models in terms of overall performance. Extensive analyses further validate the effectiveness of EP-Net. Bin Ji 0002, Shasha Li 0001, Shaoduo Gan, Jie Yu 0008, Jun Ma 0015, Huijun Liu 0003 |
COLING | 5 |
| 2022 | Topic-Grained Text Representation-Based Model for Document Retrieval
Mengxue Du, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Bin Ji 0002, Huijun Liu 0003, Wuhang Lin, Zibo Yi |
ICANN (3) | 4 |
| 2022 | KylinTune: DQN-based Energy-efficient Model for Browser in Mobile DevicesabstractBrowser is a key application for mobile devices and its power management is significant given that mobile devices are power-sensitive. Currently, dynamic voltage and frequency scaling (DVFS) and energy-aware scheduling (EAS) techniques have been implemented in mobile devices for energy savings. However, it is still challenging to achieve an energy-efficient mobile browser due to the varied content of webpages that need different resources to fetch, parser, render, etc. An ideal power governor should adjust CPU frequency dynamically according to webpage characteristics, but the current governor is configured statically and webpage-agnostic. To address the above issues, we propose KylinTune, an energy-efficient model for mobile browsers. The KylinTune is based on Deep-Q Network (DQN), a reinforcement learning technique. KylinTune learns from the browser runtime and adjusts CPU frequency to an optimal execution speed for a specific webpage based on EAS. We apply KylinTune to the Chromium browser on Google Pixel2 XL and evaluate it on the top 100 popular websites. Experimental results show that KylinTune achieves 14.51%–24% energy savings in different loading environments, with trivial quality of service (QoS) degradation. Hao Xu 0015, Long Peng 0002, Xiaodong Liu 0004, Menglin Zhang, Jun Ma 0015, Jie Yu 0008, Zibo Yi |
IPCCC | 5 |
| 2022 | Textual adversarial attacks by exchanging text-self wordsabstractAdversarial attacks expose the vulnerability of deep neural networks. Compared to image adversarial attacks, textual adversarial attacks are more challenging due to the discrete nature of texts. Recent synonym-based methods achieve the current state-of-the-art results. However, these methods introduce new words against the original text, leading to that humans easily perceive the difference between the adversarial example and the original text. Motivated by the fact that humans are usually unaware of chaotic word order in some cases, we propose exchange-attack (EA), a concise and effective word-level textual adversarial attack model. Specifically, the EA model generates adversarial examples by exchanging words of the original text itself according to the contributions that these words make regarding classification results. Intuitively, the smaller the distance between the two exchanged words, the more difficult the chaotic word order to be perceived by humans. We thus take the word distance into consideration when generating the chaotic word orders. Extensive experiments on several text classification data sets show that the EA model consistently outperforms the selected baselines in terms of averaged after-attack accuracy, modification rate, query number, and semantic similarity. And human evaluation results reveal that humans difficultly perceive the adversarial examples generated by the EA model. In addition, quantitative and qualitative analyses further validate the effectiveness of the EA model, including that the generated adversarial examples are grammatically correct and semantically preserved. Huijun Liu 0003, Jie Yu 0008, Jun Ma 0015, Shasha Li 0001, Bin Ji 0002, Zibo Yi, Miaomiao Li 0001, Long Peng 0002, Xiaodong Liu 0004 |
Int. J. Intell. Syst. | 3 |
| 2022 | Towards an Efficient and Robust Adversarial Attack Against Neural Text ClassifierabstractAdversarial attack is a serious threat to neural network-based natural language processing applications. Adversarial attack uses tiny well-crafted perturbations to mislead neural networks. While existing adversarial text attacks can achieve good attack effects, they still do not guarantee efficiency and robustness. The adversarial text attacks are more efficient if they use less perturbation to achieve a higher attack success rate. The attacks are more robust if they can achieve a higher success rate when defense strategies are applied. To improve the efficiency and robustness of the adversarial attack, we propose SMAL: Saliency Map Attack with Levenshtein-similarity. The proposed attack consists of two parts: (1) The saliency map measures the perturbation priority of each word. It considers not only the influence of each word on the classification result but also how to maintain the misled classification result to improve the robustness of the attack. (2) Levenshtein-similarity network embeds words into edit distance space. When perturbing sentences, some words are replaced by substitutions with less edit distance. This can reduce the amount of modification, which improves the efficiency of the attack. Since the words are embedded in edit distance space rather than semantic space, the semantic-based defense is not effective for this attack, which improves the robustness. The experiments show that SMAL achieves a higher attack success rate with fewer perturbations. Also, the proposed attack is better when attacking a classifier defended by adversarial training. Zibo Yi, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2022 | A novel bundling learning paradigm for named entity recognition
Bin Ji 0002, Yalong Xie, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Yun Ji, Huijun Liu 0003 |
Knowl. Based Syst. | 5 |
| 2021 | A Unified Summarization Model with Semantic Guide and Keyword Coverage Mechanism
Wuhang Lin, Jianling Li, Zibo Yi, Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015 |
ICANN (5) | 7 |
| 2021 | Many-To-Many Chinese ICD-9 Terminology Standardization Based on Neural Networks
Shasha Li 0001, Jie Yu 0008, Yusong Tan, Jun Ma 0015, Qingbo Wu 0003 |
ICIC (2) | 5 |
| 2021 | Combating Word-level Adversarial Text with Robust Adversarial TrainingabstractNLP models perform well on many tasks, but they are also easy to be fooled by adversarial examples. A small perturbation can change the output of the deep neural network model. This kind of perturbation is hard to be perceived by humans, especially adversarial examples generated by word-level adversarial attack. Character-level adversarial attack can be defended by grammar detection and word recognition. The existing word-level textual adversarial attacks are based on synonym replacement, so adversarial texts usually have correct grammar and semantics. The defense of word-level adversarial attack is more challenging. In this paper, we propose a framework which is called Robust Adversarial Training (RAT) to defend against word-level adversarial attacks. RAT enhances the model by combining adversarial training and data perturbation during training. Our experiments on two datasets show that the model based on our framework can effectively defend against word-level adversarial attacks. Compared with the existing defense methods, the model trained under RAT has a higher defense success rate on 1000 adversarial examples. In addition, the accuracy of our model on the standard testing set is also better than the existing defense methods, and the accuracy is very close to or even higher than that of the standard model. Xiaohu Du, Jie Yu 0008, Shasha Li 0001, Zibo Yi, Jun Ma 0015 |
IJCNN | 6 |
| 2021 | FastDCF: A Partial Index Based Distributed and Scalable Near-Miss Code Clone Detection Approach for Very Large Code Repositories
Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan |
PDCAT | 5 |
| 2020 | Span-based Joint Entity and Relation Extraction with Attention-based Span-specific and Contextual Semantic RepresentationsabstractSpan-based joint extraction models have shown their efficiency on entity recognition and relation extraction.These models regard text spans as candidate entities and span tuples as candidate relation tuples.Span semantic representations are shared in both entity recognition and relation extraction, while existing models cannot well capture semantics of these candidate entities and relations.To address these problems, we introduce a span-based joint extraction framework with attention-based semantic representations.Specially, attentions are utilized to calculate semantic representations, including span-specific and contextual ones.We further investigate effects of four attention variants in generating contextual semantic representations.Experiments show that our model outperforms previous systems and achieves state-of-the-art results on ACE2005, CoNLL2004 and ADE. Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003 |
COLING | 4 |
| 2020 | Research on Chinese medical named entity recognition based on collaborative cooperation of multiple neural network models
Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Jintao Tang, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003, Yun Ji |
J. Biomed. Informatics | 4 |
| 2020 | Build real-time communication for hybrid dual-OS system
Pan Dong, Zhe Jiang 0004, Alan Burns 0001, Jun Ma 0015 |
J. Syst. Archit. | 5 |
| 2019 | CLASC: A Changelog Based Automatic Code Source Classification Method for Operating System PackagesabstractOpen source represents an important way in which today's software is developed. The adoption of open source software continues to accelerate because of the great potential it offers, such as productivity improvement, cost savings and quicker innovation. While the complexity and the size of software composition grow, it becomes difficult to effectively scan and track the code source, especially for software with tremendous scale of code, such as operating systems. So far, existing work on open source components mainly focus on how to mitigate potential license incompliance, to reduce potential security risks introduced by open source vulnerabilities, and to detect and match open source components in the code. To ensure code traceability and manageability for large scale mixed-source operating system, we believe it is beneficial to automatically distinguish sources of the system code in the granularity of software packages and manage them separately. However, according to the literature, there is a lack of relevant work in this area. In this paper, we first classify the packages into three categories in terms of code source from the perspective of OS developers and maintainers. Then we propose CLASC, an efficient code source classification algorithm. With the capability of package info extraction and analysis, CLASC can classify software packages into the defined categories according to their changelog info. And we design and implement KyAnalyzer, a Web-based package management and code source analysis platform. It provides automatic code source analyzing services and is capable of managing OS packages differentially according to their different categories of code source with CLASC incorporated as a component of it. Experimental results show the correctness and efficiency of the Web-enabled package source classifier. Yi Ren 0008, Jianbo Guan, Jun Ma 0015, Yusong Tan, Qingbo Wu 0003 |
APSEC | 3 |
| 2019 | Work-in-Progress: Real-Time RPC for Hybrid Dual-OS SystemabstractFor the power and space sensitive systems such as automotive/avionic computers, an important trend is isolating and integrating multiple Operating Systems (OSs) in one physical platform, which is named as hybrid multi-OS system. Generally, in a commonly used hybrid dual-OS system, a RTOS (realtime operating system) and a GPOS (general-purpose operating system) are integrated. Cooperation (among the OSs) is a vital feature of a hybrid system to obtain the necessary capabilities, and inter-OS communication is the key. However, it is difficult to satisfy the real-time metrics of inter-OS communication required by the RTOS, due to the uncertainty in communication maintenance and the time-sharing policy of the GPOS. This paper aims to build a time predictable and secure RPC mechanism (i.e., the primary and critical communication unit in a hybrid multi-OS system). Afterwards, a real-time RPC scheme (termed RTRGRPC) is proposed, which is applied to a ready-built TrustZonebased hybrid dual-OS system (i.e., TZDKS). RTRG-RPC achieves accurate time control through three mechanisms: SGI message transforming, interrupt handler RPC servicing, and priorityswapping. Evaluations show that RTRG-RPC can achieve realtime predictability and can also reduce priority inversion. Pan Dong, Zhe Jiang 0004, Alan Burns 0001, Jun Ma 0015 |
RTSS | 5 |
| 2016 | A novel optimization scheme for caching in locality-aware P2P networksabstractDeploying cache has been generally adopted by Internet service providers (ISPs) to mitigate P2P traffic in recent years. Most traditional caching algorithms are designed for locality-unaware P2P networks, which mainly consider the requested frequency of contents as the principle of caching policies. However, in more prevalent locality-aware conditions with biased neighbor-selection policies, the existing caching schemes can hardly optimize the situation. In this paper we show that, what need to be cached in locality-aware conditions are the contents that can not be well provided by local neighbors, rather than the contents which are requested most frequently. Therefore, states of local neighbors should be taken into consideration in caching policies. We first present a new model in which P2P cache and locality-aware neighbor selection work together. We focus on inter-ISP traffic and available bandwidth of users in order to benefit both ISPs and users. Based on the mathematical model, a novel caching algorithm is proposed which considers replacement and allocation policies together. According to trace-driven simulations, the proposed algorithm outperforms other two representative caching algorithms in various scenarios. Shaoduo Gan, Jiexin Zhang 0001, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003 |
ISCC | 5 |
| 2016 | An Optimized DHT for Linux Package DistributionabstractThe rapid rising of Linux users requires P2P, an efficient content transport method, to distribute packages. Different from traditional streaming P2P systems, a P2P package distribution system is hazarded by the special characteristics of small package size and hot packages. Due to the small package size, DHT search performance, which is rarely considered in traditional P2P system, comes to be an important factor in package distribution process. To improve the DHT search performance in this circumstance, we propose four kinds of optimizations as follow. Firstly, Fast-Response is proposed to eliminate useless searches after having found the target pair. Secondly, LRU Cache is proposed to reduce search hops on the same package. Thirdly, Leap Cache is proposed to reduce the cache redundancy. Finally, Probability Cache, gathering those optimizing above and considering hot packages in addition, is proposed to get increase of cache hit rate and overall efficiency improvement. We simulate our optimizations using PeerSim platform. The results show that all the four optimizations get considerable improvement in performance. With the best situation of Probability Cache, 86.22% delay time of original Kademlia is saved. Qi Zhang 0028, Jie Yu 0008, Lei Luo 0002, Jun Ma 0015, Qingbo Wu 0003, Shasha Li 0001 |
ISPDC | 4 |
| 2016 | ERPC: An Edge-Resources Based Framework to Reduce Bandwidth Cost in the Personal Cloud
Shaoduo Gan, Jie Yu 0008, Xiaoling Li 0002, Jun Ma 0015, Lei Luo 0002, Qingbo Wu 0003, Shasha Li 0001 |
WAIM (2) | 4 |
| 2013 | Efficient revocation in ciphertext-policy attribute-based encryption based cryptographic cloud storageabstractIt is secure for customers to store and share their sensitive data in the cryptographic cloud storage. However, the revocation operation is a sure performance killer in the cryptographic access control system. To optimize the revocation procedure, we present a new efficient revocation scheme which is efficient, secure, and unassisted. In this scheme, the original data are first divided into a number of slices, and then published to the cloud storage. When a revocation occurs, the data owner needs only to retrieve one slice, and re-encrypt and re-publish it. Thus, the revocation process is accelerated by affecting only one slice instead of the whole data. We have applied the efficient revocation scheme to the ciphertext-policy attribute-based encryption (CP-ABE) based cryptographic cloud storage. The security analysis shows that our scheme is computationally secure. The theoretically evaluated and experimentally measured performance results show that the efficient revocation scheme can reduce the data owner’s workload if the revocation occurs frequently. Zhiying Wang 0003, Jun Ma 0015, Jiangjiang Wu, Songzhu Mei, Jiangchun Ren |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | SWHash: An Efficient Data Integrity Verification Scheme Appropriate for USB Flash DiskabstractData integrity verification is utmost important in trusted computing and Merkle trees are usually employed in implementation. However, the efficiency of data authentication is regarded as the main bottleneck in performance. In this paper, we propose an efficient data authentication protocol appropriate for a USB flash disk, named UTrustDisk (a trust-based intelligent disk). In our scheme, verification is speed up by using WH universal hash function and speculative caching. WH algorithm can hash message into a short digest at a high speed and the collision probability is almost negligible. Speculative caching will cache the potential hot chunks which can reduce the memory bandwidth pollution. In our experiments the success rate of speculation reaches 94.5% because the UTrustDisk's access mode is usually sequentially. The comparative experiment results show that SWHash average write throughput is 44.8% higher than NH scheme and 316% higher than SHA-1 scheme. Zhiying Wang 0003, Jiangjiang Wu, Songzhu Mei, Jiangchun Ren, Jun Ma 0015 |
TrustCom | 6 |
| 2008 | Research of a Secure File System for Protection of Intellectual Property RightabstractThis paper analyses the architecture of current secure file system and the security needs for protection of intellectual property rights and especially the major problems of it. Then we propose a secure data container model based on data encapsulation from the realization concept of virtual file system (VFS) in Linux. Based on this model, we design and implant a secure file system IPR-SFS on Windows platform for protection of intellectual property which achieves perfect combination between data encryption and access control. Compared with the previous systems, the IPR-SFS file system is more convenient and flexible, safe and scalable, also comparable to the existing file. Jun Ma 0015, Jiangchun Ren, Zhiying Wang 0003, Yaokai Zhu |
WAIM | 1 |