Xiaodong Liu 0004

dblp:65/622-4 · DBLP profile ↗
← Back
35ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-9800-6886ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 4 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Efficient 4-bit Quantized Inference for LLMs on RISC-V via RVV-Based GGUF Weight Layout Reconfiguration
Long Peng 0002, Xiaodong Liu 0004, Jie Yu 0008
ICIC (5)5
2026 SPAR: Step-wise Path Dispatching and Asymmetric Re-routing for Efficient MoE Inference
Qingxiao Zhang, Xiaopeng Li 0006, Jinzhu Kong, Xiaodong Liu 0004, Bin Ji 0002, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008
ICIC (26)4
2026 AdaSSF: A Token-Efficient Adaptive Framework for Document-Level Knowledge Graph Extraction via Semantic-Spatial Fusion
Long Peng 0002, Xiaodong Liu 0004, Jie Yu 0008
KSEM (3)4
2026 EMSEdit: Efficient Multi-Step Meta-Learning-based Model Editing
abstract
Large Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted parameter modifications, with meta-learning-based model editing (MLME) demonstrating strong effectiveness and efficiency. However, we find that MLME struggles in low-data regimes and incurs high training costs due to the use of KL divergence. To address these issues, we propose $\textbf{E}$fficient $\textbf{M}$ulti-$\textbf{S}$tep $\textbf{Edit (EMSEdit)}$, which leverages multi-step backpropagation (MSBP) to effectively capture gradient-activation mapping patterns within editing samples, performs multi-step edits per sample to enhance editing performance under limited data, and introduces norm-based regularization to preserve unedited knowledge while improving training efficiency. Experiments on two datasets and three LLMs show that EMSEdit consistently outperforms state-of-the-art methods in both sequential and batch editing. Moreover, MSBP can be seamlessly integrated into existing approaches to yield additional performance gains. Further experiments on a multi-hop reasoning editing task demonstrate EMSEdit's robustness in handling complex edits, while ablation studies validate the contribution of each design component. Our code is available at https://github.com/xpq-tech/emsedit.
Xiaopeng Li 0006, Shasha Li 0001, Xi Wang 0018, Shezheng Song, Bin Ji 0002, Shangwen Wang, Jun Ma 0015, Xiaodong Liu 0004, Mina Liu, Jie Yu 0008
WWW8
2026 Emp: enhance memory in data pruning
Jinying Xiao, Ping Li 0034, Jie Nie, Bin Ji 0002, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Qingbo Wu 0003, Jie Yu 0008
Data Min. Knowl. Discov.6
2025 Towards Verifiable Text Generation with Generative Agent
abstract
Text generation with citations makes it easy to verify the factuality of Large Language Models’ (LLMs) generations. Existing one-step generation studies expose distinct shortages in answer refinement and in-context demonstration matching. In light of these challenges, we propose R2-MGA, a Retrieval and Reflection Memory-augmented Generative Agent. Specifically, it first retrieves the memory bank to obtain the best-matched memory snippet, then reflects the retrieved snippet as a reasoning rationale, next combines the snippet and the rationale as the best-matched in-context demonstration. Additionally, it is capable of in-depth answer refinement with two specifically designed modules. We evaluate R2-MGA across five LLMs on the ALCE benchmark. The results reveal R2-MGA’ exceptional capabilities in text generation with citations. In particular, compared to the selected baselines, it delivers up to +58.8% and +154.7% relative performance gains on answer correctness and citation quality, respectively. Extensive analyses strongly support the motivations of R2-MGA.
Bin Ji 0002, Huijun Liu 0003, Mingzhe Du, Shasha Li 0001, Xiaodong Liu 0004, Jun Ma 0015, Jie Yu 0008, See-Kiong Ng
AAAI5
2025 SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering
abstract
The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability.
Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Huijun Liu 0003, Bin Ji 0002, Xi Wang 0018, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004
AAAI9
2025 Cross-Modal Reasoning-Based Unsupervised Multi-modal Entity Linking
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008
DASFAA (3)5
2025 Accelerating LLM Inference on RISC-V Edge Devices via Vector Extension Optimization
Long Peng 0002, Wenzhu Wang, Ke Li 0026, Binrui Zeng, Jie Yu 0008, Xiaodong Liu 0004
ICIC (3)7
2025 Multi-modal Entity Linking Model Based on Knowledge Distillation
Yongtao Tang, Shasha Li 0001, Jun Ma 0015, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008
ICIC (24)5
2025 Model Editing for LLMs4Code: How Far are we?
abstract
Large Language Models for Code (LLMs4Code) have been found to exhibit outstanding performance in the software engineering domain, especially the remarkable performance in coding tasks. However, even the most advanced LLMs4Code can inevitably contain incorrect or outdated code knowledge. Due to the high cost of training LLMs4Code, it is impractical to re-train the models for fixing these problematic code knowledge. Model editing is a new technical field for effectively and efficiently correcting erroneous knowledge in LLMs, where various model editing techniques and benchmarks have been proposed recently. Despite that, a comprehensive study that thoroughly compares and analyzes the performance of the state-of-the-art model editing techniques for adapting the knowledge within LLMs4Code across various code-related tasks is notably absent. To bridge this gap, we perform the first systematic study on applying state-of-the-art model editing approaches to repair the inaccuracy of LLMs4Code. To that end, we introduce a benchmark named CLMEEval, which consists of two datasets, i.e., CoNaLa-Edit (CNLE) with 21K+ code generation samples and CodeSearchNet-Edit (CSNE) with 16K+ code summarization samples. With the help of CLMEEval, we evaluate six advanced model editing techniques on three LLMs4Code: CodeLlama (7B), CodeQwen1.5 (7B), and Stable-Code (3B). Our findings include that the external memorization-based GRACE approach achieves the best knowledge editing effectiveness and specificity (the editing does not influence untargeted knowledge), while generalization (whether the editing can generalize to other semantically-identical inputs) is a universal challenge for existing techniques. Furthermore, building on in-depth case analysis, we introduce an enhanced version of GRACE called A-GRACE, which incorporates contrastive learning to better capture the semantics of the inputs. Results demonstrate that A-GRACE notably enhances generalization while maintaining similar levels of effectiveness and specificity compared to the vanilla GRACE.
Xiaopeng Li 0006, Shangwen Wang, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004, Bin Ji 0002
ICSE6
2025 LSAQ: Layer-Specific Adaptive Quantization for Large Language Model Deployment
abstract
As Large Language Models (LLMs) demonstrate exceptional performance across various domains, deploying LLMs on edge devices has emerged as a new trend. Quantization techniques, which reduce the size and memory requirements of LLMs, are effective for deploying LLMs on resource-limited edge devices. However, existing one-size-fits-all quantization methods often fail to dynamically adjust the memory requirements of LLMs, limiting their applications to practical edge devices with various computation resources. To tackle this issue, we propose Layer-Specific Adaptive Quantization (LSAQ), a system for adaptive quantization and dynamic deployment of LLMs based on layer importance. Specifically, LSAQ evaluates the importance of LLMs’ neural layers by constructing top-k token sets from the inputs and outputs of each layer and calculating their Jaccard similarity. Based on layer importance, our system adaptively adjusts quantization strategies in real time according to the computation resource of edge devices, which applies higher quantization precision to layers with higher importance, and vice versa. Experimental results show that LSAQ consistently outperforms the selected quantization baselines in terms of perplexity and zero-shot tasks. Additionally, it can devise appropriate quantization schemes for different usage scenarios to facilitate the deployment of LLMs.
Binrui Zeng, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Xiaopeng Li 0006, Shangwen Wang, Xinran Hong, Yongtao Tang
IJCNN3
2025 A Survey of AI Inference Technologies for On-Device Systems
abstract
In recent years, artificial intelligence(AI) technologies represented by foundation models have experienced rapid development. Concurrently, On-device AI inference has become the primary approach for intelligent technology applications, offering advantages such as low latency, high security, and personalization. However, due to the limited resources of on-device systems, on-device AI inference faces new challenges, including improving computational efficiency, optimizing task parallelism, and model optimization. This survey addresses these challenges from a software and algorithmic perspective, focusing on three key areas: Operator Computation: Explores methods to accelerate matrix multiplication and convolution, as well as techniques like operator fusion and vectorized computation. Task Inference: Analyzes heterogeneous and distributed computing, memory allocation, and energy-efficient tuning to improve the parallel execution and energy efficiency of inference tasks. AI Models: Covers model compression, lookup table quantization, and model architecture design to reduce computational complexity and storage requirements. By analyzing these areas, the survey aims to improve inference speed, reduce resource dependency, and provide insights into the future trends of on-device AI technology.
Wenzhu Wang, Ke Li 0026, Bin Ji 0002, Xiaodong Liu 0004, Jie Yu 0008, Qingbo Wu 0003
IEEE Internet Things J.4
2024 A Learning-Based and Network-Aware Power Management for Mobile Devices
abstract
This paper proposes a deep reinforcement learning-based power management method for mobile devices. By learning the load characteristics of the device under different usage scenarios and considering the influence of network conditions on power consumption, the CPU and GPU frequencies are dynamically adjusted for multiple application scenarios. At the same time, a “SLIDER” adjustment strategy is proposed, and combined with the system default adjustment strategy, which reduces the difficulty of adjustment and more fully utilizes the middle adjustable frequency of the CPU. The proposed method reduces power consumption by 5.3%-18% compared to state-of-the-art method.
Jiangjie Huang, Long Peng 0002, Xiaodong Liu 0004, Jie Yu 0008, Wenzhu Wang
COMPSAC4
2024 DIM: Dynamic Integration of Multimodal Entity Linking with Large Language Model
Shezheng Song, Shasha Li 0001, Jie Yu 0008, Shan Zhao 0002, Xiaopeng Li 0006, Jun Ma 0015, Xiaodong Liu 0004, Xiaoguang Mao
PRCV (5)7
2023 WMWatcher: Preventing Workload-Related Misconfigurations in Production Environment
abstract
Among the misconfigurations with increasing preva-lence and severity in recent years, workload-related misconfigu-rations, i.e. misconfigurations under certain workloads with valid configuration values, account for a significant portion. Since the runtime constraints of configuration parameters are influenced by workloads, piror researches could not handle workload-related misconfigurations at present. To solve the situation mentioned above, we conducted an empirical study on how configuration variables interact with other program variables, and summarized five handling type of the interactions happen in branch statements. Based on the study, we proposed WMWatcher to help system admins to prevent workload-related misconfigurations in production environment. WMWatcher infers the runtime constraints of configuration parameters under certain workload by instrumenting probes in source code and monitoring the corresponding status. The experiments on seven open-source software systems proved that WMWatcher could automatically instrument proper probes while bringing only 2.33% extra runtime overhead at most. And the case study demonstrates the effectiveness of WMWatcher in preventing workload-related misconfigurations in real-world scenarios.
Shulin Zhou, Zhijie Jiang, Shanshan Li 0001, Xiaodong Liu 0004, Zhouyang Jia, Yuanliang Zhang, Jun Ma 0015, Haibo Mi
APSEC4
2023 ConfTainter: Static Taint Analysis For Configuration Options
abstract
The prevalence and severity of software configuration-induced issues have driven the design and development of a number of detection and diagnosis techniques. Many of these techniques need to perform static taint analysis on configuration-related variables to analyze the data flow, control flow, and execution paths given by configuration options. However, existing taint analysis or static slicer tools are not suitable for configuration analysis due to the complex effects of configuration on program behaviors. In this experience paper, we conducted an empirical study on the propagation policy of configuration options. We concluded four rules of how configurations affect program behaviors, among which implicit data-flow and control-flow propagation are often ignored by existing tools. We report our experience designing and implementing a taint analysis infrastructure for configurations, ConfTainter. It can support various kinds of configuration analysis, e.g., explicit or implicit analysis for data or control flow. Based on the infrastructure, researchers and developers can easily implement analysis techniques for different configuration-related targets, e.g., misconfiguration detection. We evaluated the effectiveness of ConfTainter on 5 popular open-source systems. The result shows that the accuracy rate of data- and control-flow analysis is 96.1% and 97.7%, and the recall rate is 94.2% and 95.5%, respectively. We also apply ConfTainter to two types of configuration-related tasks: misconfiguration detection and configuration-related bug detection. The result shows that ConfTainter is highly applicable for configuration-related tasks with a few lines of code.
Teng Wang 0004, Haochen He, Xiaodong Liu 0004, Shanshan Li 0001, Zhouyang Jia, Yu Jiang 0001, Qing Liao 0001, Wang Li 0003
ASE3
2022 KylinTune: DQN-based Energy-efficient Model for Browser in Mobile Devices
abstract
Browser is a key application for mobile devices and its power management is significant given that mobile devices are power-sensitive. Currently, dynamic voltage and frequency scaling (DVFS) and energy-aware scheduling (EAS) techniques have been implemented in mobile devices for energy savings. However, it is still challenging to achieve an energy-efficient mobile browser due to the varied content of webpages that need different resources to fetch, parser, render, etc. An ideal power governor should adjust CPU frequency dynamically according to webpage characteristics, but the current governor is configured statically and webpage-agnostic. To address the above issues, we propose KylinTune, an energy-efficient model for mobile browsers. The KylinTune is based on Deep-Q Network (DQN), a reinforcement learning technique. KylinTune learns from the browser runtime and adjusts CPU frequency to an optimal execution speed for a specific webpage based on EAS. We apply KylinTune to the Chromium browser on Google Pixel2 XL and evaluate it on the top 100 popular websites. Experimental results show that KylinTune achieves 14.51%–24% energy savings in different loading environments, with trivial quality of service (QoS) degradation.
Hao Xu 0015, Long Peng 0002, Xiaodong Liu 0004, Menglin Zhang, Jun Ma 0015, Jie Yu 0008, Zibo Yi
IPCCC3
2022 Textual adversarial attacks by exchanging text-self words
abstract
Adversarial attacks expose the vulnerability of deep neural networks. Compared to image adversarial attacks, textual adversarial attacks are more challenging due to the discrete nature of texts. Recent synonym-based methods achieve the current state-of-the-art results. However, these methods introduce new words against the original text, leading to that humans easily perceive the difference between the adversarial example and the original text. Motivated by the fact that humans are usually unaware of chaotic word order in some cases, we propose exchange-attack (EA), a concise and effective word-level textual adversarial attack model. Specifically, the EA model generates adversarial examples by exchanging words of the original text itself according to the contributions that these words make regarding classification results. Intuitively, the smaller the distance between the two exchanged words, the more difficult the chaotic word order to be perceived by humans. We thus take the word distance into consideration when generating the chaotic word orders. Extensive experiments on several text classification data sets show that the EA model consistently outperforms the selected baselines in terms of averaged after-attack accuracy, modification rate, query number, and semantic similarity. And human evaluation results reveal that humans difficultly perceive the adversarial examples generated by the EA model. In addition, quantitative and qualitative analyses further validate the effectiveness of the EA model, including that the generated adversarial examples are grammatically correct and semantically preserved.
Huijun Liu 0003, Jie Yu 0008, Jun Ma 0015, Shasha Li 0001, Bin Ji 0002, Zibo Yi, Miaomiao Li 0001, Long Peng 0002, Xiaodong Liu 0004
Int. J. Intell. Syst.9
2021 DepOwl: Detecting Dependency Bugs to Prevent Compatibility Failures
abstract
Applications depend on libraries to avoid reinventing the wheel. Libraries may have incompatible changes during evolving. As a result, applications will suffer from compatibility failures. There has been much research on addressing detecting incompatible changes in libraries, or helping applications co-evolve with the libraries. The existing solution helps the latest application version work well against the latest library version as an afterthought. However, end users have already been suffering from the failures and have to wait for new versions. In this paper, we propose DepOwl, a practical tool helping users prevent compatibility failures. The key idea is to avoid using incompatible versions from the very beginning. We evaluated DepOwl on 38 known compatibility failures from StackOverflow, and DepOwl can prevent 35 of them. We also evaluated DepOwl using the software repository shipped with Ubuntu-19.10. DepOwl detected 77 unknown dependency bugs, which may lead to compatibility failures.
Zhouyang Jia, Shanshan Li 0001, Tingting Yu 0001, Erci Xu, Xiaodong Liu 0004, Ji Wang 0001, Xiangke Liao
ICSE6
2021 ConfInLog: Leveraging Software Logs to Infer Configuration Constraints
abstract
Misconfigurations have become the dominant causes of software failures in recent years, drawing tremendous attention for their increasing prevalence and severity. Configuration constraints can preemptively avoid misconfiguration by defining the conditions that configuration options should satisfy. Documentation is the main source of configuration constraints, but it might be incomplete or inconsistent with the source code. In this regard, prior researches have focused on obtaining configuration constraints from software source code through static analysis. However, the difficulty in pointer analysis and context comprehension prevents them from collecting accurate and comprehensive constraints. In this paper, we observed that software logs often contain configuration constraints. We conducted an empirical study and summarized patterns of configuration-related log messages. Guided by the study, we designed and implemented ConfInLog, a static tool to infer configuration constraints from log messages. ConfInLog first selects configuration-related log messages from source code by using the summarized patterns, then infers constraints from log messages based on the summarized natural language patterns. To evaluate the effectiveness of ConfInLog, we applied our tool on seven popular open-source software systems. ConfInLog successfully inferred 22~163 constraints, in which 59.5%~ 61.6% could not be inferred by the state-of-the-art work. Finally, we submitted 67 documentation patches regarding the constraints inferred by ConfInLog. The constraints in 29 patches have been confirmed by the developers, among which 10 patches have been accepted.
Shulin Zhou, Xiaodong Liu 0004, Shanshan Li 0001, Zhouyang Jia, Yuanliang Zhang, Teng Wang 0004, Wang Li 0003, Xiangke Liao
ICPC2
2019 Detecting Error-Handling Bugs without Error Specification Input
abstract
Most software systems frequently encounter errors when interacting with their environments. When errors occur, error-handling code must execute flawlessly to facilitate system recovery. Implementing correct error handling is repetitive but non-trivial, and developers often inadvertently introduce bugs into error-handling code. Existing tools require correct error specifications to detect error-handling bugs. Manually generating error specifications is error-prone and tedious, while automatically mining error specifications is hard to achieve a satisfying accuracy. In this paper, we propose EH-Miner, a novel and practical tool that can automatically detect error-handling bugs without the need for error specifications. Given a function, EH-Miner mines its error-handling rules when the function is frequently checked by an equivalent condition, and handled by the same action. We applied EH-Miner to 117 applications across 15 software domains. EH-Miner mined error-handling rules with the precision of 91.1% and the recall of 46.9%. We reported 142 bugs to developers, and 106 bugs had been confirmed and fixed at the time of writing. We further applied EH-Miner to Linux kernel, and reported 68 bugs for kernel-4.17, of which 42 had been confirmed or fixed.
Zhouyang Jia, Shanshan Li 0001, Tingting Yu 0001, Xiangke Liao, Ji Wang 0001, Xiaodong Liu 0004, Yunhuai Liu
ASE6
2018 Relax: Automatic Contention Detection and Resolution for Configuration Related Performance Tuning
abstract
As the scale and complexity of software expands, the issue of software performance is attracting increasing attention. The causes of performance problems mainly fall into two categories: software bugs and the resource contention among multiple software programs. Software bugs are usually caused by inefficient or unnecessary computation in source code. However, the performance problems caused by resource contention among multiple software programs are usually ignored by most researchers. Unlike software bugs, resource contention is not a bug; as a result, it is difficult to identify the concrete reason for a performance problem given that they share the same symptoms, such as long response time or low system throughput. In this paper, we investigate the performance problems caused by resource contention from a configuration perspective. By studying the response time distribution of software as the workload changes, we find that there is an inflection point of response time with the change of workload. Based on our observations, we design and implement a tool, Relax, to automatically detect and resolve resource contention. Relax combines resource request delay at the inflection point and the system resource usage rate to identify the performance problems caused by resource contention. Moreover, inspired by the congestion control algorithm in computer networks, Relax uses the square-increase and multiplicative-decrease method to adjust the resource-related configurations so as to resolve the contention. Our experiments show that Relax can effectively detect and resolve resource contention, and shorten the total software response time by 15.8% ~ 22.8%.
Zhimin Feng, Shanshan Li 0001, Xiangke Liao, Xiaodong Liu 0004, Shulin Zhou
APSEC4
2018 MisconfDoctor: Diagnosing Misconfiguration via Log-Based Configuration Testing
abstract
As software configurations continue to grow in complexity, misconfiguration has become one of major causes of software failure. Software configuration errors can have catastrophic consequences, seriously affecting the normal use of software and quality of service. And misconfiguration diagnosis faces many challenges, such as path-explosion problems and incomplete statistical data. Our study of the log that is generated in response to misconfigurations by six widely used pieces of software highlights some interesting characteristics. These observations have influenced the design of MisconfDoctor, a misconfiguration diagnosis tool via log-based configuration testing. Through comprehensive misconfiguration testing, MisconfDoctor first extracts log features for every misconfiguration and builds a feature database. When a system misconfiguration occurs, MisconfDoctor suggests potential misconfigurations by calculating the similarity of the new exception log to the feature database. We use manual and real-world error cases from Httpd, MySQL and PostgreSQL in order to evaluate the effectiveness of the tool. Experimental results demonstrate that the tool's accuracy reaches 85% when applied to manual-error cases, and 78% for real-world cases.
Teng Wang 0004, Xiaodong Liu 0004, Shanshan Li 0001, Xiangke Liao, Wang Li 0003, Qing Liao 0001
QRS2
2018 SMARTLOG: Place error log statement by deep understanding of log intention
abstract
Failure-diagnosis logs can dramatically reduce the system recovery time when software systems fail. Log automation tools can assist developers to write high quality log code. In traditional designs of log automation tools, they define log placement rules by extracting syntax features or summarizing code patterns. These approaches are, however, limited since the log placements are far beyond those rules but are according to the intention of software code. To overcome these limitations, we design and implement SmartLog, an intention-aware log automation tool. To describe the intention of log statements, we propose the Intention Description Model (IDM). SmartLog then explores the intention of existing logs and mines log rules from equivalent intentions. We conduct the experiments based on 6 real-world open-source projects. Experimental results show that SmartLog improves the accuracy of log placement by 43% and 16% compared with two state-of-the-art works. For 86 real-world patches aimed to add logs, 57% of them can be covered by SmartLog, while the overhead of all additional logs is less than 1%.
Zhouyang Jia, Shanshan Li 0001, Xiaodong Liu 0004, Xiangke Liao, Yunhuai Liu
SANER3
2018 Do You Really Know How to Configure Your Software? Configuration Constraints in Source Code May Help
abstract
Misconfigurations have become one of the major causes of software failures because of their increasing prevalence and severity. The complexity of configurations and users' lack of domain knowledge are the main reasons for massive misconfigurations. Users usually identify and diagnose misconfigurations by making a comparison against the conditions that configuration options should satisfy, which we refer to as configuration constraints; however, sometimes it is hard for users to accomplish this work. Some work has been done on obtaining configuration constraints, especially from source code; nevertheless, only part of the situation has been considered, such as if-statement code snippets, limiting its help in misconfiguration diagnosis. In order to better extract configuration constraints for users' guidance and misconfiguration diagnosis, we carried out a comprehensive manual study on the existence and variance of the configuration constraints in the source code of five different pieces of widely used open-source software. Three categories of findings are summarized based on our study, namely the general statistics, the general features of specific kinds of constraints, and the obstacles to the automatic extraction of configuration constraints. With these findings, we proposed several suggestions to maximize the automatic extraction of configuration constraints. The results show that our suggestions could improve the extraction of configuration constraints compared to existing methods.
Xiangke Liao, Shulin Zhou, Shanshan Li 0001, Zhouyang Jia, Xiaodong Liu 0004, Haochen He
IEEE Trans. Reliab.5
2017 Easier Said Than Done: Diagnosing Misconfiguration via Configuration Constraints Analysis: A Study of the Variance of Configuration Constraints in Source Code
abstract
Misconfigurations have drawn tremendous attention for their increasing prevalence and severity, and the main causes are the complexity of configurations as well as the lack of domain knowledge for software. To diagnose misconfigurations, one typical approach is to find out the conditions that configuration options should satisfy, which we refer to as configuration constraints. Current researches only handled part of the situations of configuration constraints in source code, which provide only limited help for misconfiguration diagnosis. To better extract configuration constraints, we conduct a comprehensive manual study on the existence and variance of the configuration constraints in source code from five pieces of popular open-source software. We summarized several findings from different aspects, including the general statistics about configuration constraints, the general features for specific configurations, and the obstacles in extraction of configuration constraints. Based on the findings, we propose several suggestions to maximize the automation of constraints extraction.
Shulin Zhou, Shanshan Li 0001, Xiaodong Liu 0004, Si Zheng 0003, Xiangke Liao, Yun Xiong
EASE3
2017 IdenEH: Identify error-handling code snippets in large-scale software
abstract
Error-handling (EH) code snippets are widely used for troubleshooting in software projects. Analyzing these snippets help to better understand how developers handle errors. However, the identification of such error-handling code snippets from the large-scale software is non-trivial, since traditional methods meet a challenge of scalability. In this paper, we analyze a large number of error-handling code snippets and get same interesting and useful observations. We extract seven features according to these observations. Based on these features, we design an automatic approach to identify error-handling codes using static program analysis and machine learning algorithms. Finally, we evaluate this approach and select the optimal feature subset from all feature combinations. Our evaluation demonstrates the high F-Score of up to 0.85 in identifying error-handling code snippets.
Shanshan Li 0001, Zhouyang Jia, Xiaodong Liu 0004, Bin Lin 0011, Xiangke Liao
ICCSA (7)4
2016 Towards Efficient Influence Maximization for Evolving Social Networks
Xiaodong Liu 0004, Xiangke Liao, Shanshan Li 0001, Bin Lin 0011
APWeb (1)1
2015 HeMatch: A redundancy layout placement scheme for erasure-coded storages in practical heterogeneous failure patterns
Shanshan Li 0001, Xiangke Liao, Shaoliang Peng, Xiaodong Liu 0004, Zhouyang Jia
Sci. China Inf. Sci.5
2014 PathZip: A lightweight scheme for tracing packet path in wireless sensor networks
Xiaopei Lu, Dezun Dong, Xiangke Liao, Shanshan Li 0001, Xiaodong Liu 0004
Comput. Networks5
2014 Leach: an automatic learning cache for inline primary deduplication system
Bin Lin 0011, Shanshan Li 0001, Xiangke Liao, Xiaodong Liu 0004
Frontiers Comput. Sci.5
2014 IMGPU: GPU-Accelerated Influence Maximization in Large-Scale Social Networks
abstract
Influence Maximization aims to find the top-$(K)$ influential individuals to maximize the influence spread within a social network, which remains an important yet challenging problem. Proven to be NP-hard, the influence maximization problem attracts tremendous studies. Though there exist basic greedy algorithms which may provide good approximation to optimal result, they mainly suffer from low computational efficiency and excessively long execution time, limiting the application to large-scale social networks. In this paper, we present IMGPU, a novel framework to accelerate the influence maximization by leveraging the parallel processing capability of graphics processing unit (GPU). We first improve the existing greedy algorithms and design a bottom-up traversal algorithm with GPU implementation, which contains inherent parallelism. To best fit the proposed influence maximization algorithm with the GPU architecture, we further develop an adaptive K-level combination method to maximize the parallelism and reorganize the influence graph to minimize the potential divergence. We carry out comprehensive experiments with both real-world and sythetic social network traces and demonstrate that with IMGPU framework, we are able to outperform the state-of-the-art influence maximization algorithm up to a factor of 60, and show potential to scale up to extraordinarily large-scale networks.
Xiaodong Liu 0004, Mo Li 0001, Shanshan Li 0001, Shaoliang Peng, Xiangke Liao, Xiaopei Lu
IEEE Trans. Parallel Distributed Syst.1
2014 Know by a handful the whole sack: efficient sampling for top-k influential user identification in large graphs
Xiaodong Liu 0004, Shanshan Li 0001, Xiangke Liao, Shaoliang Peng, Zhiyin Kong
World Wide Web1
2013 The architecture and traffic management of wireless collaborated hybrid data center network
abstract
This paper introduces a novel wireless collaborated hybrid data center architecture called RF-HYBRID that could optimize the effect of wireless transmission while reduce the complexity of wired network. RF-HYBRID improves throughput and packet delivery latency through flexible wireless detours and shortcuts, with a comprehensive routing and congestion control method.
Xiangke Liao, Shanshan Li 0001, Shaoliang Peng, Xiaodong Liu 0004, Bin Lin 0011
SIGCOMM5