EDBT 2026 Demo / reviewers in the wild / expert
Xiaohong Su
dblp:98/1257
· DBLP profile ↗
57ranked-venue papers
3as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 25 · 12 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Security and privacy · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-CheckingabstractXiaocheng Zhang, Xi Wang, Yifei Lu, Jianing Wang, Zhuangzhuang Ye, Mengjiao Bao, Peng Yan, Xiaohong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaocheng Zhang, Zhuangzhuang Ye, Mengjiao Bao, Xiaohong Su |
ACL (1) | 8 |
| 2026 | Shield Broken: Black-Box Adversarial Attacks on LLM-Based Vulnerability DetectorsabstractVulnerability detection is critical for ensuring software security. Although deep learning (DL) methods, particularly those employing large language models (LLMs), have shown strong performance in automating vulnerability identification, they remain susceptible to adversarial examples, which are carefully crafted inputs with subtle perturbations designed to evade detection. Existing adversarial attack methods often require access to model architectures or confidence scores, making them impractical for real-world black-box systems. In this paper, we propose SVulAttack, a novel label-only adversarial attack framework targeting LLM-based vulnerability detectors. Our key innovation lies in a similarity-based strategy that estimates statement importance and model confidence, thereby enabling more effective selection of semantic-preserving code perturbations. SVulAttack combines this strategy with a transformation component and a search component, based on either greedy or genetic algorithms, to effectively identify and apply optimal combinations of transformations. We evaluate SVulAttack on open-source models (LineVul, StagedVulBERT, Code Llama, Deepseek-Coder) and closed-source models (GPT-5 nano, GPT-4o, GPT-4o-mini, Claude Sonnet 4). Results show that SVulAttack significantly outperforms existing label-only black-box attack methods. For example, against LineVul, our method with genetic algorithm achieves an attack success rate of 49.0%, improving over DIP and CODA by 150.0% and 240.3%, respectively. Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2025 | Effective Code Membership Inference for Code Completion Models via Adversarial PromptsabstractMembership inference attacks (MIAs) on code completion models offer an effective way to assess privacy risks by inferring whether a given code snippet was part of the training data. Existing black- and gray-box MIAs rely on expensive surrogate models or manually crafted heuristic rules, which limit their ability to capture the nuanced memorization patterns exhibited by over-parameterized code language models. To address these challenges, we propose AdvPrompt-MIA, a method specifically designed for code completion models, combining code-specific adversarial perturbations with deep learning. The core novelty of our method lies in designing a series of adversarial prompts that induce variations in the victim code model’s output. By comparing these outputs with the ground-truth completion, we construct feature vectors to train a classifier that automatically distinguishes member from non-member samples. This design allows our method to capture richer memorization patterns and accurately infer training set membership. We conduct comprehensive evaluations on widely adopted models, such as Code Llama 7B, over the APPS and HumanEval benchmarks. The results show that our approach consistently outperforms state-of-the-art baselines, with AUC gains of up to 102%. In addition, our method exhibits strong transferability across different models and datasets, underscoring its practical utility and generalizability. Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
ASE | 5 |
| 2025 | Visual Modeling and Simulation of AUTOSAR Application Layer Models Using ModelicaabstractAs automotive electronic architectures grow increasingly complex and software development costs escalate, AUTOSAR plays a critical role in standardizing and enhancing the reusability of automotive controllers. However, existing AUTOSAR application layer modeling tools, such as the Simulink AUTOSAR Blockset, primarily adopt causal modeling paradigms, which constrain flexibility in capturing intricate system interactions. Additionally, their proprietary nature limits model accessibility, hindering cross-platform collaboration and multi-domain integration. Modelica, an object-oriented, equation-based modeling language, is particularly well-suited for multi-domain simulations due to its acausal modeling capabilities and strong support for component reuse. This paper proposes a Modelica-based visual modeling approach for AUTOSAR application layer models. Specifically, it establishes encapsulation rules for representing AUTOSAR constructs in Modelica and develops an open-source AUTOSAR model library, facilitating industry collaboration and accelerating rapid prototyping. A formal mathematical representation of AUTOSAR models is introduced to enhance both expressiveness and verifiability. Furthermore, a structured visual modeling methodology is presented to lower the development barrier for AUTOSAR application layer modeling. Comparative analysis with Simulink's AUTOSAR Blockset demonstrates that the proposed approach successfully integrates Modelica's multi-domain modeling capabilities into the AUTOSAR workflow while ensuring simulation consistency with Simulink. To the best of our knowledge, this work represents the first integration of AUTOSAR within Modelica's multi-domain simulation framework. Compared to Simulink, Modelica's acausal modeling paradigm enables more flexible system representations, its open ecosystem supports cross-platform collaboration, and its multi-domain integration enhances interoperability. Beyond the automotive domain, the proposed approach can also be applied to controller design in other industries, further demonstrating its potential for cross-disciplinary adoption. Peihao Yang, Tiantian Wang 0001, Xiaohong Su |
MODELS | 4 |
| 2025 | Knowledge-guided large language models are trustworthy API recommenders
Hongwei Wei, Xiaohong Su, Weining Zheng, Wenxing Tao, Yuqian Kuang |
Autom. Softw. Eng. | 2 |
| 2025 | VDExplainer: Sequential decision-making and probability sampling guided statement-level explanation for vulnerability detection
Weining Zheng, Xiaohong Su, Hongwei Wei, Wenxin Tao |
Comput. Secur. | 2 |
| 2025 | Transformer-based statement level vulnerability detection by cross-modal fine-grained features capture
Wenxin Tao, Xiaohong Su, Yekun Ke, Hongwei Wei |
Knowl. Based Syst. | 2 |
| 2025 | Enhancing Fine-Grained Vulnerability Detection With Reinforcement LearningabstractThe rapid growth of vulnerabilities has significantly accelerated the development of automated vulnerability detection methods, especially those based on data-driven models. However, most of them primarily focus on extracting accurate code representations while overlooking the complex vulnerability patterns among vulnerable statements, thereby leaving room for improvement. To overcome this limitation, we present a novel reinforcement learning framework (RLFD) for detecting vulnerabilities at a fine-grained level.RLFDredefines the detection task as a sequential decision-making process and then employs reinforcement learning to automatically learn vulnerability-relevant structures from code snippets. Moreover, by designing reward functions aligned with fine-grained evaluation metrics,RLFDfocuses on the co-existence relations among statements from a global perspective, enabling the model to capture complex interactions that lead to vulnerabilities. Additionally, the framework utilizes CodeBERT-HLS for code representation, ensuring consistency with the state-of-the-art method while highlighting the improvements brought by the proposed reinforcement learning-based approach. Comprehensive experiments show that our method achieves a locating precision (IoU) of 69.7% and a Top-5% Acc of 67.7% on thebig_vuldataset, outperforming the state-of-the-art method by an overall 3.4% improvement in IoU. Notably, our method achieves up to a 19.7% increase in IoU for specific categories, e.g., CWE-416 (use-after-free). Zhichen Qu, Christoph Treude, Xiaohong Su, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2024 | LI4: Label-Infused Iterative Information Interacting Based Fact Verification in Question-answering DialogueabstractFact verification constitutes a pivotal application in the effort to combat the dissemination of disinformation, a concern that has recently garnered considerable attention. However, previous studies in the field of fact verification, particularly those focused on question-answering dialogue, have exhibited limitations, such as failing to fully exploit the potential of question structures and ignoring relevant label information during the verification process. In this paper, we introduce Label-Infused Iterative Information Interacting (LI4), a novel approach designed for the task of question-answering dialogue based fact verification. LI4 consists of two meticulously designed components, namely the Iterative Information Refining and Filtering Module (IIRF) and the Fact Label Embedding Module (FLEM). The IIRF uses the Interactive Gating Mechanism to iteratively filter out the noise of question and evidence, concurrently refining the claim information. The FLEM is conceived to strengthen the understanding ability of the model towards labels by injecting label knowledge. We evaluate the performance of the proposed LI4 on HEALTHVER, FAVIQ, and COLLOQUIAL. The experimental results confirm that our LI4 model attains remarkable progress, manifesting as a new state-of-the-art performance. Xiaocheng Zhang, Guoping Zhao, Xiaohong Su |
LREC/COLING | 4 |
| 2024 | SVulDetector: Vulnerability detection based on similarity using tree-based attention and weighted graph embedding mechanisms
Weining Zheng, Xiaohong Su, Hongwei Wei, Wenxin Tao |
Comput. Secur. | 2 |
| 2024 | Joint reconstruction and deidentification for mobile identity anonymization
Hyeongbok Kim, Lingling Zhao, Zhiqi Pang, Xiaohong Su, Jin Suk Lee |
Multim. Tools Appl. | 4 |
| 2024 | StagedVulBERT: Multigranular Vulnerability Detection With a Novel Pretrained Code ModelabstractThe emergence of pre-trained model-based vulnerability detection methods has significantly advanced the field of automated vulnerability detection. However, these methods still face several challenges, such as difficulty in learning effective feature representations of statements for fine-grained predictions and struggling to process overly long code sequences. To address these issues, this study introduces StagedVulBERT, a novel vulnerability detection framework that leverages a pre-trained code language model and employs a coarse-to-fine strategy. The key innovation and contribution of our research lies in the development of the CodeBERT-HLS component within our framework, specialized in hierarchical, layered, and semantic encoding. This component is designed to capture semantics at both the token and statement levels simultaneously, which is crucial for achieving more accurate multi-granular vulnerability detection. Additionally, CodeBERT-HLS efficiently processes longer code token sequences, making it more suited to real-world vulnerability detection. Comprehensive experiments demonstrate that our method enhances the performance of vulnerability detection at both coarse- and fine-grained levels. Specifically, in coarse-grained vulnerability detection, StagedVulBERT achieves an F1 score of 92.26%, marking a 6.58% improvement over the best-performing methods. At the fine-grained level, our method achieves a Top-5% accuracy of 65.69%, which outperforms the state-of-the-art methods by up to 75.17%. Yujian Zhang, Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2023 | A Graph Neural Network-Based Smart Contract Vulnerability Detection Method with Artificial Rule
Ziyue Wei, Weining Zheng, Xiaohong Su, Wenxin Tao, Tiantian Wang 0001 |
ICANN (4) | 3 |
| 2023 | Documentation-Guided API Sequence Search without Worrying about the Text-API Semantic GapabstractDevelopers often search for application programming interfaces (APIs) and their usage patterns to speed up the efficiency of software development. This paper focuses on the API sequence search task, which refers to using a function-relevant textual query to search for API sequences mined from open-source software repositories that can implement this function. However, the severe semantic gap between text and API makes it challenging to discover the correspondence between natural language queries and desired API sequences. Therefore, we propose a method called documentation-guided API sequence search (DGAS), through which we do not need to worry about the semantic gap between text and API. Specifically, DGAS consists of documentation-guided cross-modal attention (DGCA) and documentation-guided cross-modal matching (DGCM). DGCA calculates the cross-modal attention map using features extracted from the same modality (i.e., API documentation sequence and textual query) instead of from different modalities (i.e., API sequence and textual query) to bridge the semantic gap during the cross-modal attention phase. Besides, DGCM takes API documentation as supplementary information of API sequence to bridge the semantic gap during the cross-modal matching phase. We use the API documentation to extend the existing dataset for API sequence generation to construct a dataset for API sequence search to evaluate DGAS. Experimental results show that DGAS outperforms the baseline methods. Hongwei Wei, Xiaohong Su, Weining Zheng, Wenxin Tao |
SANER | 2 |
| 2023 | Vulnerability detection through cross-modal feature enhancement and fusion
Wenxin Tao, Xiaohong Su, Jiayuan Wan, Hongwei Wei, Weining Zheng |
Comput. Secur. | 2 |
| 2023 | Does Deep Learning improve the performance of duplicate bug report detection? An empirical study
Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
J. Syst. Softw. | 2 |
| 2023 | Semantic-aware deidentification generative adversarial networks for identity anonymizationabstractAbstract Privacy protection in the computer vision field has attracted increasing attention. Generative adversarial network-based methods have been explored for identity anonymization, but they do not take into consideration semantic information of images, which may result in unrealistic or flawed facial results. In this paper, we propose a Semantic-aware De-identification Generative Adversarial Network (SDGAN) model for identity anonymization. To retain the facial expression effectively, we extract the facial semantic image using the edge-aware graph representation network to constraint the position, shape and relationship of generated facial key features. Then the semantic image is injected into the generator together with the randomly selected identity information for de-Identification. To ensure the generation quality and realistic-looking results, we adopt the SPADE architecture to improve the generation ability of conditional GAN. Meanwhile, we design a hybrid identity discriminator composed of an image quality analysis module, a VGG-based perceptual loss function, and a contrastive identity loss to enhance both the generation quality and ID anonymization. A comparison with the state-of-the-art baselines demonstrates that our model achieves significantly improved de-identification (De-ID) performance and provides more reliable and realistic-looking generated faces. Our code and data are available on https://github.com/kimhyeongbok/SDGAN Hyeongbok Kim, Zhiqi Pang, Lingling Zhao, Xiaohong Su, Jin Suk Lee |
Multim. Tools Appl. | 4 |
| 2023 | A Hypothesis Testing-based Framework for Software Cross-modal Retrieval in Heterogeneous Semantic SpacesabstractSoftware cross-modal retrieval is a popular yet challenging direction, such as bug localization and code search. Previous studies generally map natural language texts and codes into a homogeneous semantic space for similarity measurement. However, it is not easy to accurately capture their similar semantics in a homogeneous semantic space due to the semantic gap. Therefore, we propose to map the multi-modal data into heterogeneous semantic spaces to capture their unique semantics. Specifically, we propose a novel software cross-modal retrieval framework named Deep Hypothesis Testing (DeepHT). In DeepHT, to capture the unique semantics of the code’s control flow structure, all control flow paths (CFPs) in the control flow graph are mapped to a CFP sample set in the sample space. Meanwhile, the text is mapped to a CFP correlation distribution in the distribution space to model its correlation with different CFPs. The matching score is calculated according to how well the sample set obeys the distribution using hypothesis testing. The experimental results on two text-to-code retrieval tasks (i.e., bug localization and code search) and two code-to-text retrieval tasks (i.e., vulnerability knowledge retrieval and historical patch retrieval) show that DeepHT outperforms the baseline methods. Hongwei Wei, Xiaohong Su, Weining Zheng, Wenxin Tao |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | Dynamically Relative Position Encoding-Based Transformer for Automatic Code EditabstractAdapting deep learning (DL) techniques to automate nontrivial coding activities, such as code documentation and defect detection, has been intensively studied recently. Learning to predict code changes is one of the popular and essential investigations. Prior studies have shown that DL techniques, such as neural machine translation (NMT), can benefit meaningful code changes, including bug fixing and code refactoring. However, NMT models may encounter bottleneck when modeling long sequences; thus, they are limited in accurately predicting code changes. In this article, we design a Transformer-based approach, considering that the Transformer has proven effective in capturing long-term dependencies. Specifically, we propose a novel model named DTrans. For better incorporating the local structure of code, i.e., statement-level information in this article, DTrans is designed with dynamically relative position encoding in the multihead attention of the Transformer. Experiments on benchmark datasets demonstrate that DTrans can more accurately generate patches than the state-of-the-art methods, increasing the performance by at least 5.45–46.57% in terms of the exact match metric on different datasets. Moreover, DTrans can locate the lines to change with 1.75–24.21% higher accuracy than the existing methods. Shiyi Qi, Cuiyun Gao 0001, Xiaohong Su, Shuzheng Gao, Zibin Zheng, Chuanyi Liu |
IEEE Trans. Reliab. | 4 |
| 2022 | Fault localization based on wide & deep learning model by mining software behavior
Tiantian Wang 0001, HaiLong Yu, Kechao Wang, Xiaohong Su |
Future Gener. Comput. Syst. | 4 |
| 2022 | Golden Mutator Recommendation Based on Mutation Pattern MiningabstractMutation testing is widely used in the research of evaluation and optimization of test set quality, and has been paid attention to the study of bug localization and fixing. But one inherent problem of mutation testing is huge computation cost. Selective mutation is an important method to reduce mutation testing cost. However, the existing selective mutation researches indicate that there is no universal selection strategy. This paper proposes a method which integrates with historical software data mining to recommend suitable mutator for program under test. The basic idea of the method is in the software version control system, any pair of nonfixed and post-fixed programs can be original program and mutant of each other. The edit performed during bug fixing contains mutators, and such mutator can guide buggy program to correct program, hence it is called Golden Mutator. The basic method is to compare the buggy files and the corresponding fixed files in the version control system, obtain historical faulty statements and their fix editing operations, thereby accumulating the bug-fix instance base, and then mining mutation patterns from it. Before mutating the target statement, we first find out faulty instances similar to the target statement, and use their fixing edits to match with the mutation pattern, so as to get the golden mutator applicable to the target statement. This paper uses Defects4j dataset in the test, verifies the accuracy of the proposed recommendation method, and further uses the method in bug localization. Compared to fixedly selected mutators, when applying the golden mutator recommended by using the proposed method, the average accuracy of bug localization is higher. Dan Gong, Tiantian Wang 0001, Xiaohong Su |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2022 | Equivalent Mutants Detection Based on Weighted Software Behavior GraphabstractThe equivalent mutants problem is one of the crucial problems in mutation testing. In consequence of its existence, the effectiveness of mutation testing is underestimated. In addition, it will produce a certain amount of useless overhead. Equivalent mutants cannot be detected by any test input. The existing works mostly focus on static analysis to detect, or avoid generating, the equivalent mutants. The essence of these methods is to use prior knowledge to establish some rules of program equivalence. However, (1) it needs a lot of professional labor to sort out the equivalence rules, and (2) only a small part of the rules can be determined in advance, because of the diversity of mutation operators and mutation targets. Consequently, the best result reported so far is 50% of the equivalent mutants can be detected. Since it is generally believed that manual judgment of program equivalence is the most reliable, this paper proposes a novel method to automatically detect equivalent mutants by tracing program behavior like the professionals. The weighted software behavior graph is utilized in the detection of equivalent mutants for the first time. This method can not only figure out different execution paths, but also be sensitive to execution frequency. By comparing the weighted software behavior graphs of an alive mutant and its original program, we are able to examine more precisely whether the alive mutant is the same as the original program, in terms of the state of infection and/or the propagation. Evaluation results on an open dataset of manually evaluated equivalent mutants show that our approach can detect 77.5% of all the equivalent mutants, which is much higher than the existing static methods. Dan Gong, Tiantian Wang 0001, Xiaohong Su, Yanhang Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2022 | Hierarchical semantic-aware neural code representation
Xiaohong Su, Christoph Treude, Tiantian Wang 0001 |
J. Syst. Softw. | 2 |
| 2021 | Vu1SPG: Vulnerability detection based on slice property graph representation learningabstractVulnerability detection is an important issue in software security. Although various data-driven vulnerability detection methods have been proposed, the task remains challenging since the diversity and complexity of real-world vulnerable code in syntax and semantics make it difficult to extract vulnerable features with regular deep learning models, especially in analyzing a large program. Moreover, the fact that real-world vulnerable codes contain a lot of redundant information unrelated to vulnerabilities will further aggravate the above problem. To mitigate such challenges, we define a novel code representation named Slice Property Graph (SPG), and then propose VulSPG, a new vulnerability detection approach using the improved R-GCN model with triple attention mechanism to identify potential vulnerabilities in SPG. Our approach has at least two advantages over other methods. First, our proposed SPG can reflect the rich semantics and explicit structural information that may be relevance to vulnerabilities, while eliminating as much irrelevant information as possible to reduce the complexity of graph. Second, VulSPG incorporates triple attention mechanism in R-GCNs to achieve more effective learning of vulnerability patterns from SPG. We have extensively evaluated VulSPG on two large-scale datasets with programs from SARD and real-world projects. Experimental results prove the effectiveness and efficiency of VulSPG. Weining Zheng, Xiaohong Su |
ISSRE | 3 |
| 2020 | LTRWES: A new framework for security bug report detection
Pengcheng Lu, Xiaohong Su, Tiantian Wang 0001 |
Inf. Softw. Technol. | 3 |
| 2019 | Invariant based fault localization by analyzing error propagation
Tiantian Wang 0001, Kechao Wang, Xiaohong Su, Lei Zhang 0036 |
Future Gener. Comput. Syst. | 3 |
| 2019 | Automatic debugging of operator errors based on efficient mutation analysis
Tiantian Wang 0001, Jiahuan Xu, Xiaohong Su, ChenShi Li, Yang Chi |
Multim. Tools Appl. | 3 |
| 2018 | The Delta Generalized Labeled Multi-Bernoulli Filter for Cell TrackingabstractCell tracking automatically in time-lapse image sequences is important for understanding the dynamic pattern of micro-cell. In this paper, we present a novel method for tracking cell with shape feature based on the delta generalized labeled multi-Bernoulli (delta-GLMB) filter which is of great research significance. The delta-GLMB filter with cell shape parameters can improve the tracking accuracy. This approach is evaluated and compared with raw detection using the generalized optimal sub-pattern assignment (GOSPA) metric on real N2DH-SIM cell sequences. Experiment results show that the delta-GLMB filter can provide the shape information as well as the better estimation than raw detection and KTH method. Chunmei Shi, Junjie Wang 0005, Lingling Zhao, Xiaohong Su, Guangshun Jiang |
BIBE | 4 |
| 2018 | An Empirical Study on Software Defect Prediction Using Over-Sampling by SMOTEabstractSoftware defect prediction suffers from the class-imbalance. Solving the class-imbalance is more important for improving the prediction performance. SMOTE is a useful over-sampling method which solves the class-imbalance. In this paper, we study about some problems that faced in software defect prediction using SMOTE algorithm. We perform experiments for investigating how they, the percentage of appended minority class and the number of nearest neighbors, influence the prediction performance, and compare the performance of classifiers. We use paired t-test to test the statistical significance of results. Also, we introduce the effectiveness and ineffectiveness of over-sampling, and evaluation criteria for evaluating if an over-sampling is effective or not. We use those concepts to evaluate the results in accordance with the evaluation criteria for the effectiveness of over-sampling. The results show that they, the percentage of appended minority class and the number of nearest neighbors, influence the prediction performance, and show that the over-sampling by SMOTE is effective in several classifiers. CholMyong Pak, Tiantian Wang 0001, Xiaohong Su |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2017 | Interactive WCET Prediction with Warning for Timeout RiskabstractWorst case execution time (WCET) analysis is essential for exposing timeliness defects when developing hard real-time systems. However, it is too late to fix timeliness defects cheaply since developers generally perform WCET analysis in a final verification phase. To help developers quickly identify real timeliness defects in an early programming phase, a novel interactive WCET prediction with warning for timeout risk is proposed. The novelty is that the approach not only fast estimates WCET based on a control flow tree (CFT), but also assesses the estimated WCET with a trusted level by a lightweight false path analysis. According to the trusted levels, corresponding warnings will be triggered once the estimated WCET exceeds a preset safe threshold. Hence developers can identify real timeliness defects more timely and efficiently. To this end, we first analyze the reasons of the overestimation of CFT-based WCET calculation; then we propose a trusted level model of timeout risks; for recognizing the structural patterns of timeout risks, we develop a risk data counting algorithm; and we also give some tactics for applying our approach more effectively. Experimental results show that our approach has almost the same running speed compared with the fast and interactive WCET analysis, but it saves more time in identifying real timeliness defects. Fanqi Meng, Xiaohong Su, Zhaoyang Qu |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2017 | Predicting change consistency in a clone group
Fanlong Zhang, Siau-Cheng Khoo, Xiaohong Su |
J. Syst. Softw. | 3 |
| 2017 | Greening Software Requirements Change Management Strategy Based on Nash EquilibriumabstractRecently, green computing has become more and more important in software engineering (SE), which can be achieved by effectively recycling the software system and utilizing the computing resources. However, the requirement change may lead to unnecessary labor and time cost. Moreover, it may also result in the waste of hardware and computing resources once unreasonable requirements are realized. Thus, to perform green computing in SE, it is necessary to propose effective strategies to manage the requirement change. For this decision-making problem, game theoretical methods can be feasible solutions. In this paper, we propose a novel requirement change management approach based on game theory. Specifically, we model the problem as a game between the stakeholders and the developer and devise the payoff matrix between different strategies of the players. We then propose a Nash equilibrium-based game theoretical algorithm to manage requirement change. The evaluation results show that, compared to the exhaustive algorithm, our method not only can achieve almost the same optimal results but also can significantly reduce the computational time complexity. Thus, our method is feasible for a lot of requirement changes and can facilitate the green computing targets from the perspective of software engineering. Zhixiang Tong, Xiaohong Su, Longzhu Cen, Tiantian Wang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2016 | Optimization and improvements of a Moodle-Based online learning system for C programmingabstractIt is important for students to solve problems with specific requirements in the programming teaching. Our teaching system is a Moodle-based interactive teaching platform for C programming. Its online judging system can grade students code automatically. It plays an extremely important role in programming language teaching. This paper is devoted to optimizing and improving the system. We firstly analyze the five problems in the system according to the feedback from the teachers and students: 1) logical errors in students programs cannot be located; 2) a cheating that directly outputs answers cannot be detected; 3) evaluation results lack statistics and visualization; 4) the code submitting procedure is complicated; 5) the feedback of incorrect answers is not detailed. In order to solve these problems, we employ a fault localization algorithm, revise the evaluation logic, introduce third-party visualization plug-ins and refactor the system, respectively. Detailed and exact solutions are given also. After optimizing and improving the system, the user experience is significantly improved. It is convenient for the student to find and correct errors in their programs. Also, it is easier for teachers to acquire valuable feedback and master students' learning situation. Xiaohong Su, Jing Qiu 0003, Tiantian Wang 0001, Lingling Zhao |
FIE | 1 |
| 2016 | Predicting Consistent Clone ChangeabstractCode clones, being an inevitable by-product of rapid software development, can impact software quality. The introduction of code clone groups and clone genealogies enable software developers to be aware of the presence of and changes to clones as a collective group, they also allow developers to understand how clone groups evolve throughout software life cycle. Due to similarity in codes within a clone group, a change in one piece of the code may require developers to make changes to other clones in the group. Failure in making consistent change to a clone group when necessary is commonly known as "clone consistency-defect", which can adversely impact software reliability. We propose an approach to predict clone consistency-requirement at the time when changes have been made to a clone group. Our predictor is a Bayesian network implemented in WEKA. We build a variant of clone genealogies to collect all consistent/inconsistent changes to clone groups, and extract three sets of attributes from clone groups as input for predicting consistent clone change. These three sets are: code attributes, context attributes and evolution attributes. We conduct experiments on three open source projects. These experiments show that our approach has high precision and recall in predicting clone consistency-requirement. This holistic approach can aid developers in maintaining code clone changes, and avoid potential clone consistency-defect, which can improve the software quality and reliability. Fanlong Zhang, Siau-Cheng Khoo, Xiaohong Su |
ISSRE | 3 |
| 2016 | A Novel Multi Stage Cooperative Path Re-planning Method for Multi UAV
Xiaohong Su, Lingling Zhao, Yan-hang Zhang |
PRICAI | 1 |
| 2016 | RaceTracker: Effective and efficient detection of data racesabstractData races are common concurrency bugs in multi-threaded programs and many of them can cause server failures. They are difficult to be detected or verified due to some non-deterministic interleavings. Numerous static and dynamic program analysis techniques have been proposed to detect data races. However, some of detectors may report large amount of false races and some of them may miss lots of true races. This paper proposes RaceTracker that combines static and dynamic techniques to detect data races effectively and efficiently. First, we use current static detectors to produce potential races and identify ad-hoc synchronizations. Second, we instrument code locations corresponding to potential races to try best to expose them in controlled thread interleavings. Meanwhile, we apply the hybrid dynamic detection techniques to detect races in case they are exposed. Finally, we also apply the hybrid dynamic detection techniques to prune benign and false races mainly caused by ad-hoc synchronization. In order to increase the chances to trigger real race conditions in dynamic analysis, we use some strategies to control the schedule of threads. We have implemented our tool RaceTracker in the dynamic binary instrumentation framework PIN and evaluated with nearly 100 small data race programs from google data-race-test suit and some real-world concurrent applications from SPLASH-2 and Maple. Evaluations show that RaceTracker can identify more data races effectively compared with prior pure dynamic race verifiers and detectors. Meanwhile, compared with grouping verifier, RaceTracker only executes the program twice and reduces the average runtime overhead by 64%. Xiaohong Su, Peijun Ma |
SNPD | 3 |
| 2016 | Using Reduced Execution Flow Graph to Identify Library Functions in Binary CodeabstractDiscontinuity and polymorphism of a library function create two challenges for library function identification, which is a key technique in reverse engineering. A new hybrid representation of dependence graph and control flow graph called Execution Flow Graph (EFG) is introduced to describe the semantics of binary code. Library function identification turns to be a subgraph isomorphism testing problem since the EFG of a library function instance is isomorphic to the sub-EFG of this library function. Subgraph isomorphism detection is time-consuming. Thus, we introduce a new representation called Reduced Execution Flow Graph (REFG) based on EFG to speed up the isomorphism testing. We have proved that EFGs are subgraph isomorphic as long as their corresponding REFGs are subgraph isomorphic. The high efficiency of the REFG approach in subgraph isomorphism detection comes from fewer nodes and edges in REFGs and new lossless filters for excluding the unmatched subgraphs before detection. Experimental results show that precisions of both the EFG and REFG approaches are higher than the state-of-the-art tool and the REFG approach sharply decreases the processing time of the EFG approach with consistent precision and recall. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
IEEE Trans. Software Eng. | 2 |
| 2015 | Identifying and Understanding Self-Checksumming Defenses in SoftwareabstractSoftware self-checksumming is widely used as an anti-tampering mechanism for protecting intellectual property and deterring piracy. This makes it important to understand the strengths and weaknesses of various approaches to self-checksumming. This paper describes a dynamic information-flow-based attack that aims to identify and understand self-checksumming behavior in software. Our approach is applicable to a wide class of self chesumming defenses and the information obtained can be used to determine how the checksumming defenses may be bypassed. Experiments using a prototype implementation of our ideas indicate that our approach can successfully identify self-checksumming behavior in (our implementations of) proposals from the research literature. Jing Qiu 0003, Babak Yadegari, Brian Johannesmeyer, Saumya K. Debray, Xiaohong Su |
CODASPY | 5 |
| 2015 | Motivating students with new mechanisms of online assignments and examination to meet the MOOC challenges for programmingabstractThe advent of massive open online courses (MOOC) poses challenges for teaching and learning programming. This paper has analyzed these challenges and thereby proposed a self-motivating learning platform for students in the introductory programming course. Novel mechanisms of online assignments and examination have been introduced. Our platform provides functions for self-motivating learning and practicing in MOOC, which makes it distinguish from the others. For example, self-paced timetable with supervision, self-motivated exercise contents, exercise market, and relative ranking. The automatic grading approach is also a highlight. Programs even with syntactic or semantic errors can be automatically graded. Our platform gains popularity among both students and teachers. The platform has been used together with a programming MOOC. This course is ranked as the third most popular courses among over 500 courses. The platform has also been used by more than 100 other universities. The application of the platform in both MOOC and the traditional classroom courses has shown that students' self-motivation in learning programming has been greatly promoted, and their practical skills have also been significantly improved. Xiaohong Su, Tiantian Wang 0001, Jing Qiu 0003, Lingling Zhao |
FIE | 1 |
| 2015 | Interest-driven and innovation-oriented practice for programming courseabstractIn order to maximize the motivation of students in the programming practice, this paper offers an analysis on the core factors of practice case motivating students put in effort in programming practice, namely, "interest", "usability", and "hierarchy". Furthermore, we present typical practice cases which are carefully designed according to the motivating factors and give a description on the implementation and experience of our programming practice course at Harbin Institute of Technology. The designed programming practice can not only train the students' practical programming skills but also enhance their self-regulated learning skills, creativity and self-efficacy. Lingling Zhao, Xiaohong Su, Tiantian Wang 0001, Yongfeng Yuan |
FIE | 2 |
| 2015 | Nonparametric background model based clutter map for X-band marine radarabstractIn a radar system, clutter means any echoes which are not scattered by the wanted target. Usually, the radar clutter map stores an average level for each point or cell in the range-azimuth coordinates as the reference value. A target is then detected in a range-azimuth region if the new echo value in that region exceeds the average background level. In visual computing domain, background model is employed for foreground segmentation, motion detection, salient feature detection, etc. A few kinds of background models are built on each pixel of the sequential images, including the recently popular non-parametric model. In this paper, we proposes a non-parametric background model to implement the radar clutter map. A set of intensity values, which are selected in the past in each pixel location, are stored as the initial clutter model. Then, the new value of a pixel is classified as foreground stemming from a moving target, if the value was stronger than those of the reference samples in the recorded set Clutter model updating is based on randomly choosing in temporal and substituting background pixel values in spatial. Proposed model is proved to be efficient and effective in a moving vehicle detecting application. Also it is compared to a broadly used clutter map method which employs averaging in temporal. Yi Zhou 0011, Jidong Suo, Xiaohong Su, Limei Liu |
ICIP | 5 |
| 2015 | Library functions identification in binary code by using graph isomorphism testingsabstractLibrary functions identification is a key technique in reverse engineering. Discontinuity and polymorphism of inline and optimized library functions in binary code create a difficult challenge for library functions identification. To solve this problem, a novel approach is developed to identify library functions. First, we introduce execution dependence graphs (EDGs) to describe the behavior characteristics of binary code. Then, by finding similar EDG subgraphs in target functions, we identify both full and inline library functions. Experimental results from the prototype tool show that the proposed method is not only capable of identifying inline functions but is also more efficient and precise than the current methods for identifying full library functions. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
SANER | 2 |
| 2015 | State dependency probabilistic model for fault localization
Dandan Gong, Xiaohong Su, Tiantian Wang 0001, Peijun Ma |
Inf. Softw. Technol. | 2 |
| 2015 | Identifying functions in binary code with reverse extended control flow graphsabstractAbstract In binary code analysis, current function identification approaches are challenged by functions without explicit call sites and handcrafted assembly without standard prologues/epilogues. We propose a new function representation called a reverse extended control flow graph (RECFG) and a RECFG‐based method for identifying functions in stripped binary code. A function has at least one return instruction (an instruction that makes the control flow leave a function). Therefore, return instructions are more reliable than the function prologues and epilogues used by traditional methods. We first build RECFGs from any values that can be interpreted as return instructions in a code range. Then, for each independent RECFG, the multiple‐decision method chooses a subgraph as the control flow graph of a function. A prototype tool is developed for evaluation on seven open source applications, 138 binaries in MASM32 code examples, and 292 binaries in Windows XP SP3. Experimental results show that the proposed method can identify functions that cannot be identified by current methods with high precision and stable recall. Copyright © 2015 John Wiley & Sons, Ltd. Jing Qiu 0003, Xiaohong Su, Peijun Ma |
J. Softw. Evol. Process. | 2 |
| 2014 | Multi-object Tracking Based on Particle Probability Hypothesis Density Tracker in Microscopic VideoabstractResearch on biological objects requires tracking hundreds of micro-objects from the microscopy video. We propose an automated tracking framework to extract trajectories of micro-objects. This framework uses a particle probability hypothesis density (PF-PHD) tracker to implement a recursive Bayesian state estimation and trajectories association. In the framework, an ellipse target model is presented to describe the micro-objects with shape parameters instead of point-like targets. Furthermore, an orientation and positional constraint model is developed to deal with the data association of crossing trajectories in multitarget tracking. Using this framework, a significantly larger number of tracks are obtained than manual tracking. The experiments on simulated image sequences of microtubule movement are performed in order to evaluate the proposed PF-PHD tracking method. Chunmei Shi, Lingling Zhao, Peijun Ma, Xiaohong Su, Junjie Wang 0005, Chiping Zhang |
BIBE | 4 |
| 2014 | Detection of semantically similar code
Tiantian Wang 0001, Kechao Wang, Xiaohong Su, Peijun Ma |
Frontiers Comput. Sci. | 3 |
| 2014 | Corrigendum to: "SPAPE: A semantic-preserving amorphous procedure extraction method for near-miss clones": [J. Syst. Softw. 86 (2013) 2077-2093]
Yixin Bian, Akif Günes Koru, Xiaohong Su, Peijun Ma |
J. Syst. Softw. | 3 |
| 2013 | A test-suite reduction approach to improving fault-localization effectiveness
Dandan Gong, Tiantian Wang 0001, Xiaohong Su, Peijun Ma |
Comput. Lang. Syst. Struct. | 3 |
| 2013 | SPAPE: A semantic-preserving amorphous procedure extraction method for near-miss clones
Yixin Bian, Akif Günes Koru, Xiaohong Su, Peijun Ma |
J. Syst. Softw. | 3 |
| 2013 | Robust feature selection based on regularized brownboost loss
Qinghua Hu, Peijun Ma, Xiaohong Su |
Knowl. Based Syst. | 4 |
| 2012 | A STPHD-Based Multi-sensor Fusion Method
Zhenwei Lu, Lingling Zhao, Xiaohong Su, Peijun Ma |
ICONIP (3) | 3 |
| 2012 | Improvements of a two-in-one image secret sharing scheme based on gray mixing model
Peng Li 0050, Peijun Ma, Xiaohong Su, Ching-Nung Yang |
J. Vis. Commun. Image Represent. | 3 |
| 2011 | Entropy on interval-valued intuitionistic fuzzy sets and its application in multi-attribute decision making
Peijun Ma, Xiaohong Su, Chiping Zhang |
FUSION | 3 |
| 2011 | Selective Track Fusion
Peijun Ma, Xiaohong Su |
ICONIP (3) | 3 |
| 2010 | A new multi-target state estimation algorithm for PHD particle filter
Lingling Zhao, Peijun Ma, Xiaohong Su |
FUSION | 3 |
| 2007 | Distributed Learning Strategy Based on Chips for Classification with Large-Scale DatasetabstractLearning with very large-scale datasets is always necessary when handling real problems using artificial neural networks. However, it is still an open question how to balance computing efficiency and learning stability, when traditional neural networks spend a large amount of running time and memory to solve a problem with large-scale learning dataset. In this paper, we report the first evaluation of neural network distributed-learning strategies in large-scale classification over protein secondary structure. Our accomplishments include: (1) an architecture analysis on distributed-learning, (2) the development of scalable distributed system for large-scale dataset classification, (3) the description of a novel distributed-learning strategy based on chips, (4) a theoretical analysis of distributed-learning strategies for structure-distributed and data-distributed, (5) an investigation and experimental evaluation of distributed-learning strategy based-on chips with respect to time complexity and their effect on the classification accuracy of artificial neural networks. It is demonstrated that the novel distributed-learning strategy is better-balanced in parallel computing efficiency and stability as compared with the previous algorithms. The application of the protein secondary structure prediction demonstrates that this method is feasible and effective in practical applications. Bo Yang 0013, Xiaohong Su |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | Semantic similarity-based grading of student programs
Tiantian Wang 0001, Xiaohong Su, Peijun Ma |
Inf. Softw. Technol. | 2 |