Hailong Sun 0001

dblp:21/1509-1 · DBLP profile ↗
← Back
133ranked-venue papers
8as first author
52since 2021 · last 2026
0000-0001-7654-5574ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 48 · 3 first-author · 21 since 2021Artificial intelligence and machine learning · 23 · 13 since 2021Databases, data management, data science and information retrieval · 22 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 5 since 2021Systems, architecture and hardware · 10 · 2 first-author · 1 since 2021Computer networks · 9 · 1 first-author · 1 since 2021Security and privacy · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Attribution Analysis-based Concept Alignment: A Human-in-the-loop Data Debugging Framework
abstract
Ensuring consistently high-quality training data is essential for developing reliable machine learning systems. Recent research demonstrates that incorporating human supervision into training set debugging effectively improves model performance, especially for text classification tasks. However, such methods often prove inapplicable to image understanding tasks, where inherently unstructured pixel data presents challenges in understanding and correcting biases. Inspired by human-AI alignment, we introduce AACA (Attribution Analysis-based Concept Alignment), a human-in-the-loop framework that mitigates bias in the training set by aligning the concepts used by humans and AI during the decision-making process. Specifically, AACA comprises two primary stages: interpretable data bug discovery and targeted data augmentation. During the data bug discovery stage, AACA identifies confounded and valid concepts to explain why prediction failure occurs and what concept the model should focus, using interpretability methods and human annotation. In the stage of targeted data augmentation, AACA adopts these concept-level attributions as clues to synthesize debugging instances via text-to-image generative model. The initial model is then retrained on the augmented set to correct prediction failures. Comparative experiments conducted on crowdsourced annotations and real-world datasets demonstrate that AACA can accurately identifies data bugs and effectively repairs prediction failures, thereby significantly improving prediction performance.
Lei Chai, Hailong Sun 0001, Jingxuan Xu
AAAI3
2026 Efficient Bug Detection by Inferring Implicit API Contract of Pointer State Transition
Xingjing Deng, Xiang Gao 0012, Hailong Sun 0001
DSN4
2026 Quality control in open-ended crowdsourcing: a survey
abstract
Abstract Crowdsourcing provides a flexible approach for leveraging human intelligence to solve large-scale problems, gaining widespread acceptance in domains like intelligent information processing, social decision-making, and crowd ideation. However, the uncertainty of participants significantly compromises the answer quality, sparking substantial research interest. Existing surveys predominantly concentrate on quality control in Boolean tasks, which are generally formulated as simple label classification, ranking, or numerical prediction. Ubiquitous open-ended tasks like question-answering, translation, and semantic segmentation have not been sufficiently discussed. These tasks usually have large to infinite answer spaces and non-unique acceptable answers, posing significant challenges for quality assurance. This survey focuses on quality control methods applicable to open-ended tasks in crowdsourcing. We propose a two-tiered framework to categorize related works. The first tier presents a comprehensive overview of the quality model, covering essential aspects including tasks, workers, answers, and the system. The second tier further refines this classification by breaking it down into more detailed categories: ‘quality dimensions’, ‘evaluation metrics’, and ‘design decisions’. This breakdown provides deeper insights into the internal structure of the quality control model for each aspect. We thoroughly investigate how these quality control methods are implemented in state-of-the-art works and discuss key challenges and potential future research directions.
Lei Chai, Hailong Sun 0001
Frontiers Comput. Sci.2
2026 Fed4Fed: A Privacy-Preserving Federated Statistical Approach for Evaluating Federated Learning Models
abstract
With the widespread application of federated learning in healthcare scenarios, ensuring performance fairness of disease diagnosis models across different medical institutions (clients) has attracted increasing attention. However, accurately evaluating whether models achieve this goal is equally critical yet faces numerous challenges: on one hand, each client can only evaluate the global model based on its own limited private data, which easily leads to performance estimation bias; on the other hand, due to data privacy constraints, clients cannot know the model's performance at other institutions, making it difficult to determine whether the global model truly achieves cross-client performance fairness. To address this, this paper proposes theFed4Fedfederated evaluation framework, which can more accurately evaluate the global model's performance in actual deployment while protecting data privacy by combining private data from multiple clients, and rigorously infer model performance fairness based on statistical hypothesis testing. Specifically,Fed4Feddraws inspiration from federated learning principles to collaboratively utilize multi-party private data while protecting data privacy. Second, it innovatively introduces Bootstrap methods and statistical inference strategies to construct and analyze the statistical distribution of model performance, reducing the randomness of performance evaluation. Third, based on statistical homogeneity testing theory, two fairness testing methods are designed to provide theoretical guarantees for evaluating performance fairness. Finally, experiments on synthetic datasets as well as four types of multi-modal real datasets including CIFAR-10, MNIST, Fashion-MNIST, and SST demonstrate that:Fed4Fedeffectively overcomes the limitations of existing evaluation methods, with a fairness misjudgment rate below 5%, an average confidence interval coverage rate of 94.28% for performance, and robust performance across different degrees of non-independent and identically distributed (non-IID) scenarios.
Zhongchi Wang, Hailong Sun 0001, Wei Ni 0001, Xiang Gao 0012
IEEE Trans. Dependable Secur. Comput.2
2026 NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
abstract
With the growing incorporation of deep neural network (DNN) models into modern software systems, the prohibitive construction costs of DNN models have become a significant challenge in software development. To address this challenge, model reuse has been widely applied to reduce model training costs; however, indiscriminately reusing an entire model may incur significant inference overhead. Consequently, DNN modularization—borrowing the idea of modularization in software engineering—has increasingly gained attention, enabling module reuse by decomposing a DNN model into modules. In particular, the emerging modularizing-while-training (MwT) paradigm, which outperforms modularizing-after-training by incorporating modularization into the model’s training process, has been demonstrated as a more effective approach for DNN modularization. However, existing MwT approaches focus on small-scale convolutional neural network (CNN) models at the convolutional kernel level. They struggle to handle diverse DNNs and large-scale models, particularly Transformer-based models, which consistently achieve state-of-the-art results across various tasks. To address these limitations, we propose NeMo, a scalable and more generalizable MwT approach. NeMo operates at the neuron level—a fundamental component common to all DNNs—thereby ensuring applicability to Transformers and various DNN architectures. Moreover, we design a contrastive learning-based modular training method, equipped with an effective composite loss function, hence being scalable to large-scale models. Comprehensive experiments on two Transformer-based models and four CNN models across two widely used classification datasets demonstrate NeMo’s superiority over the state-of-the-art MwT method. Results show average performance gains of 1.72% in module classification accuracy and a 58.10% reduction in module size. Our findings demonstrate that NeMo exhibits efficacy across both CNN and large-scale Transformer-based models. Moreover, a case study based on open source projects demonstrates the potential benefits of NeMo in practical scenarios, offering a promising approach for achieving scalable and generalizable DNN modularization.
Xiaohan Bi, Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Yue Yu 0001, Xiaojun Liang
ACM Trans. Softw. Eng. Methodol.3
2026 Reference-Based Retrieval-Augmented Unit Test Generation
abstract
Automated unit test generation has been widely studied, with Large Language Models (LLMs) recently showing significant potential. LLMs like GPT-4, trained in vast text and code data, excel in various code-related tasks, including unit test generation. However, existing LLM-based approaches often focus solely on the context within the code itself, such as referenced variables, while neglecting broader task-specific contexts, such as the utility of referring to existing tests of relevant methods in unit test generation. Moreover, in the context of unit test generation, these tools prioritize high code coverage, often at the expense of practical usability, correctness, and maintainability. In response, we propose Reference-Based Retrieval Augmentation , a novel mechanism that extends LLM-based Retrieval-Augmented Generation (RAG) to retrieve relevant information by considering task-specific context. In the unit test generation task, for a given focal method, the reference relationships is defined as the reusability or referentiality of tests between the focal method and other methods. To generate high-quality unit tests for the focal method, the test reference relationships are then used to retrieve relevant methods and their existing unit tests. Specifically, we account for the unique structure of unit tests by dividing the test generation process into Given , When , and Then phases. When generating unit tests for a focal method, we retrieve pre-existing tests of other relevant methods, which can provide valuable insights for any of the Given , When , and Then phases. We implement this approach in a tool called RefTest , which sequentially performs preprocessing, test reference retrieval, and unit test generation, using an incremental strategy in which newly generated tests guide the creation of subsequent ones. We evaluated RefTest on 12 open source projects with 1,515 methods, and the results demonstrate that RefTest consistently outperforms existing tools in terms of correctness, completeness, and maintainability of the generated tests.
Yuanzhang Lin, Xiang Gao 0012, Hailong Sun 0001, Yuan Yuan 0004
ACM Trans. Softw. Eng. Methodol.5
2026 Scalable Large-Scale Multi-Granularity Code Clone Detection via Clustering Search and Pre-Trained Models
abstract
Code cloning is a common phenomenon in software development, which reduces developers’ programming efforts but also poses risks of defect inheritance. Clone detection locates exact or similar pieces of code within or between software systems. With the amount of source code increasing steadily, efficient and large-scale clone detection has become a necessity. Moreover, code clones may occur at various levels of code granularity, e.g., file, function, and block level, which pose more challenges for efficient clone detection. Although numerous methods have been proposed to detect code clones at different granularities, they often suffer from low detection efficiency, false positive results and are typically limited to identifying clones at a specific granularity. In this paper, we introduce an efficient clone detection, named MGCD, to detect code clones among large-scale codebases. Specifically, we embed function-level code into vectors using a pre-trained model and perform clustering search with the IVF Flat algorithm to identify clone candidates. These candidates are then filtered through an entropy-based method to enhance accuracy and avoid false positive results. Moreover, we leverage the information from function-level clone detection results to further conduct file and block level clone detection. We evaluate our approach on the BigCloneBench benchmark. Experimental results show that our approach only takes 0.23 ms to search clone results among 800,000 functions and achieves high precision and recall.
Yifan An, Xiang Gao 0012, Hailong Sun 0001
IEEE Trans. Software Eng.4
2025 Poster: Black-box Attacks on Multimodal Large Language Models through Adversarial ICC Profiles
abstract
Despite their remarkable performance on vision-language tasks, multimodal large language models (MLLMs) remain vulnerable to adversarial examples. However, most existing attacks rely on gradient-based pixel perturbations and require white-box access to model parameters. In this paper, we propose ICCAdv, a novel black-box attack that requires no access to model parameters or gradients. The core idea of ICCAdv is to exploit the discrepancy between human and model perception of images during input processing. This discrepancy arises from the color management process, as human observers perceive rendered images based on ICC profile transformations, whereas most MLLMs circumvent this process and operate directly on raw RGB values. By embedding adversarial ICC profiles into image files, ICCAdv manipulates the perceived color semantics of MLLMs while preserving the natural visual appearance for human observers. Preliminary experiments indicate that ICCAdv can effectively attack state-of-the-art MLLMs while maintaining a natural visual appearance to human observers.
Chengbin Sun, Hailong Sun 0001, Guancheng Li, Jiashuo Liang
CCS2
2025 UICOMPASS: UI Map Guided Mobile Task Automation via Adaptive Action Generation
abstract
Mobile task automation is an emerging technology that leverages AI to automatically execute routine tasks by users' commands on mobile devices like Android, thus enhancing efficiency and productivity.While large language models (LLMs) excel at general mobile tasks through training on massive datasets, they struggle with app-specific workflows.To solve this problem, we designed UI Map, a structured representation of target app's UI information.We further propose a UI Map-guided LLM-based approach UICOMPASS to automate mobile tasks.Specifically, UICOMPASS first leverages static analysis and LLMs to automatically build UI Map from either source codes of apps or byte codes (i.e., APK packages).During task execution, UICOMPASS mines the task-relevant information from UI Map to feed into the LLMs, generates a planned path, and adaptively adjusts the path based on the actual app state and action history.Experimental results demonstrate that UICOMPASS achieves a 14.52% higher task executing success rate than SOTA approaches.Even when only APK is available, UICOMPASS maintains superior performance, demonstrating its applicability to closed-source apps.
Yuanzhang Lin, He Rui, Qingao Dong, Mingyi Zhou, Xiang Gao 0012, Hailong Sun 0001
EMNLP8
2025 Backdoor Defense via Enhanced Splitting and Trap Isolation
Hongrui Yu, Wanyu Lin, Jian Chen 0046, Hailong Sun 0001, Chengbin Sun
ICCV5
2025 CABS: Conflict-Aware and Balanced Sparsification for Enhancing Model Merging
abstract
Model merging based on task vectors, i.e., the parameter differences between fine-tuned models and a shared base model, provides an efficient way to integrate multiple task-specific models into a multitask model without retraining. Recent works have endeavored to address the conflicts between task vectors, one of the significant challenges faced by model merging, through sparsification; however, two issues significantly limit their performance: high parameter overlap and unbalanced weight distribution. To address these issues, we propose a simple yet effective framework called CABS (Conflict-Aware and Balanced Sparsification), consisting of Conflict-Aware Sparsification (CA) and Balanced Sparsification (BS). CA reduces parameter overlap by applying masks during sequential pruning, ensuring that each task vector retains distinct, non-overlapping parameters. BS leverages $n$:$m$ pruning to preserve critical weights while maintaining an even distribution across layers. Our comprehensive experiments demonstrate that CABS outperforms state-of-the-art methods across diverse tasks and model sizes.
Zongzhen Yang, Binhang Qi, Hailong Sun 0001, Wenrui Long, Ruobing Zhao, Xiang Gao 0012
ICML3
2025 Code Property Graph Meets Typestate: A Scalable Framework to Behavioral Bug Detection
abstract
Behavioral bugs caused by incorrect state changes are particularly challenging to identify because they depend on specific code execution paths. While code property graph (CPG) combine multiple code views through abstract syntax trees (AST), their built-in redundancy from syntax details and fixed connection rules make them hard to scale-a major problem when analyzing large software systems. We introduce QVoG, a new framework that improves CPG by combining graphbased code analysis with state behavior checking. Our main innovation lies in simplifying the CPG at the statement level by consolidating control and data flows into meaningful code blocks and optimizing the edges. This approach reduces the graph size by more than 10 times compared to AST-based methods while maintaining accuracy. This lightweight design allows easy integration of state tracking, where we match object lifecycle rules to simplified CPG connections using replaceable patterns. The combination of streamlined graphs and state-aware analysis helps QVoG effectively find difficult-to-identify behavioral bugs, successfully detecting 25 issues (including 17 confirmed cases and 2 official CVE) in real-world projects. Importantly, QVoG analyzes raw source code without requiring compilation and supports projects exceeding 1 million lines of code.
Xingjing Deng, Zhengyao Liu, Xitong Zhong, Shuo Hong, Yixin Yang 0006, Xiang Gao 0012, Xuhui Yan, Hailong Sun 0001
ICSME8
2025 Enhanced Vulnerability Localization: Harmonizing Task-Specific Tuning and General LLM Prompting
abstract
Large Language Models (LLMs) have shown significant potential for vulnerability localization in software security. However, current LLM-based approaches face a critical dilemma: direct application of general-purpose LLMs lacks crucial domainspecific expertise, while fine-tuning suffers from limited robustness when faced with unfamiliar data. These problems result in subpar performance in vulnerability localization and weak generalization capabilities. To address these limitations, we introduce ENVUL, a novel domain adaptation framework for vulnerability localization. ENVUL improves vulnerability localization by synergizing enhanced task-specific tuning with prompt engineering of general-purpose LLMs. ENVUL incorporates three key innovations for addressing two problems: (1) how to optimize fine-tuning for localization task, and (2) when to wisely choose tuning and prompting. To solve the first problem, we introduce: (a). a context Consolidator that captures rich statement-level code semantic, improving the model's understanding of code context; (b). a semantic Indicator employing attention rectification to highlight patterns indicative of vulnerabilities, focusing the model on critical security signals. To solve the second problem, we introduce a dynamic routing mechanism based on joint-representation similarity analysis that strategically delegates tasks between the fine-tuned model and the general LLM. It ensures ENVUL's robust performance across diverse real-world vulnerability types. Real-world evaluations demonstrate ENVUL's robust expertise in outperforming state-of-the-art vulnerability localization baselines, achieving absolute improvements of$\mathbf{2 2. 7 \% - 3 0. 3 \%}$in top-1 accuracy. Notably, ENVul exhibits exceptional generalization, achieving$\mathbf{4 3. 6 \% - 5 0 \%}$higher accuracy on unfamiliar vulnerability types.
Wentong Tian, Yuanzhang Lin, Xiang Gao 0012, Hailong Sun 0001
ICSME4
2025 Turning Swords into Shields: Defense Against Adversarial Examples by Using Trojan Attacks
Chengbin Sun, Hailong Sun 0001
PRCV (4)2
2025 Enhancing Automated Vulnerability Repair Through Dependency Embedding and Pattern Store
abstract
In recent years, the proliferation of software vulnerabilities has significantly increased the complexities and costs associated with manual remediation efforts. Although AI-based methods for automated vulnerability repair are gaining traction, many existing approaches have two limitations: 1) treat code as a sequence of tokens, neglecting critical structural information like control flow and data flow, and 2) do not fully utilize the repair patterns of vulnerabilities. To address these limitations, we introduce FAVOR, an innovative tool that utilizes both the vulnerable function's code and its control flow graph (CFG) as inputs. FAVOR incorporates a dependency embedding module to capture structural and dependency information and leverages CodeT5, a state-of-the-art model pre-trained for code generation tasks. To further enhance the repair process, we introduce a pattern store that uses KNN search to retrieve similar past repair patterns, which helps guide the model toward generating more contextually accurate patches. In our experiments, FAVOR, trained on a dataset of 6548 faulty C/C++ functions, repaired 45 more vulnerabilities compared to VULREPAIR, demonstrating improved accuracy and efficiency in automated vulnerability repair.
Qingao Dong, Yuanzhang Lin, Hailong Sun 0001, Xiang Gao 0012
SANER3
2025 Deep learning-based software engineering: progress, challenges, and opportunities
abstract
Abstract Researchers have recently achieved significant advances in deep learning techniques, which in turn has substantially advanced other research disciplines, such as natural language processing, image processing, speech recognition, and software engineering. Various deep learning techniques have been successfully employed to facilitate software engineering tasks, including code generation, software refactoring, and fault localization. Many studies have also been presented in top conferences and journals, demonstrating the applications of deep learning techniques in resolving various software engineering tasks. However, although several surveys have provided overall pictures of the application of deep learning techniques in software engineering, they focus more on learning techniques, that is, what kind of deep learning techniques are employed and how deep models are trained or fine-tuned for software engineering tasks. We still lack surveys explaining the advances of subareas in software engineering driven by deep learning techniques, as well as challenges and opportunities in each subarea. To this end, in this study, we present the first task-oriented survey on deep learning-based software engineering. It covers twelve major software engineering subareas significantly impacted by deep learning techniques. Such subareas spread out through the whole lifecycle of software development and maintenance, including requirements engineering, software development, testing, maintenance, and developer collaboration. As we believe that deep learning may provide an opportunity to revolutionize the whole discipline of software engineering, providing one survey covering as many subareas as possible in software engineering can help future research push forward the frontier of deep learning-based software engineering more systematically. For each of the selected subareas, we highlight the major advances achieved by applying deep learning techniques with pointers to the available datasets in such a subarea. We also discuss the challenges and opportunities concerning each of the surveyed software engineering subareas.
Xiangping Chen, Xing Hu 0008, Yuan Huang 0002, He Jiang 0001, Weixing Ji, Yanjie Jiang, Yanyan Jiang 0001, Bo Liu 0094, Hui Liu 0003, Xiaoli Lian, Guozhu Meng, Xin Peng 0001, Hailong Sun 0001, Lin Shi 0006, Bo Wang 0050, Chong Wang 0013, Jifeng Xuan, Xin Xia 0001, Yibiao Yang, Yixin Yang 0006, Li Zhang 0029, Yuming Zhou, Lu Zhang 0023
Sci. China Inf. Sci.14
2025 Understanding vulnerabilities in software supply chains
Xiang Gao 0012, Hailong Sun 0001
Empir. Softw. Eng.3
2025 Unlearning Attacks for Regression Learning
abstract
Recently, the machine unlearning has emerged as a popular method for efficiently erasing the impact of personal data in machine learning (ML) models upon the data owner's removal request. However, few studies take into consideration the security concerns that may exist in the unlearning process. In this article, we propose the first unlearning attack dubbed unlearning attack for regression learning (UnAR) to deliberately influence the predictive behavior of the target sample against regression learning models. The central concept of UnAR revolves around misleading the regression model into erasing the information associated with the influential samples for the target sample. Observing that the influential samples for target data are generally located far away from the regression plane, we thus propose two novel methods, known as influential sample selection (ISS) and influential sample unlearning (ISU), to identify and subsequently eliminate the lineage of the influential samples. By doing so, we can substantially introduce bias into the prediction pertaining to the target sample, yielding the deliberate manipulation for the user adversely. We extensively evaluate UnAR on five public datasets, and the experimental results indicate our attacks can achieve prediction deviations over 35% by unlearning only 0.5% data as the influential samples.
Jian Chen 0046, Wenlong Shi, Wanyu Lin, Chen Wang 0011, Wei Liu 0004, Hailong Sun 0001, Gaoyang Liu
IEEE Trans. Neural Networks Learn. Syst.6
2025 Toward Open-World Domain Adaptation via Iteratively Contrastive Learning and Clustering
abstract
The open-set domain adaptation (DA) aims to address both covariate shift and category shift between a labeled source domain and an unlabeled target domain. Nevertheless, existing open-set DA methods always ignore the demand for discovering novel classes that are not present in the source domain and simply reject them as "unknown" sets without further exploration, which motivates us to understand the unknown sets more specifically. In this article, we present a more challenging open-world DA problem that recognizes seen classes while discovering novel classes in the target domain. To address this problem, we propose a novel framework that converts this problem into a clustering task via contrastive learning to learn pairwise relationships among the instances. More specifically, our method consists of two iterative steps. The semi-supervised clustering step clusters the unlabeled target data and separates it into seen and novel classes. In the contrastive learning step, based on the cluster assignments, we design tailored contrastive losses that learn pairwise relationships to reduce domain discrepancy and discover novel classes. Our method can be optimized as an example of expectation maximization (EM). We establish several baselines by extending related work. Our method obtains the superior performance on five public datasets, benchmarking this challenging setting for future research.
Jingzheng Li, Hailong Sun 0001, Jiyi Li, Shikui Wei
IEEE Trans. Neural Networks Learn. Syst.2
2024 RA3: A Human-in-the-loop Framework for Interpreting and Improving Image Captioning with Relation-Aware Attribution Analysis
abstract
Interpreting model behavior is crucial for model evaluation and optimization. Recent research demonstrates that incorporating human intelligence into the learning process effectively improve the interpretability and performance of the machine learning models, especially for simple classification tasks. However, the image captioning task has not received much attention. Such complex sequential tasks generally contain semantic relationships between different concepts, which pose challenges for interpreting model behavior and developing optimization methods. In this paper, we present RA3(Relation-Aware Attribution Analysis), a human-in-the-loop framework, for improving the interpretability, and further boosting the performance of the image captioning model. Specifically, we first engage human participants in two types of annotation tasks to identify what the model actually focuses on (model attribution) and what it should focus on (human rationale) at the conceptual level, supported by machine learning interpretability methods. Then, we identify and filter hard instances based on relation-aware model attribution for both validating the quality of the explanation and eliminating low-quality captions (this process is also considered as a kind of data debugging). We subsequently designed an explanation loss that penalizes the difference between model attribution and human rationale to optimize the model's behavior for improving caption quality. Through extensive experiments on crowdsourced annotations and MSCOCO, the experiment results indicate that the explanations produced by RA3can accurately describe the model's behavior, effectively identify difficult instances, and significantly improve the caption quality.
Lei Chai, Hailong Sun 0001, Jingzheng Li
ICDE3
2024 Modularizing while Training: A New Paradigm for Modularizing DNN Models
abstract
Deep neural network (DNN) models have become increasingly crucial components of intelligent software systems. However, training a DNN model is typically expensive in terms of both time and computational resources. To address this issue, recent research has focused on reusing existing DNN models - borrowing the concept of software reuse in software engineering. However, reusing an entire model could cause extra overhead or inherit the weaknesses from the undesired functionalities. Hence, existing work proposes to decompose an already trained model into modules, i.e., modularizing-after-training, to enable module reuse. Since the trained models are not built for modularization, modularizing-after-training may incur huge overhead and model accuracy loss. In this paper, we propose a novel approach that incorporates modularization into the model training process, i.e., modularizing-while-training (MwT). We train a model to be structurally modular through two loss functions that optimize intra-module cohesion and inter-module coupling. We have implemented the proposed approach for modularizing Convolutional Neural Network (CNN) models. The evaluation results on representative models demonstrate that MwT outperforms the existing state-of-the-art modularizing-after-training approach. Specifically, the accuracy loss caused by MwT is only 1.13 percentage points, which is less than that of the existing approach. The kernel retention rate of the modules generated by MwT is only 14.58%, with a reduction of 74.31% over the existing approach. Furthermore, the total time cost required for training and modularizing is only 108 minutes, which is half the time required by the existing approach. Our work demonstrates that MwT is a new and more effective paradigm for realizing DNN model modularization, offering a fresh perspective on achieving model reuse.
Binhang Qi, Hailong Sun 0001, Hongyu Zhang 0002, Ruobing Zhao, Xiang Gao 0012
ICSE2
2024 Investigating White-Box Attacks for On-Device Models
abstract
Numerous mobile apps have leveraged deep learning capabilities. However, on-device models are vulnerable to attacks as they can be easily extracted from their corresponding mobile apps. Although the structure and parameters information of these models can be accessed, existing on-device attacking approaches only generate black-box attacks (i.e., indirect white-box attacks), which are less effective and efficient than white-box strategies. This is because mobile deep learning (DL) frameworks like TensorFlow Lite (TFLite) do not support gradient computing (referred to as non-debuggable models), which is necessary for white-box attacking algorithms. Thus, we argue that existing findings may underestimate the harm-fulness of on-device attacks. To validate this, we systematically analyze the difficulties of transforming the on-device model to its debuggable version and propose a Reverse Engineering framework for On-device Models (REOM), which automatically reverses the compiled on-device TFLite model to its debuggable version, enabling attackers to launch white-box attacks. Our empirical results show that our approach is effective in achieving automated transformation (i.e., 92.6%) among 244 TFLite models. Compared with previous attacks using surrogate models, REOM enables attackers to achieve higher attack success rates (10.23%→89.03%) with a hundred times smaller attack perturbations (1.0→0.01). Our findings emphasize the need for developers to carefully consider their model deployment strategies, and use white-box methods to evaluate the vulnerability of on-device models. Our artifacts 1 are available.
Mingyi Zhou, Xiang Gao 0012, Jing Wu 0021, Kui Liu 0001, Hailong Sun 0001, Li Li 0029
ICSE5
2024 API Misuse Detection via Probabilistic Graphical Model
abstract
API misuses can cause a range of issues in software development, including program crashes, bugs, and vulnerabilities. Different approaches have been developed to automatically detect API misuses by checking the program against usage rules extracted from extensive codebase or API documents. However, these mined rules may not be precise or complete, leading to high false positive/negative rates. In this paper, we propose a novel solution to this problem by representing the mined API usage rules as a probabilistic graphical model, where each rule's probability value represents its trustworthiness of being correct. Our approach automatically constructs probabilistic usage rules by mining codebase and documents, and aggregating knowledge from different sources. Here, the usage rules obtained from the codebase initialize the probabilistic model, while the knowledge from the documents serves as a supplement for adjusting and complementing the probabilities accordingly. We evaluate our approach on the MuBench benchmark. Experimental results show that our approach achieves 42.0% precision and 54.5% recall, significantly outperforming state-of-the-art approaches.
Wentong Tian, Xiang Gao 0012, Hailong Sun 0001, Li Li 0029
ISSTA4
2024 FedEvalFair: A Privacy-Preserving and Statistically Grounded Federated Fairness Evaluation Framework
abstract
Federated learning has rapidly gained attention in the industrial sector due to its significant advantages in protecting privacy. However, ensuring the fairness of federated learning models post-deployment presents a challenge in practical applications. Given that clients typically rely on limited private datasets to assess model fairness, this constrains their ability to make accurate judgments about the fairness of the model. To address this issue, we propose an innovative evaluation framework, FedEvalFair, which integrates private data from multiple clients to comprehensively assess the fairness of models in actual deployment without compromising data privacy. Firstly, FedEvalFair draws on the concept of federated learning to achieve a comprehensive assessment while protecting privacy. Secondly, based on the statistical concept of "estimating the population from the sample", FedEvalFair is capable of estimating the fairness performance of the model in real-world settings from a limited data sample. Thirdly, we have designed a flexible two-stage evaluation strategy based on statistical hypothesis testing. We verified the theoretical performance and sensitivity to fairness variations of FedEvalFair using Monte Carlo simulations, demonstrating the superior performance of its two-stage evaluation strategy. Additionally, we validated the effectiveness of the FedEvalFair method on real-world datasets, including UCI Adult and eICU, and demonstrated its stability in dealing with real-world data distribution changes compared to traditional evaluation methods.
Zhongchi Wang, Hailong Sun 0001
ACM Multimedia2
2024 ModelGalaxy: A Versatile Model Retrieval Platform
abstract
With the growing number of available machine learning models and the emergence of model-sharing platforms, model reuse has become a significant approach to harnessing the power of artificial intelligence. One of the key issues to realizing model reuse resides in efficiently and accurately finding the target models that meet user needs from a model repository. However, the existing popular model-sharing platforms (e.g., Hugging Face) mainly support model retrieval based on model name matching and task filtering. If not familiar with the platform or specific models, users may suffer from low retrieval efficiency and a less user-friendly interaction experience. To address these issues, we have developed ModelGalaxy, a versatile model retrieval platform supporting multiple model retrieval methods, including keyword-based search, dataset-based search, and user-task-centric search. Moreover, ModelGalaxy leverages the power of large language models to provide users with easily retrieving and using models. Our source code is available at https://github.com/zwl906711886/ModelGalaxy.
Wenling Zhang, Zhaotian Li, Hailong Sun 0001, Xiang Gao 0012, Xudong Liu 0001
SIGIR4
2024 Investigating and Detecting Silent Bugs in PyTorch Programs
abstract
Deep Learning (DL) has been widely applied in various fields. Unlike traditional software, DL programs possess the “black box” characteristic that can make it challenging for developers to debug when anomalous behaviors arise. In particular, silent bugs, a type of bugs in DL programs, can lead to erroneous behaviors without causing system crashes or suspensions, and they do not display error messages to users. This makes silent bugs more difficult for developers to discover, locate, and fix. In this paper, we present the first detailed study of silent bugs in PyTorch programs. We collect 14,523 posts from the official PyTorch forum and use a LLM-based semi-automated approach to filter the silent bugs. By analyzing the symptoms, root causes, and patterns of silent bugs, we have derived several important findings and implications: (1) most silent bugs cause abnormal outputs, which requires the design of more flexible test oracles to detect them, (2) the wide range of symptoms and root causes do not necessarily have one-to-one correspondences, which makes detecting and debugging silent bugs more challenging, (3) silent bugs exhibit common bug patterns, such as redundant, missing, or misplaced operations. Building upon these findings, we design and implement an extensible rule-based tool PYSIASSIST to help developer debug and resolve silent bugs. Evaluation results show that Pysiassist achieves 92.4% precision and 85.3% recall, outperforming existing techniques.
Shuo Hong, Hailong Sun 0001, Xiang Gao 0012, Shin Hwei Tan
SANER2
2024 Reducing False Positives of Static Bug Detectors Through Code Representation Learning
abstract
With the increasing significance of software correctness and security, automatic static analysis tools (ASATs) play a more and more important role in software development due to their ability and scalability. However, compared to dynamic analysis methods, static tools often suffer from the severe problem of generating high false positive rates, due to their analysis mechanisms. To alleviate the false positive problem, many approaches have been proposed, which focus on manually extracted features from code snippets and then prioritize real warnings by means of statistics or machine learning techniques. However, manual encoded features are insufficient to achieve satisfactory performance across different datasets. In this study, we focus on exploring the effectiveness of various code representation learning (CRL) techniques in understanding the semantics of warnings generated by ASATs. In particular, our large-scale empirical study not only reveals that CRL models can effectively differentiate buggy code snippets (i.e., containing warnings detected by ASATs) from clean ones (the median of F1-score reaches 87.3 % for binary classification, and reaches 77.4 % for multi-class classification), they are also promising in identifying false positive warnings (the F1-score of best performer is 75.6%). Such findings drive us to further design a novel approach named PRI SM, to PRIoritize Static warnings based on aggregating multiple CRL Models to reduce the false positives generated by existing ASATs. Extensive evaluations demonstrate that our designed approach can outperform existing baselines significantly.
Yixin Yang 0006, Ming Wen 0001, Xiang Gao 0012, Hailong Sun 0001
SANER5
2024 Target Structure Learning Framework for Unsupervised Multi-Class Domain Adaptation
abstract
Unsupervised multi-class domain adaptation (multi-class UDA) has recently been proposed to fill the gap between empirically practical methods for multi-class classification and well-founded theory with the setting of binary classification. Nevertheless, the multi-class UDA methods use model predictions to characterize the disagreement of multi-class scoring hypotheses, which is used to optimize the divergence between domain distributions. Such self-training manner may bring inaccurate model predictions, which would damage the target structure due to the absence of labels of the target domain, leading to sub-optimal performance. On the other hand, this disagreement between multi-class scoring hypotheses does not involve the relationships among all of the multiple classes. It causes that multi-class UDA cannot properly connect the advanced practical UDA methods that consider class-conditional distribution alignment. Thus, we propose to exploit the target structure information and then incorporate it into multi-class UDA to achieve class-conditional distribution alignment. We theoretically and experimentally explain the importance of accurate target structure information to reduce the expected error on the target domain. Notably, our method achieves state-of-the-art results on three commonly-used benchmarks with different scales. In addition, using the target structure information, we propose a variant to cope with noisy open-world source domains such as noisy labels and out-of-distribution samples, enhancing the robustness of our method. The source code is available at https://github.com/jingzhengli/Multi_Class_UDA .
Jingzheng Li, Hailong Sun 0001, Lei Chai, Jiyi Li
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Reusing Convolutional Neural Network Models through Modularization and Composition
abstract
With the widespread success of deep learning technologies, many trained deep neural network (DNN) models are now publicly available. However, directly reusing the public DNN models for new tasks often fails due to mismatching functionality or performance. Inspired by the notion of modularization and composition in software reuse, we investigate the possibility of improving the reusability of DNN models in a more fine-grained manner. Specifically, we propose two modularization approaches named CNNSplitter and GradSplitter, which can decompose a trained convolutional neural network (CNN) model for N -class classification into N small reusable modules. Each module recognizes one of the N classes and contains a part of the convolution kernels of the trained CNN model. Then, the resulting modules can be reused to patch existing CNN models or build new CNN models through composition. The main difference between CNNSplitter and GradSplitter lies in their search methods: the former relies on a genetic algorithm to explore search space, while the latter utilizes a gradient-based search method. Our experiments with three representative CNNs on three widely used public datasets demonstrate the effectiveness of the proposed approaches. Compared with CNNSplitter, GradSplitter incurs less accuracy loss, produces much smaller modules (19.88% fewer kernels), and achieves better results on patching weak models. In particular, experiments on GradSplitter show that (1) by patching weak models, the average improvement in terms of precision, recall, and F1-score is 17.13%, 4.95%, and 11.47%, respectively, and (2) for a new task, compared with the models trained from scratch, reusing modules achieves similar accuracy (the average loss of accuracy is only 2.46%) without a costly training process. Our approaches provide a viable solution to the rapid development and improvement of CNN models.
Binhang Qi, Hailong Sun 0001, Hongyu Zhang 0002, Xiang Gao 0012
ACM Trans. Softw. Eng. Methodol.2
2024 MTL-TRANSFER: Leveraging Multi-task Learning and Transferred Knowledge for Improving Fault Localization and Program Repair
abstract
Fault localization (FL) and automated program repair (APR) are two main tasks of automatic software debugging. Compared with traditional methods, deep learning-based approaches have been demonstrated to achieve better performance in FL and APR tasks. However, the existing deep learning-based FL methods ignore the deep semantic features or only consider simple code representations. And for APR tasks, existing template-based APR methods are weak in selecting the correct fix templates for more effective program repair, which are also not able to synthesize patches via the embedded end-to-end code modification knowledge obtained by training models on large-scale bug-fix code pairs. Moreover, in most of FL and APR methods, the model designs and training phases are performed separately, leading to ineffective sharing of updated parameters and extracted knowledge during the training process. This limitation hinders the further improvement in the performance of FL and APR tasks. To solve the above problems, we propose a novel approach called MTL-TRANSFER, which leverages a multi-task learning strategy to extract deep semantic features and transferred knowledge from different perspectives. First, we construct a large-scale open-source bug datasets and implement 11 multi-task learning models for bug detection and patch generation sub-tasks on 11 commonly used bug types, as well as one multi-classifier to learn the relevant semantics for the subsequent fix template selection task. Second, an MLP-based ranking model is leveraged to fuse spectrum-based, mutation-based and semantic-based features to generate a sorted list of suspicious statements. Third, we combine the patches generated by the neural patch generation sub-task from the multi-task learning strategy with the optimized fix template selecting order gained from the multi-classifier mentioned above. Finally, the more accurate FL results, the optimized fix template selecting order, and the expanded patch candidates are combined together to further enhance the overall performance of APR tasks. Our extensive experiments on widely-used benchmark Defects4J show that MTL-TRANSFER outperforms all baselines in FL and APR tasks, proving the effectiveness of our approach. Compared with our previously proposed FL method TRANSFER-FL (which is also the state-of-the-art statement-level FL method), MTL-TRANSFER increases the faults hit by 8/11/12 on Top-1/3/5 metrics (92/159/183 in total). And on APR tasks, the number of successfully repaired bugs of MTL-TRANSFER under the perfect localization setting reaches 75, which is 8 more than our previous APR method TRANSFER-PR. Furthermore, another experiment to simulate the actual repair scenarios shows that MTL-TRANSFER can successfully repair 15 and 9 more bugs (56 in total) compared with TBar and TRANSFER, which demonstrates the effectiveness of the combination of our optimized FL and APR components.
Xu Wang 0007, Xiangxin Meng, Hongliang Cao, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001, Chunming Hu
ACM Trans. Softw. Eng. Methodol.6
2023 AutoMRM: A Model Retrieval Method Based on Multimodal Query and Meta-learning
abstract
With more and more Deep Neural Network (DNN) models are publicly available on model sharing platforms (e.g., HuggingFace), model reuse has become a promising way in practice to improve the efficiency of DNN model construction by avoiding the costs of model training. To that end, a pivotal step for model reuse is model retrieval, which facilitates discovering suitable models from a model hub that match the requirements of users. However, the existing model retrieval methods have inadequate performance and efficiency, since they focus on matching user requirements with the model names, and thus cannot work well for high-dimensional data such as images. In this paper, we propose a user-task-centric multimodal model retrieval method named AutoMRM. AutoMRM can retrieve DNN models suitable for the user's task according to both the dataset and description of the task. Moreover, AutoMRM utilizes meta-learning to retrieve models for previously unseen task queries. Specifically, given a task, AutoMRM extracts the latent meta-features from the dataset and description for training meta-learners offline and obtaining the representation of user task queries online. Experimental results demonstrate that AutoMRM outperforms existing model retrieval methods including the state-of-the-art method in both effectiveness and efficiency.
Zhaotian Li, Binhang Qi, Hailong Sun 0001, Xiang Gao 0012
CIKM3
2023 Learning from Noisy Crowd Labels with Logics
abstract
This paper explores the integration of symbolic logic knowledge into deep neural networks for learning from noisy crowd labels. We introduce Logic-guided Learning from Noisy Crowd Labels (Logic-LNCL), an EM-alike iterative logic knowledge distillation framework that learns from both noisy labeled data and logic rules of interest. Unlike traditional EM methods, our framework contains a "pseudo-E-step" that distills from the logic rules a new type of learning target, which is then used in the "pseudo-M-step" for training the classifier. Extensive evaluations on two real-world datasets for text sentiment classification and named entity recognition demonstrate that the proposed framework improves the state-of-the-art and provides a new solution to learning from noisy crowd labels.
Hailong Sun 0001, Haoqian He
ICDE2
2023 Template-based Neural Program Repair
abstract
In recent years, template-based and NMT-based automated program repair methods have been widely studied and achieved promising results. However, there are still disadvantages in both methods. The template-based methods cannot fix the bugs whose types are beyond the capabilities of the templates and only use the syntax information to guide the patch synthesis, while the NMT-based methods intend to generate the small range of fixed code for better performance and may suffer from the OOV (Out-of-vocabulary) problem. To solve these problems, we propose a novel template-based neural program repair approach called TENURE to combine the template-based and NMT- based methods. First, we build two large-scale datasets for 35 fix templates from template-based method and one special fix template (single-line code generation) from NMT-based method, respectively. Second, the encoder-decoder models are adopted to learn deep semantic features for generating patch intermediate representations (IRs) for different templates. The optimized copy mechanism is also used to alleviate the OOV problem. Third, based on the combined patch IRs for different templates, three tools are developed to recover real patches from the patch IRs, replace the unknown tokens, and filter the patch candidates with compilation errors by leveraging the project-specific information. On Defects4J-vl.2, TENURE can fix 79 bugs and 52 bugs with perfect and Ochiai fault localization, respectively. It is able to repair 50 and 32 bugs as well on Defects4J-v2.0. Compared with the existing template-based and NMT-based studies, TENURE achieves the best performance in all experiments.
Xiangxin Meng, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001, Chunming Hu
ICSE4
2023 Reusing Deep Neural Network Models through Model Re-engineering
abstract
Training deep neural network (DNN) models, which has become an important task in today's software development, is often costly in terms of computational resources and time. With the inspiration of software reuse, building DNN models through reusing existing ones has gained increasing attention recently. Prior approaches to DNN model reuse have two main limitations: 1) reusing the entire model, while only a small part of the model's functionalities (labels) are required, would cause much overhead (e.g., computational and time costs for inference), and 2) model reuse would inherit the defects and weaknesses of the reused model, and hence put the new system under threats of security attack. To solve the above problem, we propose SeaM, a tool that re-engineers a trained DNN model to improve its reusability. Specifically, given a target problem and a trained model, SeaM utilizes a gradient-based search method to search for the model's weights that are relevant to the target problem. The re-engineered model that only retains the relevant weights is then reused to solve the target problem. Evaluation results on widely-used models show that the re-engineered models produced by SeaM only contain 10.11% weights of the original models, resulting 42.41% reduction in terms of inference time. For the target problem, the re-engineered models even outperform the original models in classification accuracy by 5.85%. Moreover, reusing the re-engineered models inherits an average of 57% fewer defects than reusing the entire model. We believe our approach to reducing reuse overhead and defect inheritance is one important step forward for practical model reuse.
Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Hongyu Zhang 0002, Zhaotian Li, Xudong Liu 0001
ICSE2
2023 Black-Box Data Poisoning Attacks on Crowdsourcing
abstract
Understanding the vulnerability of label aggregation against data poisoning attacks is key to ensuring data quality in crowdsourced label collection. State-of-the-art attack mechanisms generally assume full knowledge of the aggregation models while failing to consider the flexibility of malicious workers in selecting which instances to label. Such a setup limits the applicability of the attack mechanisms and impedes further improvement of their success rate. This paper introduces a black-box data poisoning attack framework that finds the optimal strategies for instance selection and labeling to attack unknown label aggregation models in crowdsourcing. We formulate the attack problem on top of a generic formalization of label aggregation models and then introduce a substitution approach that attacks a substitute aggregation model in replacement of the unknown model. Through extensive validation on multiple real-world datasets, we demonstrate the effectiveness of both instance selection and model substitution in improving the success rate of attacks.
Yongqiang Yang, Dingqi Yang, Hailong Sun 0001
IJCAI4
2023 Detecting Condition-Related Bugs with Control Flow Graph Neural Network
abstract
Automated bug detection is essential for high-quality software development and has attracted much attention over the years. Among the various bugs, previous studies show that the condition expressions are quite error-prone and the condition-related bugs are commonly found in practice. Traditional approaches to automated bug detection are usually limited to compilable code and require tedious manual effort. Recent deep learning-based work tends to learn general syntactic features based on Abstract Syntax Tree (AST) or apply the existing Graph Neural Networks over program graphs. However, AST-based neural models may miss important control flow information of source code, and existing Graph Neural Networks for bug detection tend to learn local neighbourhood structure information. Generally, the condition-related bugs are highly influenced by control flow knowledge, therefore we propose a novel CFG-based Graph Neural Network (CFGNN) to automatically detect condition-related bugs, which includes a graph-structured LSTM unit to efficiently learn the control flow knowledge and long-distance context information. We also adopt the API-usage attention mechanism to leverage the API knowledge. To evaluate the proposed approach, we collect real-world bugs in popular GitHub repositories and build a large-scale condition-related bug dataset. The experimental results show that our proposed approach significantly outperforms the state-of-the-art methods for detecting condition-related bugs.
Jian Zhang 0087, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001, Chunming Hu, Yang Liu 0003
ISSTA4
2023 Neural-Hidden-CRF: A Robust Weakly-Supervised Sequence Labeler
abstract
We propose a neuralized undirected graphical model called Neural-Hidden-CRF to solve the weakly-supervised sequence labeling problem. Under the umbrella of undirected graphical theory, the proposed Neural-Hidden-CRF embedded with a hidden CRF layer models the variables of word sequence, latent ground truth sequence, and weak label sequence with the global perspective that undirected graphical models particularly enjoy. In Neural-Hidden-CRF, we can capitalize on the powerful language model BERT or other deep models to provide rich contextual semantic knowledge to the latent ground truth sequence, and use the hidden CRF layer to capture the internal label dependencies. Neural-Hidden-CRF is conceptually simple and empirically powerful. It obtains new state-of-the-art results on one crowdsourcing benchmark and three weak-supervision benchmarks, including outperforming the recent advanced model CHMM by 2.80 F1 points and 2.23 F1 points in average generalization and inference performance, respectively.
Hailong Sun 0001, Wanhao Zhang, Chunyi Xu, Qianren Mao
KDD2
2023 LiFT: Transfer Learning in Vision-Language Models for Downstream Adaptation and Generalization
abstract
Pre-trained Vision-Language Models (VLMs) on large-scale image-text pairs, e.g., CLIP, have shown promising performance on zero-shot knowledge transfer. Recently, fine-tuning pre-trained VLMs to downstream few-shot classification with limited image annotation data yields significant gains. However, there are two limitations. First, most of the methods for fine-tuning VLMs only update newly added parameters while keeping the whole VLM frozen. Thus, it remains unclear how to directly update the VLM itself. Second, fine-tuning VLMs to a specific set of base classes would deteriorate the well-learned representation space such that the VLMs generalize poorly on novel classes. To address these issues, we first propose Layer-wise Fine-Tuning (LiFT) which achieves average gains of 3.9%, 4.3%, 4.2% and 4.5% on base classes under 2-, 4-, 8- and 16-shot respectively compared to the baseline CoOp over 11 datasets. Alternatively, we provide a parameter-efficient LiFT-Adapter exhibiting favorable performance while updating only 1.66% of total parameters. Further, we design scalable LiFT-NCD to identify both base classes and novel classes, which boosts the accuracy by an average of 5.01% over zero-shot generalization of CLIP, exploring the potential of VLMs in discovering novel classes.
Jingzheng Li, Hailong Sun 0001
ACM Multimedia2
2023 NaCL: noise-robust cross-domain contrastive learning for unsupervised domain adaptation
Jingzheng Li, Hailong Sun 0001
Mach. Learn.2
2023 Beyond confusion matrix: learning from multiple annotators with awareness of instance features
Jingzheng Li, Hailong Sun 0001, Jiyi Li
Mach. Learn.2
2022 Adversarial Learning from Crowds
abstract
Learning from Crowds (LFC) seeks to induce a high-quality classifier from training instances, which are linked to a range of possible noisy annotations from crowdsourcing workers under their various levels of skills and their own preconditions. Recent studies on LFC focus on designing new methods to improve the performance of the classifier trained from crowdsourced labeled data. To this day, however, there remain under-explored security aspects of LFC systems. In this work, we seek to bridge this gap. We first show that LFC models are vulnerable to adversarial examples---small changes to input data can cause classifiers to make prediction mistakes. Second, we propose an approach, A-LFC for training a robust classifier from crowdsourced labeled data. Our empirical results on three real-world datasets show that the proposed approach can substantially improve the performance of the trained classifier even with the existence of adversarial examples. On average, A-LFC has 10.05% and 11.34% higher test robustness than the state-of-the-art in the white-box and black-box attack settings, respectively.
Hailong Sun 0001, Yongqiang Yang
AAAI2
2022 Improving Fault Localization and Program Repair with Deep Semantic Features and Transferred Knowledge
abstract
Automatic software debugging mainly includes two tasks of fault localization and automated program repair. Compared with the traditional spectrum-based and mutation-based methods, deep learning-based methods are proposed to achieve better performance for fault localization. However, the existing methods ignore the deep semantic features or only consider simple code representations. They do not leverage the existing bug-related knowledge from large-scale open-source projects either. In addition, existing template-based program repair techniques can incorporate project specific information better than deep-learning approaches. However, they are weak in selecting the fix templates for efficient program repair. In this work, we propose a novel approach called TRANSFER, which leverages the deep semantic features and transferred knowledge from open-source data to improve fault localization and program repair. First, we build two large-scale open-source bug datasets and design 11 BiLSTM-based binary classifiers and a BiLSTM-based multi-classifier to learn deep semantic features of statements for fault localization and program repair, respectively. Second, we combine semantic-based, spectrum-based and mutation-based features and use an MLP-based model for fault localization. Third, the semantic-based features are leveraged to rank the fix templates for program repair. Our extensive experiments on widely-used benchmark De-fects4J show that TRANSFER outperforms all baselines in fault localization, and is better than existing deep-learning methods in automated program repair. Compared with the typical template-based work TBar, TRANSFER can correctly repair 6 more bugs (47 in total) on Defects4J.
Xiangxin Meng, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001
ICSE4
2022 Patching Weak Convolutional Neural Network Models through Modularization and Composition
abstract
Despite great success in many applications, deep neural networks are not always robust in practice. For instance, a convolutional neuron network (CNN) model for classification tasks often performs unsatisfactorily in classifying some particular classes of objects. In this work, we are concerned with patching the weak part of a CNN model instead of improving it through the costly retraining of the entire model. Inspired by the fundamental concepts of modularization and composition in software engineering, we propose a compressed modularization approach, CNNSplitter, which decomposes a strong CNN model for N-class classification into N smaller CNN modules. Each module is a sub-model containing a part of the convolution kernels of the strong model. To patch a weak CNN model that performs unsatisfactorily on a target class (TC), we compose the weak CNN model with the corresponding module obtained from a strong CNN model. The ability of the weak CNN model to recognize the TC can thus be improved through patching. Moreover, the ability to recognize non-TCs is also improved, as the samples misclassified as TC could be classified as non-TCs correctly. Experimental results with two representative CNNs on three widely-used datasets show that the averaged improvement on the TC in terms of precision and recall are 12.54% and 2.14%, respectively. Moreover, patching improves the accuracy of non-TCs by 1.18%. The results demonstrate that CNNSplitter can patch a weak CNN model through modularization and composition, thus providing a new solution for developing robust CNN models.
Binhang Qi, Hailong Sun 0001, Xiang Gao 0012, Hongyu Zhang 0002
ASE2
2022 Correct Twice at Once: Learning to Correct Noisy Labels for Robust Deep Learning
abstract
Deep Neural Networks (DNNs) have shown impressive performance on large-scale training data with high-quality annotations. However, the collected annotations inevitably contain inaccurate labels in consideration of time and money budget, which causes DNNs to generalize poorly on the test set. To combat noisy labels in deep learning, the label correction methods are dedicated to simultaneously updating model parameters and correcting noisy labels, in which the noisy labels are usually corrected based on model predictions, the topological structures of data, or the aggregation of multiple models. However, such self-training manner cannot guarantee that the direction of label correction is always reliable. In view of this, we propose a novel label correction method to supervise and guide the process of label correction. In particular, the proposed label correction is an online two-fold process at each iteration only through back-propagation. The first label correction minimizes the empirical risk on noisy training data using noise-tolerant loss function, and the second label correction adopts a meta-learning paradigm to rectify the direction of first label correction so that the model can perform optimally in the evaluation procedure. Extensive experiments demonstrate the effectiveness of the proposed method on synthetic datasets with varying noise types and noise rates. Notably, our method achieves test accuracy of 77.37% on the real-world Clothing1M dataset.
Jingzheng Li, Hailong Sun 0001
ACM Multimedia2
2022 A Collaboration-Aware Approach to Profiling Developer Expertise with Cross-Community Data
abstract
Developer expertise is an important factor that should be considered in various software development activities. And it is challenging to accurately profile the expertise of developers as their activities often disperse across different online communities, such as Community Question Answering sites (e.g., Stack Overflow) and Open Source Software platforms (e.g., GitHub). In this regard, early work mainly considers a single community while recent studies are starting to profile developers with cross-community data. However, few works consider the collaborative interactions among developers in evaluating developer expertise across communities. In this work, we propose a collaboration-aware approach to profiling developer expertise using cross-community data by taking into consideration developers’ contributions, collaborative interactions, and the dynamic changes of expertise. Specifically, we are concerned with the common developers in GitHub and Stack Overflow. First, we propose a time-sensitive model to characterize the developer’s expertise in the two communities and integrate the results to generate basic expertise profiles. Second, we build a developer network by analyzing the collaborative interactions among the developers of the two communities. Finally, we apply the topic-sensitive PageRank algorithm to incorporate developer relationships into expertise profiling. Results of extensive experiments on a large number of common developers of GitHub and Stack Overflow demonstrate the effectiveness of our approach.
Xiaotao Song, Jiafei Yan, Yuexin Huang, Hailong Sun 0001, Hongyu Zhang 0002
QRS4
2022 Incorporating pixel proximity into answer aggregation for crowdsourced image segmentation
Hailong Sun 0001
CCF Trans. Pervasive Comput. Interact.3
2022 An error consistency based approach to answer aggregation in open-ended crowdsourcing
abstract
Crowdsourcing plays a vital role in today’s AI industry. However, existing crowdsourcing research mainly focuses on those simple tasks that are often formulated as label classification, while complex open-ended tasks such as question answering and translation have not received much attention. Such tasks usually have open solution spaces and non-unique true answers, which pose great challenges for designing effective crowdsourcing algorithms. In this work, we are concerned specifically with complex text annotation crowdsourcing tasks, where each answer of a task is in the form of free text. We propose an error consistency-based approach to inferring a satisfying result from a set of open-ended answers. First, each answer is represented with two vectors that capture the local word collocation and the global sentence semantics respectively. Second, the true answer is approximated by the sum of the answer vectors weighted by the reciprocals of their respective errors. Third, an algorithm called AEC (Aggregation based on Error Consistency) is designed to infer the aggregated result by maximizing the consistency of the errors of an answer in two vector spaces. Experimental results on two datasets demonstrate the effectiveness of our approach.
Lei Chai, Hailong Sun 0001, Zizhe Wang
Inf. Sci.2
2022 DreamLoc: A Deep Relevance Matching-Based Framework for bug Localization
abstract
To improve the software debugging efficiency, bug localization techniques have been developed to automatically locate buggy files based on bug reports. Traditional information retrieval-based bug localization cannot deal with the lexical mismatch, thus its performance is limited. In recent years, some deep learning models have been proposed to learn the semantics of bug reports and source files to bridge the lexical gap. However, their accuracy is still limited as building accurate semantic representations of bug reports and source files is very challenging. Recently, relevance matching was proposed to identify whether a document is relevant to a given query by considering both local matching and global matching. In this work, we propose a novel framework DreamLoc, which utilizes a relevance matching model to locate buggy files. Specifically, DreamLoc conducts the local matching by employing an attention-based mechanism to calculate the matching scores between bug report terms and code snippets. It also conducts the global matching by employing a gating mechanism to aggregate results of local matching and obtain the final matching score between a bug report and a source file. Since the local matching considers the relevance between each word and the global matching differentiates the importance of words, DreamLoc can effectively model the characteristics of bug reports and source files. Experimental results on five benchmark datasets show that DreamLoc outperforms five state-of-the-art models. For example, compared with DeepLoc, a recently proposed approach, the evaluation measures Accuracy@10, MAP, and MRR are improved by 6.4%, 7.4%, and 7.2%, respectively.
Binhang Qi, Hailong Sun 0001, Wei Yuan 0011, Hongyu Zhang 0002, Xiangxin Meng
IEEE Trans. Reliab.2
2021 Teaching Active Human Learners
Zizhe Wang, Hailong Sun 0001
AAAI2
2021 Incorporating Multiple Features to Predict Bug Fixing Time with Neural Networks
abstract
Debugging is a well-known time-consuming task, and knowing how long it would take to resolve bugs is of great importance for allocating the limited resources in a software development team. However, it is challenging to predict bug fixing time since fixing bugs is subject to a plethora of uncertain factors such as types of bugs, program complexity and developers' abilities. Existing work mainly focuses on developers' activities in a bug lifecycle and ignores other important factors. In light of the limitations of existing work, we propose a novel approach to predicting the bug fixing time by incorporating a comprehensive set of relevant features. Specifically, we consider four types of features including developers' activities, developers' sentiments, semantics of bugs, and efforts caused by understanding and analyzing source code, and design particular neural networks to take advantage of these features and make them work efficiently. Experimental results on four real application datasets demonstrate that on the one hand, our approach outperforms the state-of-the-art by over 5% in accuracy and 7.3% in F1-score on average; on the other hand, each type of the features considered in our approach plays an important part in the prediction.
Wei Yuan 0011, Yuan Xiong, Hailong Sun 0001, Xudong Liu 0001
ICSME3
2021 Model learning: a survey of foundations, tools and applications
Shahbaz Ali, Hailong Sun 0001, Yongwang Zhao
Frontiers Comput. Sci.2
2021 Find truth in the hands of the few: acquiring specific knowledge with crowdsourcing
Tao Han 0003, Hailong Sun 0001, Yangqiu Song, Yili Fang, Xudong Liu 0001
Frontiers Comput. Sci.2
2020 DependLoc: A Dependency-based Framework For Bug Localization
abstract
As software systems are becoming larger and more complex, debugging poses great challenges to software developers and maintainers. Among various efforts on easing the burden of debugging, bug localization techniques are developed to help locate where a bug occurs in source code files automatically. Information retrieval and deep neural network techniques are often adopted in existing research to achieve bug localization through capturing the textual or semantic similarity between bug reports and source code files. At the same time, some domain- specific eatures in software engineering are also utilized to locate the buggy files. However, the dependency relationship between classes (A depends on$B$if$A$references B) is not considered or utilized by existing approaches. In this work, we propose a novel framework DependLoc for bug localization which leverages the dependency relationship among source code files. DependLoc is based on the observation that buggy files may not be highly similar to a bug report but have a dependency relationship with one or more files that are quite similar to the bug report. DependLoc adopts a customized Ant Colony algorithm to quantify the intrinsic dependency relationship (called reference heat) and designs a segment-based encoder to learn this feature. Experimental results on six widely-used benchmark datasets for bug localization show that our approach outperforms the state-of-the-art methods, and Accuracy@10 is improved by 4% on average.
Wei Yuan 0011, Binhang Qi, Hailong Sun 0001, Xudong Liu 0001
APSEC3
2020 Robust Adversarial Active Learning with a Novel Diversity Constraint
abstract
Active learning adopts an iterative process that prioritizes the labeling of the most informative samples. However, in many real-world applications, the training data usually contains Out-of-Distribution (OoD) samples that affect the robustness of the trained models. Unfortunately, most existing active learning approaches focus on the uncertainty measure, thus have a bias towards selecting OoD samples. In this paper, we propose a robust adversarial active learning method that performs well on datasets with OoD samples. First, we incorporate recent advances in adversarial networks into an active learning framework to select the samples that are most dissimilar to the labeled pool. Secondly, we design a novel loss function based on Earth-Mover (EM) distance, which makes the model training more stable. Moreover, we propose a novel diversity constraint learned from feature space that penalizes the OoD samples. Experimental evaluation results on the datasets of varying size demonstrate the effectiveness of our approach.
Chengbin Sun, Hailong Sun 0001, Xudong Liu 0001
IEEE BigData2
2020 Retrieval-based neural source code summarization
abstract
Source code summarization aims to automatically generate concise summaries of source code in natural language texts, in order to help developers better understand and maintain source code. Traditional work generates a source code summary by utilizing information retrieval techniques, which select terms from original source code or adapt summaries of similar code snippets. Recent studies adopt Neural Machine Translation techniques and generate summaries from code snippets using encoder-decoder neural networks. The neural-based approaches prefer the high-frequency words in the corpus and have trouble with the low-frequency ones. In this paper, we propose a retrieval-based neural source code summarization approach where we enhance the neural model with the most similar code snippets retrieved from the training set. Our approach can take advantages of both neural and retrieval-based techniques. Specifically, we first train an attentional encoder-decoder model based on the code snippets and the summaries in the training set; Second, given one input code snippet for testing, we retrieve its two most similar code snippets in the training set from the aspects of syntax and semantics, respectively; Third, we encode the input and two retrieved code snippets, and predict the summary by fusing them during decoding. We conduct extensive experiments to evaluate our approach and the experimental results show that our proposed approach can improve the state-of-the-art methods.
Jian Zhang 0087, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001
ICSE4
2020 Structured Probabilistic End-to-End Learning from Crowds
abstract
End-to-end learning from crowds has recently been introduced as an EM-free approach to training deep neural networks directly from noisy crowdsourced annotations. It models the relationship between true labels and annotations with a specific type of neural layer, termed as the crowd layer, which can be trained using pure backpropagation. Parameters of the crowd layer, however, can hardly be interpreted as annotator reliability, as compared with the more principled probabilistic approach. The lack of probabilistic interpretation further prevents extensions of the approach to account for important factors of annotation processes, e.g., instance difficulty. This paper presents SpeeLFC, a structured probabilistic model that incorporates the constraints of probability axioms for parameters of the crowd layer, which allows to explicitly model annotator reliability while benefiting from the end-to-end training of neural networks. Moreover, we propose SpeeLFC-D, which further takes into account instance difficulty. Extensive validation on real-world datasets shows that our methods improve the state-of-the-art.
Hailong Sun 0001, Tao Han 0003, Xudong Liu 0001, Jie Yang 0028
IJCAI3
2020 Best Answerers Prediction With Topic Based GAT In Q&A Sites
abstract
Q&A communities are playing an important role in online knowledge sharing, where a large number of users with various knowledge background make tremendous contributions to solving many technical problems based on crowd intelligence. However, as new questions are increasingly posted, it is a non-trivial issue to find a matching answerer for each question. As a result, many questions fail to receive satisfying answers in time. This paper addresses the problem by predicting the best answerer for the new question. Many existing efforts are devoted to predicting the best answerer mainly by calculating the textual similarity between questions and a user’s historical post documents. Some works consider other features, such as the similarity of tags between questions and users, the average quality of a user’s historical answers, and so on. But few works consider interaction within the community. In recent years, works that take account of the interaction between community items (such as GCN and GAT) have made considerable progress in graph mining tasks like item recommendation, node representation, node classification, and link prediction. This kind of graph mining method can easily leverage interactive information in the community and encode it in an easy-to-use way which is very helpful for downstream tasks such as recommendation. However, questions that need to be recommended to answerers are new coming ones and with no interaction with any other node in the community yet. How to make reasonable use of collaborative information to improve recommendation performance is a real challenge. In this paper, we use the interactive information between candidate answerers and combine text information to make our best answerer recommendation. There are two main parts, LDA(Latent Dirichlet Allocation) topic model is used to capture the text information and graph attention networks (GATs) for interaction. We evaluated our approach on a real dataset from Stack Exchange. The result shows that our approach outperforms all the baseline methods.
Yuexin Huang, Hailong Sun 0001
Internetware2
2020 Learning to Handle Exceptions
abstract
Exception handling is an important built-in feature of many modern programming languages such as Java. It allows developers to deal with abnormal or unexpected conditions that may occur at runtime in advance by using try-catch blocks. Missing or improper implementation of exception handling can cause catastrophic consequences such as system crash. However, previous studies reveal that developers are unwilling or feel it hard to adopt exception handling mechanism, and tend to ignore it until a system failure forces them to do so. To help developers with exception handling, existing work produces recommendations such as code examples and exception types, which still requires developers to localize the try blocks and modify the catch block code to fit the context. In this paper, we propose a novel neural approach to automated exception handling, which can predict locations of try blocks and automatically generate the complete catch blocks. We collect a large number of Java methods from GitHub and conduct experiments to evaluate our approach. The evaluation results, including quantitative measurement and human evaluation, show that our approach is highly effective and outperforms all baselines. Our work makes one step further towards automated exception handling.
Jian Zhang 0087, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Yanjun Pu, Xudong Liu 0001
ASE4
2020 Predicting Crowdsourcing Worker Performance with Knowledge Tracing
Zizhe Wang, Hailong Sun 0001, Tao Han 0003
KSEM (2)2
2020 Developer recommendation for Topcoder through a meta-learning based policy model
Hailong Sun 0001, Hongyu Zhang 0002
Empir. Softw. Eng.2
2020 How are distributed bugs diagnosed and fixed through system logs?
Wei Yuan 0011, Shan Lu 0001, Hailong Sun 0001, Xudong Liu 0001
Inf. Softw. Technol.3
2020 CONAN: A framework for detecting and handling collusion in crowdsourcing
Hailong Sun 0001, Yili Fang, Xudong Liu 0001
Inf. Sci.2
2019 A novel neural source code representation based on abstract syntax tree
abstract
Exploiting machine learning techniques for analyzing programs has attracted much attention. One key problem is how to represent code fragments well for follow-up analysis. Traditional information retrieval based methods often treat programs as natural language texts, which could miss important semantic information of source code. Recently, state-of-the-art studies demonstrate that abstract syntax tree (AST) based neural models can better represent source code. However, the sizes of ASTs are usually large and the existing models are prone to the long-term dependency problem. In this paper, we propose a novel AST-based Neural Network (ASTNN) for source code representation. Unlike existing models that work on entire ASTs, ASTNN splits each large AST into a sequence of small statement trees, and encodes the statement trees to vectors by capturing the lexical and syntactical knowledge of statements. Based on the sequence of statement vectors, a bidirectional RNN model is used to leverage the naturalness of statements and finally produce the vector representation of a code fragment. We have applied our neural network based source code representation method to two common program comprehension tasks: source code classification and code clone detection. Experimental results on the two tasks indicate that our model is superior to state-of-the-art approaches.
Jian Zhang 0087, Xu Wang 0007, Hongyu Zhang 0002, Hailong Sun 0001, Xudong Liu 0001
ICSE4
2018 Automatically Generating API Usage Patterns from Natural Language Queries
abstract
Automatically generating code from natural language query is a very promising but much challenging direction. Existing approaches either try to generate the whole code or only predict a small part of critical code elements such as API sequence. Meanwhile, API usage patterns, including APIs and API-related control-flow statements, have the moderate complexity, but can provide enough code framework information and are very helpful for developers to implement various functionalities. Therefore, in this work, we study the problem of generating API usage patterns, represent API usage patterns by one special constrained tree API-MCTree and design one new API-MCTree decoder for automatically transforming natural language queries to API usage patterns, which can leverage both the difference of control-flow statement types and the syntactic knowledge of API usage patterns. We evaluate our model with annotated code snippets in real Java projects collected from GitHub, and the experimental results show that our approach is effective and outperforms the related approaches.
Yanfei Tian, Xu Wang 0007, Hailong Sun 0001, Chunbo Guo, Xudong Liu 0001
APSEC3
2018 On the Cost Complexity of Crowdsourcing
abstract
Existing efforts mainly use empirical analysis to evaluate the effectiveness of crowdsourcing methods, which is often unreliable across experimental settings. Consequently, it is of great importance to study theoretical methods. This work, for the first time, defines the cost complexity of crowdsourcing, and presents two theorems to compute the cost complexity. Our theorems provide a general theoretical method to model the trade-off between costs and quality, which can be used to evaluate and design crowdsourcing algorithms, and characterize the complexity of crowdsourcing problems. Moreover, following our theorems, we prove a set of corollaries that can obtain existing theoretical results for special cases. We have verified our work theoretically and empirically.
Yili Fang, Hailong Sun 0001, Jinpeng Huai
IJCAI2
2018 A Hybrid Approach to Developer Recommendation Based on Multi-relationship
abstract
The development of open source community has brought changes to software engineering. Different from traditional software development, developers are free to participate in various projects, QAs and forums under the Internet environment. The quality of participants' contribution largely decides the schedule of projects and the quality of QA answers. However, finding a right participant from numerous website users can be challenging. Therefore, evaluation and recommendation of a person are particularly important. Nonetheless, the developers' multi-relationship is an important behaviour of open source community users, and needs more research work for describing developers' online activities. To this end, we establish a multi-relational network model that describes developers' online behaviours. The capability model and interest model are established for each developer. Furthermore, we utilize a hybrid approach that combines the content and the multi-relationship together to realize the recommendation of respondents for a question. Finally, we have conducted experiments on datasets of blogs, QAs and forum of CSDN. The result proves the effectiveness of our approach.
Yifei Da, Hailong Sun 0001, Xudong Liu 0001
Internetware2
2018 Profiling Developer Expertise across Software Communities with Heterogeneous Information Network Analysis
abstract
Knowing developer expertise is critical for achieving effective task allocation. However, it is of great challenge to accurately profile the expertise of developers over the Internet as their activities often disperse across different online communities. In this regard, the existing works either merely concern a single community, or simply sum up the expertise in individual communities. The former suffers from low accuracy due to incomplete data, while the latter impractically assumes that developer expertise is completely independent and irrelavant across communities. To overcome those limitations, we propose a new approach to profile developer expertise across software communities through heterogeneous information network (HIN) analysis. A HIN is first built by analyzing the developer activities in various communities, where nodes represent objects like developers and skills, and edges represent the relations among objects. Second, as random walk with restart (RWR) is known for its ability to capture the global structure of the whole network, we adopt RWR over the HIN to estimate the proximity of developer nodes and skill nodes, which essentially reflects developer expertise. Based on the data of 72,645 common users of GitHub and Stack Overflow, we conducted an empirical study and evaluated developer expertise using proposed approach. To evaluate the effect of our approach, we use the obtained expertise to estimate the competency of developers in answering the questions posted in Stack Overflow. The experimental results demonstrate the superiority of our approach over existing methods.
Jiafei Yan, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001, Xiaotao Song
Internetware2
2018 Personalized teammate recommendation for crowdsourced software developers
abstract
Most crowdsourced software development platforms adopt contest paradigm to solicit contributions from the community. To attain competitiveness in complex tasks, crowdsourced software developers often choose to work with others collaboratively. However, existing crowdsourcing platforms generally assume independent contributions from developers and do not provide effective support for team formation. Prior studies on team recommendation aim at optimizing task outcomes by recommending the most suitable team for a task instead of finding appropriate collaborators for a specific person. In this work, we are concerned with teammate recommendation for crowdsourcing developers. First, we present the results of an empirical study of Kaggle, which shows that developers’personal teammate preferences are mainly affected by three factors. Second, we give a collaboration willingness model to characterize developers’ teammate preferences and formulate teammate recommendation as an optimization problem. Then we design a heuristic algorithm to find suitable teammates for a developer. Finally, we have conducted a set of experiments on a Kaggle dataset to evaluate the effectiveness of our approach.
Luting Ye, Hailong Sun 0001, Xu Wang 0007, Jiaruijue Wang
ASE2
2018 Context-aware result inference in crowdsourcing
Yili Fang, Hailong Sun 0001, Guoliang Li 0001, Richong Zhang, Jin-Peng Huai
Inf. Sci.2
2018 Collusion-Proof Result Inference in Crowdsourcing
Hailong Sun 0001, Yili Fang, Jinpeng Huai
J. Comput. Sci. Technol.2
2017 Budgeted Task Scheduling for Crowdsourced Knowledge Acquisition
abstract
Knowledge acquisition (e.g. through labeling) is one of the most successful applications in crowdsourcing. In practice, collecting as specific as possible knowledge via crowdsourcing is very useful since specific knowledge can be generalized easily if we have a knowledge base, but it is difficult to infer specific knowledge from general knowledge. Meanwhile, tasks for acquiring more specific knowledge can be more difficult for workers, thus need more answers to infer high-quality results. Given a limited budget, assigning workers to difficult tasks will be more effective for the goal of specific knowledge acquisition. However, existing crowdsourcing task scheduling cannot incorporate the specificity of workers' answers. In this paper, we present a new framework for task scheduling with the limited budget, targeting an effective solution to more specific knowledge acquisition. We propose novel criteria for evaluating the quality of specificity-dependent answers and result inference algorithms to aggregate more specific answers with budget constraints. We have implemented our framework with real crowdsourcing data and platform, and have achieved significant performance improvement compared with existing approaches.
Tao Han 0003, Hailong Sun 0001, Yangqiu Song, Zizhe Wang, Xudong Liu 0001
CIKM2
2017 Recommending crowdsourced software developers in consideration of skill improvement
abstract
Finding suitable developers for a given task is critical and challenging for successful crowdsourcing software development. In practice, the development skills will be improved as developers accomplish more development tasks. Prior studies on crowdsourcing developer recommendation do not consider the changing of skills, which can underestimate developers' skills to fulfill a task. In this work, we first conducted an empirical study of the performance of 74 developers on Topcoder. With a difficulty-weighted algorithm, we re-compute the scores of each developer by eliminating the effect of task difficulty from the performance. We find out that the skill improvement of Topcoder developers can be fitted well with the negative exponential learning curve model. Second, we design a skill prediction method based on the learning curve. Then we propose a skill improvement aware framework for recommending developers for software development with crowdsourcing.
Zizhe Wang, Hailong Sun 0001, Luting Ye
ASE2
2017 An efficient and highly available framework of data recency enhancement for eventually consistent data stores
Yu Tang 0018, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
Frontiers Comput. Sci.2
2017 Achieving convergent causal consistency and high availability for cloud storage
Yu Tang 0018, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
Future Gener. Comput. Syst.2
2017 Adaptive Result Inference for Collecting Quantitative Data With Crowdsourcing
abstract
In quantitative crowdsourcing, workers are asked to provide numerical answers. Different from categorical crowdsourcing, result aggregation in quantitative crowdsourcing is processed by combinatorially computing over all workers’ answers instead of by merely choosing one from a set of candidate answers. Therefore, existing result aggregation models for categorical crowdsourcing tasks cannot be used in quantitative crowdsourcing. Moreover, the worker ability often varies in the process of crowdsourcing with the changing of workers’ skill, willingness, efforts, etc. In this paper, we propose a probabilistic model to characterize the quantitative crowdsourcing problem by considering the changing of worker ability so as to achieve better quality control. The dynamic worker ability is obtained with Kalman filtering and smoother. We design an expectation-maximization-based inference algorithm and a dynamic worker filtering algorithm to compute the aggregated crowdsourcing result. Finally, we conducted experiments with real data on CrowdFlower and the results showed that our approach can effectively rule out low-quality workers dynamically and obtain more accurate results with less costs.
Hailong Sun 0001, Kefan Hu, Yili Fang, Yangqiu Song
IEEE Internet Things J.1
2017 Handling multi-dimensional complex queries in key-value data stores
Hailong Sun 0001, Yu Tang 0018, Xudong Liu 0001
Inf. Syst.1
2017 Improving the Quality of Crowdsourced Image Labeling via Label Similarity
Yili Fang, Hailong Sun 0001, Ting Deng
J. Comput. Sci. Technol.2
2017 Intelligent Development Environment and Software Knowledge Graph
Zeqi Lin, Yanzhen Zou, Junfeng Zhao 0001, Xuandong Li, Jun Wei 0001, Hailong Sun 0001, Gang Yin
J. Comput. Sci. Technol.7
2017 Adaptive trade-off between consistency and performance in data replication
abstract
Summary Replication is widely adopted in modern Internet applications and distributed systems to improve the reliability and performance. Though maintaining the strong consistency among replicas can guarantee the correctness of application behaviors, however, it will affect the application performance at the same time because there is a well‐known trade‐off between consistency and performance. Many real‐world applications favoring performance often choose to enforce weak consistency. Although there has been some work on flexible configuration of consistency, most focuses on design or deployment time. As the system settings constantly change during runtime, the tuning of the consistency‐performance trade‐off needs to be handled dynamically. Failing to do that will cause either underestimation or overestimation of the consistency and performance that can be achieved. Existing work does not well support the dynamic tuning of the aforementioned trade‐off in runtime, which is mainly because of the lack of an appropriate quantitative model of consistency and performance. In this work, based on our previous effort on the quantitative model of consistency and latency, we design a replication protocol, CC‐Paxos, to achieve an adaptive trade‐off between consistency and performance according to application preferences and runtime information. By design, CC‐Paxos is not bound to any specific underlying data stores. We have implemented CC‐Paxos and applied it to MySQL databases. And real experiments both within a data center and across data centers show that CC‐Paxos not only can dynamically adjust the delivered consistency in return for ensured performance but also outperforms MySQL Cluster in the case of strong consistency guarantee. Copyright © 2016 John Wiley & Sons, Ltd.
Hailong Sun 0001, Bang Xiao, Xu Wang 0007, Xudong Liu 0001
Softw. Pract. Exp.1
2016 Recommendflow: Use Topic Model to Automatically Recommend Stack Overflow Q&A in IDE
Fumin Sun, Xu Wang 0007, Hailong Sun 0001, Xudong Liu 0001
CollaborateCom3
2016 Effective Result Inference for Context-Sensitive Tasks in Crowdsourcing
Yili Fang, Hailong Sun 0001, Guoliang Li 0001, Richong Zhang, Jinpeng Huai
DASFAA (1)2
2016 Incorporating External Knowledge into Crowd Intelligence for More Specific Knowledge Acquisition
Tao Han 0003, Hailong Sun 0001, Yangqiu Song, Yili Fang, Xudong Liu 0001
IJCAI2
2016 Achieving convergent causal consistency and high availability with asynchronous replication
abstract
Nowadays, distributed data stores have become a fundamental infrastructure for large-scale Internet services, and they usually replicate data partitions to achieve high scalability and availability. To achieve better performance and availability, many Internet services embrace eventual consistency. However, stronger consistency is always desirable for system correctness. Recent studies [1][2] pay more attention to the convergent causal consistency, which is proved to be one of the strongest consistency models that can be achieved together with high availability in the presence of network partitions [3]. Convergent causal consistency couples the virtues of causal consistency and eventual consistency. As a result, convergent causal consistency not only guarantees that clients observe causality throughout, but also ensures that all replicas converge to the same state, which are critical for implementing reasonable application behaviors.
Yu Tang 0018, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001, Zhenglin Xia
IWQoS2
2016 Don't Get Caught in the Cold, Warm-up Your JVM: Understand and Eliminate JVM Warm-up Overhead in Data-Parallel Systems
David Lion, Adrian Chiu, Hailong Sun 0001, Xin Zhuang, Nikola Grcevski, Ding Yuan 0004
OSDI3
2016 Recommender systems based on ranking performance optimization
Richong Zhang, Han Bao 0005, Hailong Sun 0001, Yanghao Wang, Xudong Liu 0001
Frontiers Comput. Sci.3
2015 Spectral Label Refinement for Noisy and Missing Text Labels
abstract
With the recent growth of online content on the Web, there have been more user generated data with noisy and missing labels, e.g., social tags and voted labels from Amazon's Mechanical Turks. Most of machine learning methods, which require accurate label sets, could not be trusted when the label sets were yet unreliable. In this paper, we provide a text label refinement algorithm to adjust the labels for such noisy and missing labeled datasets. We assume that the labeled sets can be refined based on the labels with certain confidence, and the similarity between data being consistent with the labels. We propose a label smoothness ratio criterion to measure the smoothness of the labels and the consistency between labels and data. We demonstrate the effectiveness of the label refining algorithm on eight labeled document datasets, and validate that the results are useful for generating better labels.
Yangqiu Song, Chenguang Wang 0001, Ming Zhang 0004, Hailong Sun 0001, Qiang Yang 0001
AAAI4
2015 Combining Machine Learning and Crowdsourcing for Better Understanding Commodity Reviews
abstract
In e-commerce systems, customer reviews are important information for understanding market feedbacks on certain commodities. However, accurate analyzing reviews is challenging due to the complexity of natural language processing and informal descriptions in reviews. Existing methods mainly focus on studying efficient algorithms that cannot guarantee the accuracy for review analysis. Crowdsourcing can improve the accuracy of review analysis while it is subject to extra costs and low response time. In this work, we combine machine learning and crowdsourcing together for better understanding customer reviews. First, we collectively use multiple machine learning algorithms to pre-process review classification. Second, we select the reviews on which all machine learning algorithms cannot agree and assign them to humans to process. Third, the results from machine learning and crowdsourcing are aggregated to be the final analysis results. Finally, we perform real experiments with practical review data to confirm the effectiveness of our method.
Heting Wu, Hailong Sun 0001, Yili Fang, Kefan Hu, Yongqing Xie, Yangqiu Song, Xudong Liu 0001
AAAI2
2015 Truthful Incentive Mechanisms for Dynamic and Heterogeneous Tasks in Mobile Crowdsourcing
abstract
Crowdsourcing has received tremendous attention for collecting various data with the distributed smartphones of people. For the mobile crowdsourcing applications to obtain high-quality data, stimulating user participation is of paramount importance. Although many incentive mechanisms have been designed, most of them ignore the dynamic arrivals and different sensing requirements of tasks. Thus, the existing mechanisms will fail when being applied to the realistic scenario where tasks are publicized dynamically and heterogeneous with different sensing requirements of locations, time durations and sensing times. In this work, we propose two auction-based truthful mechanisms, TRIMS and TRIMG, for realistic mobile crowdsourcing under special user model and more general model, respectively. Through extensive simulations and theoretical analysis, we demonstrate that our mechanisms can satisfy the desired properties of truthfulness, individual rationality, computational efficiency with both low social cost and low total payment.
Hailong Sun 0001, Xudong Liu 0001
ICTAI2
2015 Efficient Testing of Web Services with Mobile Crowdsourcing
abstract
Nowadays, online Internet services are pervasive and can be invoked from diverse locations in anytime with multitudinous devices. Conventional testing approaches for online services like Web services are conducted by professional tester or developers and cannot simulate the real world running environment of a service. Fortunately, crowdtesting technology brings us promising hope and has acquired increasing interests and adoption because it can recruit plenty of end users to test services under real world environment with low cost. Meanwhile, improved mobile network techniques make crowdsourcing happen anywhere and anytime. In this paper, we present iTest which combines mobile crowdsourcing and web service testing together to support the performance testing of web services. iTest is a framework for service developers to submit their web services and conveniently get the test results from the crowd testers. Firstly, we analyze the key problems need to be solved in a mobile crowdtesting platform; secondly, the architecture of iTest framework and the workflow in it are presented; Thirdly, we perform experiments to illustrate that both the way to access network and tester's location influence the performance of web service, and formulate the tester selection problem as a Set Cover Problem and propose a greedy algorithm for solving this problem; Next, experimental evaluation of the tester selection algorithm is performed to illustrate its efficiency. Finally, we conclude our work and provide the directions for future work.
Minzhi Yan, Hailong Sun 0001, Xudong Liu 0001
Internetware2
2015 Combining Crowd Contributions with Machine Learning to Detect Malicious Mobile Apps
abstract
Android is undoubtedly becoming the most popular smartphone platform. The popularity of Android, unfortunately, has also made the devices become the target of malware. Most of existing malicious mobile apps feature stealthy operations such as collecting user privacy, sending premium SMS messages and making unauthorized http connections with no legal notice to the affected user. However, transmission of sensitive data cannot indicate malicious behavior because some benign applications also need sensitive data to improve the user experience. Existing malware detection approaches focus on static or dynamic analysis without crowd user contributions. In this paper, we propose a novel technique which combining crowd contributions with machine learning to detect malicious mobile apps. We model privacy transmission as user-determined and undetermined with the help of real user decisions based on crowdsourcing. We apply static analysis to extract application basic information such as permissions and suspicious API calls. Then we use dynamic instrumentation technique to trace real API calls at runtime and collect the crowd user decisions to the prompted sensitive data transmission. Finally, we employ several different learning-based algorithms, such as SVM, Bayesian Network, Decision Tree and KNN to detect malicious apps. Experiments with 100 real application samples show that our system was capable of detecting malicious mobile apps: our system can detect 85% to 97% of the malware with low false positive rate.
Dahai Yao, Hailong Sun 0001, Xudong Liu 0001
Internetware2
2015 Poster: TRIM: A Truthful Incentive Mechanism for Dynamic and Heterogeneous Tasks in Mobile Crowdsensing
abstract
Stimulating user participation is of paramount importance for mobile crowdsensing applications to obtain high-quality data. Although many incentive mechanisms have been designed, most of them ignore the dynamic arrivals and different sensing requirements of tasks. Thus, the existing mechanisms will fail when being applied to the realistic scenario where tasks are publicized dynamically and heterogeneous with different sensing requirements of locations, time durations and sensing times. In this work, we propose an auction-based truthful mechanism for realistic mobile crowdsensing. Through extensive simulations, we demonstrate that our mechanism can satisfy the desired properties of truthfulness, individual rationality, computational efficiency with both low social cost and low total payment.
Hailong Sun 0001, Xudong Liu 0001
MobiCom2
2015 On the tradeoff of availability and consistency for quorum systems in data center networks
Xu Wang 0007, Hailong Sun 0001, Ting Deng, Jinpeng Huai
Comput. Networks2
2015 Delivering Web service load testing as a service with a global cloud
abstract
Summary In this paper, we present WS‐TaaS, a Web services load testing platform built on a global platform PlanetLab. WS‐TaaS enables load testing process to be simple, transparent, and as close as possible to the real running scenarios of the target services. First, we briefly introduce the base of WS‐TaaS, Service4All. Second, we provide detailed analysis of the requirements of Web service load testing and present its conceptual architecture as well as algorithm design for improving resource utilization. Third, we present the implementation details of WS‐TaaS. Finally, we perform the evaluation of WS‐TaaS with a set of experiments based on the testing of real Web services, and the results illustrate that WS‐TaaS can efficiently facilitate the whole process of Web service load testing. Especially, comparing with existing testing tools, WS‐TaaS can obtain more effective and accurate test results. Copyright © 2014 John Wiley & Sons, Ltd.
Minzhi Yan, Hailong Sun 0001, Xudong Liu 0001, Ting Deng, Xu Wang 0007
Concurr. Comput. Pract. Exp.2
2014 A Model for Aggregating Contributions of Synergistic Crowdsourcing Workflows
abstract
One of the most important crowdsourcing topics is to study the effective quality control methods so as to reduce the cost and to guarantee the quality of task processing. As an effective approach, iterative improvement workflow is known to choose the best result from multiple workflows. However, for complex crowdsourcing tasks that consists of a certain number of subtasks under some specific constraints, but cannot be split into subtasks to be crowdsourced, the approach merely considers the best workflow without integrating the contributions of all workflows, which potentially results in extra costs for more iterations. In this paper, we propose an assembly model to integrate the best output of subtasks from different workflows. Moreover, we devise an efficient iterative method based on POMDP to improve the quality of assembled output. Empirical studies confirms the superiority of our proposed model.
Yili Fang, Hailong Sun 0001, Richong Zhang, Jinpeng Huai, Yongyi Mao
AAAI2
2014 SPKV: A Multi-dimensional Index System for Large Scale Key-Value Stores
Hailong Sun 0001, Yu Tang 0018, Xudong Liu 0001
APWeb2
2014 HARP: Towards enhancing data recency for eventually consistent data stores
abstract
To attain high performance and remain available during network partitions or node failures, modern distributed systems often sacrifice recency guarantees, which can provide a uniform view on recent versions of data items for different clients. In this work, we consider the problem of increasing the probability of data recency while preserving low response latency and maintaining high availability on top of an eventually consistent data store. To solve the problem, we propose HARP, an approach that can enhance data recency in a highly available way. Based on HARP, we implement an agent layer to detect stale reads and resolve the conflicts, and by leveraging widely deployed data store technologies, we build a data storage system. We compare the prototype system to Cassandra, and experimentally prove that our method produces low overhead (less than 10%) based on the eventually consistent configuration and, for most workloads, achieves better performance than the Cassandra's strong “read your writes” configurations.
Yu Tang 0018, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
ICPADS2
2014 A Novel Approach for API Recommendation in Mashup Development
abstract
Mashing up Web services and RESTful APIs is a novel programming approach to develop new applications. As the number of available resources is increasing rapidly, to discover potential services or APIs is getting difficult. Therefore, it is vital to relieve mashup developers of the burden of service discovery. In this paper, we propose a probabilistic model to assist mashup creators by recommending a list of APIs that may be used to compose a required mashup given descriptions of the mashup. Specifically, a relational topic model is exploited to characterize the relationship among mashups, APIs and their links. In addition, we incorporate the popularity of APIs to the model and make predictions on the links between mashups and APIs. Moreover, the statistical analysis on a public mashup platform shows the current status of mashup development and the applicability of this study. Experiments on a large service data set confirm the effectiveness of this proposed approach.
Chune Li, Richong Zhang, Jinpeng Huai, Hailong Sun 0001
ICWS4
2014 Incorporating Invocation Time in Predicting Web Service QoS via Triadic Factorization
abstract
With the development of Service-Oriented technologies, the amount of Web services grows rapidly. QoS-Aware Web service recommendation can help service users to design more efficient service-oriented systems. However, existing methods assume the QoS information for service users are all known and accurate, but in real case, there are always many missing QoS values in history records, which increase the difficulty of the missing QoS value prediction. By considering the user-service-time three dimension context information, we study a Temporal QoS-Aware Web Service Prediction Framework which aims to recommend best candidates to service user's requirements and meanwhile improve the QoS prediction accuracy. One major challenge is that how to deal with the high dimension, sparse QoS value data. Tensor which is known as multi-way array provides a natural representation for such QoS value data. Therefore, we formalize this problem as a tensor factorization model and propose a Tucker Decomposition (TD) algorithm which is able to deal with the triadic relations of user-service-time model. Extensive experiments are conducted based on our real-world QoS dataset collected on Planet-Lab, comprised of service invocation response-time values from 408 users on 5,473 Web services at 56 time periods. Comprehensive empirical studies demonstrate that our approach is more accuracy than other approaches and achieves 100X to 1000X memory space reduction.
Wancai Zhang, Hailong Sun 0001, Xudong Liu 0001, Xiaohui Guo
ICWS2
2014 A quantitative analysis of quorum system availability in data centers
abstract
Large-scale distributed storage systems often replicate data across servers and even geographically-distributed data centers for high availability, while existing theories like CAP and PACELC show that there is a tradeoff between availability and consistency. However, current practice is mainly experience-based and lacks quantitative analysis for identifying a good tradeoff between the two. In this work, we are concerned with providing a quantitative analysis on availability for widely-used quorum systems in data centers. First, a probabilistic model is presented to quantify availability for typical data center networks: 2-tier basic tree, 3-tier basic tree, fat tree and folded clos network. Second, we build the availability-consistency table and propose a set of rules to quantitatively make tradeoff between availability and consistency. Finally, with Monte Carlo based simulations, we validate our presented quantitative results and show that our approach to make tradeoff between availability and consistency is effective.
Xu Wang 0007, Hailong Sun 0001, Ting Deng, Jinpeng Huai
IWQoS2
2014 Poster: a framework for instant mobile web browsing with smart prefetching and caching
abstract
Mobile users often suffer from a slow page loading time due to intermittently connected wireless networks and increasing size of mobile Web pages. Although existing approaches of prefetching and caching are widely used to reduce the browsing latency, they may fail to work effectively because most Web page visits are singletons and the cache hit ratio is unsatisfying. In this work, we design and implement a framework to reduce Web browsing latency for mobile users with a smart prefetching strategy and caching mechanism. The prefetching strategy leverages the skSLRU model, which predicts and prefetches Web pages based on their contents with consideration of user contexts and the devices' status such as power consuming and cellular data usage. And our caching mechanism mainly consider the resources like CSS and JavaScript files shared among Web pages in a website. Moreover, instead of using RAM, we use ROM, the internal flash memory of devices to store cached resources with proper lifecycle so as to avoid the useful resources to be untimely evicted or expired and thus improve the hit ratio. Our evaluations show that 90% of the homepages and 60% of other pages are fetched before users' visiting, and the page loading time is no more than one second.
Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
MobiCom2
2014 AdaMF: Adaptive Boosting Matrix Factorization for Recommender System
Yanghao Wang, Hailong Sun 0001, Richong Zhang
WAIM2
2014 Discovering Semantic Mobility Pattern from Check-in Data
Ji Yuan, Xudong Liu 0001, Richong Zhang, Hailong Sun 0001, Xiaohui Guo, Yanghao Wang
WISE (1)4
2014 Temporal QoS-aware web service recommendation via non-negative tensor factorization
abstract
With the rapid growth of Web Service in the past decade, the issue of QoS-aware Web service recommendation is becoming more and more critical. Since the Web service QoS information collection work requires much time and effort, and is sometimes even impractical, the service QoS value is usually missing. There are some work to predict the missing QoS value using traditional collaborative filtering methods based on user-service static model. However, the QoS value is highly related to the invocation context (e.g., QoS value are various at different time). By considering the third dynamic context information, a Temporal QoS-aware Web Service Recommendation Framework is presented to predict missing QoS value under various temporal context. Further, we formalize this problem as a generalized tensor factorization model and propose a Non-negative Tensor Factorization (NTF) algorithm which is able to deal with the triadic relations of user-service-time model. Extensive experiments are conducted based on our real-world Web service QoS dataset collected on Planet-Lab, which is comprised of service invocation response-time and throughput value from 343 users on 5817 Web services at 32 time periods. The comprehensive experimental analysis shows that our approach achieves better prediction accuracy than other approaches.
Wancai Zhang, Hailong Sun 0001, Xudong Liu 0001, Xiaohui Guo
WWW2
2013 Consistency or latency? A quantitative analysis of replication systems based on replicated state machines
abstract
Existing theories like CAP and PACELC have claimed that there are tradeoffs between some pairs of performance measures in distributed replication systems, such as consistency and latency. However, current systems take a very vague view on how to balance those tradeoffs, e.g. eventual consistency. In this work, we are concerned with providing a quantitative analysis on consistency and latency for widely-used replicated state machines(RSMs). Based on our presented generic RSM model called RSM-d, probabilistic models are built to quantify consistency and latency. We show that both are affected by d, which is the number of ACKs received by the coordinator before committing a write request. And we further define a payoff model through combining the consistency and latency models. Finally, with Monte Carlo based simulation, we validate our presented models and show the effectiveness of our solutions in terms of how to obtain an optimal tradeoff between consistency and latency.
Xu Wang 0007, Hailong Sun 0001, Ting Deng, Jinpeng Huai
DSN2
2013 Towards a Scalable PaaS for Service Oriented Software
abstract
Software developers with service oriented technologies usually put a lot of efforts to deploy and manage supporting middleware and tools. Meanwhile PaaS in cloud computing aims at provide efficient support for software developers. In the light of this consideration, we have designed and implemented Service4All, a service cloud platform targeting at improve productivity of service oriented software developers. In this paper, we describe the design of SAE, a key component in Service4All, in terms of scalability. First, we present the key technical issues and architecture design of SAE. Second, we describe a software appliance based mechanism for elastic middleware management. Third, we describe a micro-kernel based AppEngine core for efficient coordination of various components. Finally, through a real application deployed on Service4All, we demonstrate the effectiveness of our solution.
Hailong Sun 0001, Xu Wang 0007, Minzhi Yan, Yu Tang 0018, Xudong Liu 0001
ICPADS1
2013 A cooperation model towards the internet of applications
abstract
As Internet is changing from a network of data into a network of functionalities, a federated Internet of applications is a natural trending topic, where every application can cooperate with each other smoothly to serve users. Cooperation can be regarded as integrating applications of different providers for users. However, existing integration techniques do not pay enough attention to multiple participants including application providers and end-users. In this study, we advocate a global cooperation model of Internet of applications for all the participants. Specifically, we propose an intermediary based model to realize the cooperation among applications. With this model, on the one hand, users can be greatly facilitated to cooperatively use applications from various providers to meet their individualized requirements; on the other hand, providers can easily enable their own applications to interact with those from other providers. Thus, the federated Internet of applications is easier to be achieved than using existing solutions. In addition, we implement the model and show some case studies which demonstrate the effectiveness of this model. In our vision, such a model is the beginning to bring the socialized relationships behind applications into the digital world.
Hailong Sun 0001, Richong Zhang, Xudong Liu 0001, Peilong Xu
Internetware2
2013 Discovering User Preference from Folksonomy
abstract
The increasing availability of socially shared media with tags annotated makes it vital for retrieval approaches to precisely detect web content topic semantic and better understand user interest. Most existing methodologies process the queries merely considering user posted keywords and retrieve media labeled with tags that are similar to query words, while ignoring users implicit interests and preferences. This fact stimulates us to develop preference discovering models to reveal the users' latent intents. In this paper, we study the problem of finding user preference and interest from folksonomy corpus and propose a preference-topic model that exploits probabilistic graphical model and Gibbs sampling algorithm to infer the user interested latent semantic topics. The experimental results show that, with the help of the proposed model, preference topics of the web content creators can be effectively discovered. In addition, two exemplified applications are discussed briefly.
Xiaohui Guo, Richong Zhang, Jinpeng Huai, Hailong Sun 0001, Xudong Liu 0001
SMC4
2013 Time-Aware Travel Attraction Recommendation
Richong Zhang, Xudong Liu 0001, Xiaohui Guo, Hailong Sun 0001, Jinpeng Huai
WISE (1)5
2013 GOS: a global optimal selection strategies for QoS-aware web services composition
Mu Li 0004, Danfeng Zhu, Ting Deng, Hailong Sun 0001, Huipeng Guo, Xudong Liu 0001
Serv. Oriented Comput. Appl.4
2013 Personalized QoS-Aware Web Service Recommendation and Visualization
abstract
With the proliferation of web services, effective QoS-based approach to service recommendation is becoming more and more important. Although service recommendation has been studied in the recent literature, the performance of existing ones is not satisfactory, since (1) previous approaches fail to consider the QoS variance according to users' locations; and (2) previous recommender systems are all black boxes providing limited information on the performance of the service candidates. In this paper, we propose a novel collaborative filtering algorithm designed for large-scale web service recommendation. Different from previous work, our approach employs the characteristic of QoS and achieves considerable improvement on the recommendation accuracy. To help service users better understand the rationale of the recommendation and remove some of the mystery, we use a recommendation visualization technique to show how a recommendation is grouped with other choices. Comprehensive experiments are conducted using more than 1.5 million QoS records of real-world web service invocations. The experimental results show the efficiency and effectiveness of our approach.
Zibin Zheng, Xudong Liu 0001, Hailong Sun 0001
IEEE Trans. Serv. Comput.5
2012 Building a TaaS Platform for Web Service Load Testing
abstract
Web services are widely known as the building blocks of typical service oriented applications. The performance of such an application system is mainly dependent on that of component web services. Thus the effective load testing of web services is of great importance to understand and improve the performance of a service oriented system. However, existing Web Service load testing tools ignore the real characteristics of the practical running environment of a web service, which leads to inaccurate test results. In this work, we present WS-TaaS, a load testing platform for web services, which enables load testing process to be as close as possible to the real running scenarios. In this way, we aim at providing testers with more accurate performance testing results than existing tools. WS-TaaS is developed on the basis of our existing Cloud PaaS platform: Service4All. First, we provide detailed analysis of the requirements of Web Service load testing and present the conceptual architecture and design of key components. Then we present the implementation details of WS-TaaS on the basis of Service4All. Finally, we perform a set of experiments based on the testing of real web services, and the experiments illustrate that WS-TaaS can efficiently facilitate the whole process of Web Service load testing.
Minzhi Yan, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
CLUSTER2
2012 WS-TaaS: A Testing as a Service Platform for Web Service Load Testing
abstract
Web services are widely known as the building blocks of typical service oriented applications. The performance of such an application system is mainly dependent on that of component web services. Thus the effective load testing of web services is of great importance to understand and improve the performance of a service oriented system. However, existing Web Service load testing tools ignore the real characteristics of the practical running environment of a web service, which leads to inaccurate test results. In this work, we present WS-TaaS, a load testing platform for web services, which enables load testing process to be as close as possible to the real running scenarios. In this way, we aim at providing testers with more accurate performance testing results than existing tools. WS-TaaS is developed on the basis of our existing Cloud PaaS platform: Service4All. First, we briefly introduce the functionalities and main components of Service4All. Second, we provide detailed analysis of the requirements of Web Service load testing and present the conceptual architecture and design of key components. Third, we present the implementation details of WS-TaaS on the basis of Service4All. Finally, we perform a set of experiments based on the testing of real web services, and the experiments illustrate that WS-TaaS can efficiently facilitate the whole process of Web Service load testing. Especially, comparing with existing testing tools, WS-TaaS can obtain more effective and accurate test results.
Minzhi Yan, Hailong Sun 0001, Xu Wang 0007, Xudong Liu 0001
ICPADS2
2012 Rep4WS: A Paxos Based Replication Framework for Building Consistent and Reliable Web Services
abstract
Web services are widely used to enable remote access to heterogeneous resources through standard interfaces and build complex applications by reusing existing component services. However, massive commodity computers, storage, network devices and complex management tasks running behind web services make them subject to outage and unable to provide continuously reliable services. To address this issue, we present a Paxos-based replication framework for building consistent and reliable web services. The framework mainly consists of a replication protocol and a set of failure tackling algorithms. First, in the replication protocol, besides keeping consistency of service replicas we introduce pipeline concurrency and RDG (Request Dependency Graph) to traditional Paxos so as to improve its performance. Second, we design failure recovery algorithms to recover the failed nodes, which guarantee that each web service has enough available replicas and thus can deliver expected reliability. Third, through an extensive set of experiments, we show that our method is effective in terms of keeping consistency and reliability of web services and it outperforms other replication methods.
Xu Wang 0007, Hailong Sun 0001, Ting Deng, Jinpeng Huai
ICWS2
2012 gTravel: a global social travel system
abstract
This paper presents an global social travel system to assist tourists in their itinerary planning, tour navigation, and travel knowledge sharing. In particular, firstly, we propose an efficient and flexible itinerary planning algorithm for organizing itineraries. Secondly, we design an intelligent tour path planning and navigation system by mining patterns from trajectories and geo-photos shared by other tourists. Thirdly, we provide a framework to monitor the status and events that may affect the predefined itineraries. Finally, the social media module enhances the opportunities of experience (itinerary, trajectory, and travelogue) sharing and helps travelers make better decisions. Concrete demonstrations of our system are also provided in this paper to show the flexibility and applicability of our system.
Richong Zhang, Xiaohui Guo, Hailong Sun 0001, Jinpeng Huai, Xudong Liu 0001
ACM Multimedia3
2012 A Tourist Itinerary Planning Approach Based on Ant Colony Algorithm
Richong Zhang, Hailong Sun 0001, Xiaohui Guo, Jinpeng Huai
WAIM3
2012 Generating Tourism Path from Trajectories and Geo-Photos
Zhixing Zeng, Richong Zhang, Xudong Liu 0001, Xiaohui Guo, Hailong Sun 0001
WISE5
2010 SOARWare: A Service Oriented Software Production and Running Environment
abstract
Service oriented computing provides a novel approach to building new software applications through the reuse of existing services. In this paper, we present SOARWare, a suite of middleware and tools, for software production and running based on Web services technologies. Basically, SOARWare consists of three major components including SOARBase, Service Oriented Software Production Line and Service Running Bus. Additionally, SOARWare provides a web-based platform for various users to access to the system functionality in a SaaS manner. We depict the design principle, system architecture and major functions of SOARWare.
Hailong Sun 0001, Xudong Liu 0001
APWeb1
2010 A User-Oriented Approach to Assessing Web Service Trustworthiness
Weinan Zhao, Hailong Sun 0001, Xudong Liu 0001, Xitong Kang
ATC2
2010 RegionKNN: A Scalable Hybrid Collaborative Filtering Algorithm for Personalized Web Service Recommendation
abstract
Several approaches to web service selection and recommendation via collaborative filtering have been studied, but seldom have these studies considered the difference between web service recommendation and product recommendation used in e-commerce sites. In this paper, we present RegionKNN, a novel hybrid collaborative filtering algorithm that is designed for large scale web service recommendation. Different from other approaches, this method employs the characteristics of QoS by building an efficient region model. Based on this model, web service recommendations will be generated quickly by using modified memory-based collaborative filtering algorithm. Experimental results demonstrate that apart from being highly scalable, RegionKNN provides considerable improvement on the recommendation accuracy by comparing with other well-known collaborative filtering algorithms.
Xudong Liu 0001, Hailong Sun 0001
ICWS4
2010 An adaptive heuristic approach for distributed QoS-based service composition
abstract
QoS-based service selection becomes a commonly accepted procedure to support rapid and dynamic web service composition. In this paper, we study the problem of QoS-based service selection in distributed QoS management environments where QoS values of alternative services are maintained by distributed QoS registries. A distributed heuristic approach is proposed to solve the problem efficiently with a high approximation ratio, and enable adaptability in distributed cross-organization environments with data privacy protection and a low cost of communication. The proposed approach consists of four stages in which variable elimination is used to reduce the size of the problem; constraint decomposition allows performing service selection independently on each QoS registry; supplementary service selection and concentrated optimization improve the approximation ratio. Performance analyses and simulation experiments show that the proposed approach performs efficiently with close-to-optimal results and fits well to distributed QoS management environments.
Jing Li 0075, Yongwang Zhao, Min Liu 0017, Hailong Sun 0001, Dianfu Ma
ISCC4
2010 Improving Performance for Decentralized Execution of Composite Web Services
abstract
Decentralized orchestration of composite web services offers performance improvements in terms of increased throughput and lower response time. However, most relevant research literature in decentralizing service composition omit the step of selecting component services, the locations of which have major impact on the amount of dataflow messages and the volume of network traffic during decentralized execution. In order to reduce the network traffic generated by data flow messages thus to improve the performance of the decentralized orchestration, this paper presents an approach to component service selection using data dependency graphs. The performance evaluation concludes that a substantial decrease in network traffic by use of our service selection approach results in a reduction of approximately 30% in execution time on average.
Xitong Kang, Xudong Liu 0001, Hailong Sun 0001, Yanjiu Huang
SERVICES3
2010 Live Instance Migration with Data Consistency in Composite Service Evolution
abstract
Composite service evolution is one of the most important challenges to deal with in the field of service composition. And how to migrate live instances to the evolved definition is a critical issue for correct service evolution. However, most of current instance migration approaches only consider control flow problem while ignoring data flow correctness during migration. In particular, this paper presents a data consistent approach to instance migration. We first propose a data dependence graph model to represent the data flow information in a composite service. Then, we propose a set of compliance criteria which relax the traditional compliance notion in instance migration through analyzing the data flow of the composite service. In this way, the number of migratable instances is increased, reducing the cost of redoing the case or service compensation. Finally, we design the DCIM algorithm to implement the data consistent instance migration, and an extensive set of simulations are performed to evaluate the algorithm.
Jianing Zou, Xudong Liu 0001, Hailong Sun 0001
SERVICES3
2010 Zebroid: using IPTV data to support STB-assisted VoD content delivery
Yih-Farn Robin Chen, Rittwik Jana, Daniel Stern, Bin Wei 0003, Mike Yang, Hailong Sun 0001, Jagadeesh M. Dyaberi
Multim. Syst.6
2009 Project GeoTV - A Three-Screen Service: Navigate on SmartPhone, Browse on PC, Watch on HDTV
abstract
GeoTV is a project that explores seamless integration of mobile phones, HDTV sets, and computers in the living room to enrich the user experience of existing services. The three- screen service allows a user to navigate a world map on a smart phone to track geo-located media RSS content that matches her personal interests. The user can show a matching video clip on her phone or direct a nearby HDTV set to play the video. In addition, the user can bring up a world map on a nearby computer screen to navigate areas of interest related to the video clip. GeoTV allows all three screens to be used for what they are best for: HDTV for high resolution video, computer screen for browsing a world map, and a smart phone for personalized control at hand to select media of interest.
Yih-Farn Robin Chen, David C. Gibbon, Rittwik Jana, Bernard Renger, Daniel Stern, Mike Yang, Bin Wei 0003, Hailong Sun 0001
CCNC8
2009 VP2P: A Virtual Machine-Based P2P Testbed for VoD Delivery
abstract
Recent advances in P2P technology have made it a viable alternative for the delivery of rich media by both content and service providers. However, to better understand and utilize P2P technologies in building a scalable media distribution platform, the nature of the underlying network must be taken into consideration. Although many P2P simulations have been conducted previously, very few were evaluated on a real P2P testbed with the properties of the physical network in mind. In this paper, we describe VP2P, a virtual machine-based P2P testbed, which supports P2P studies with the consideration of typical service provider networks. We use four fully equipped MacPro's to emulate 32 peers in the current implementation. Experiments of three P2P algorithms for video on demand (VoD) services: BitTorrent, Toast, and Zebra, were run on the testbed. We discuss the rationales behind the hardware and software architecture of VP2P, report our results, and conclude with lessons learned and future directions for VM-based testbeds.
Yih-Farn Robin Chen, Rittwik Jana, Daniel Stern, Hailong Sun 0001, Bin Wei 0003, Mike Yang
CCNC4
2009 LiveMig: An Approach to Live Instance Migration in Composite Service Evolution
abstract
Composite service evolution is one of the most important challenges to deal with in the field of service composition. In particular, this paper presents, LiveMig, an approach to live migration of composite service instance, which is a critical step for online composite service evolution. In LiveMig, a set of change operations preserving soundness is first defined. Second, a live instance state migration algorithm is proposed to determine if the state migration is allowed or not, and to compute the exact state after the migration. Finally, the correctness of LiveMig is theoretically proved, and an extensive set of simulations are performed to show its feasibility and effectiveness.
Jinpeng Huai, Hailong Sun 0001, Ting Deng
ICWS3
2009 Zebroid: using IPTV data to support peer-assisted VoD content delivery
abstract
P2P file transfers and streaming have already seen a tremendous growth in Internet applications. With the rapid growth of IPTV, the need to efficiently disseminate large volumes of Video-on-Demand (VoD) content has prompted IPTV service providers to consider peer-assisted VoD content delivery. This paper describes Zebroid, a VoD solution that uses IPTV operational data on an on-going basis to determine how to pre-position popular contents in customer set-top boxes during idle hours to allow these peers to assist the VoD server in content delivery during peak hours. Latest VoD request distribution, set-top box availability, and capacity data on network components are all taken into consideration in determining the parameters used in the striping algorithm of Zebroid. We show both by simulation and emulation on a realistic IPTV testbed that the VoD server load can be significantly reduced by more than 50-80% during peak hours by using Zebroid.
Yih-Farn Robin Chen, Rittwik Jana, Daniel Stern, Bin Wei 0003, Mike Yang, Hailong Sun 0001
NOSSDAV6
2008 RCT: A distributed tree for supporting efficient range and multi-attribute queries in grid computing
Hailong Sun 0001, Jinpeng Huai, Yunhao Liu 0001, Rajkumar Buyya
Future Gener. Comput. Syst.1
2007 QoS-aware Service Composition in Service Overlay Networks
abstract
As the amount of Web services over the Internet grows continuously, these services can be interconnected to form a service overlay network (SON). On the basis of SON, building value-added services by service composition is an effective method to satisfy the changeable functional and non-functional QoS (quality of service) requirements of customers. However, the previous research on QoS- aware service composition in SON mainly focuses on the context where services have simple interactions, and it can not support application scenarios with complex business collaboration in electronic business. In this paper, we propose the HOSSON (hierarchical service composition framework in SON) framework, which can be used to construct more general-purpose SON through describing the relations among services using business protocols. In HOSSON, business protocols instead of interactive messages are adopted to simplify the description of service composition requirements and a novel approach named protocol computing are proposed to implement service composition on demand. Furthermore, two algorithms, OSS and MCSS, are designed to support service selection for QoS-aware service composition. Finally, comprehensive simulations are conducted to evaluate the performance of algorithms.
Jinpeng Huai, Ting Deng, Hailong Sun 0001, Huipeng Guo, Zongxia Du
ICWS4
2007 ROST: Remote and hot service deployment with trustworthiness in CROWN Grid
Jinpeng Huai, Hailong Sun 0001, Chunming Hu, Yanmin Zhu 0006, Yunhao Liu 0001, Jianxin Li 0002
Future Gener. Comput. Syst.2
2006 RCT: A Self-Adaptive Overlay for Efficient Computational Resource Discovery in Grid Systems
abstract
Computational resource discovery is of great importance in grid environments. Existing approaches do not consider the characteristics of application resource requirements. We propose, Resource Category Tree (RCT), which organizes computational resources based on their characteristics represented by primary attributes (PA). RCT adopts a structure of AVL tree, with each node representing a specific range of PA values. Though RCT adopts a hierarchical structure, it does not require nodes in higher levels maintain more information than those in lower levels, which makes RCT highly scalable. RCT is featured by self-organization, load-aware self-adaptation and fault tolerance. Based on RCT, commonly used queries, such as range queries and multi-attribute queries, are well supported. We conduct performance evaluations through comprehensive simulations.
Hailong Sun 0001, Jinpeng Huai, Gongwei Fu, Yunhao Liu 0001
e-Science1
2006 CROWN: A service grid middleware with trust management mechanism
Jinpeng Huai, Chunming Hu, Jianxin Li 0002, Hailong Sun 0001, Tianyu Wo
Sci. China Ser. F Inf. Sci.4
2005 OpenSPACE: An Open Service Provisioning and Consuming Environment for Grid Computing
abstract
Our key project, CROWN (China Research and Development Environment Over Wide-area Network), aims to empower in-depth integration of resources and cooperation of researchers nationwide and worldwide using grid technologies. It adopts service oriented architecture. Current service grids do not consider the separation of services and underlying resources, potentially causing low job processing efficiency and resource utilization. In this paper, we propose a novel architecture for grid systems, called Open Service Provisioning and Consuming Environment (OpenSPACE). OpenSPACE is adopted by CROWN and is evaluated by prototype implementation based on real applications.
Hailong Sun 0001, Jinpeng Huai, Yunhao Liu 0001
e-Science1