Haifeng Shen

dblp:39/4885 · DBLP profile ↗
← Back
82ranked-venue papers
16as first author
29since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 24 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 20 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 16 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-author · 1 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Theory of computation · 2 · 2 first-author
YearPublicationVenuePosition
2025 Code Comment Inconsistency Detection and Rectification Using a Large Language Model
abstract
Comments are widely used in source code. If a comment is consistent with the code snippet it intends to annotate, it would aid code comprehension. Otherwise, Code Comment Inconsistency (CCI) is not only detrimental to the understanding of code, but more importantly, it would negatively impact the development, testing, and maintenance of software. To tackle this issue, existing research has been primarily focused on detecting inconsistencies with varied performance. It is evident that detection alone does not solve the problem; it merely paves the way for solving it. A complete solution requires detecting inconsistencies and, more importantly, rectifying them by amending comments. However, this type of work is scarce. In this paper, we contribute C4RLLaMA, a fine-tuned large language model based on the open-source CodeLLaMA. It not only has the ability to rectify inconsistencies by correcting relevant comment content but also outperforms state-of-the-art approaches in detecting inconsistencies. Experiments with various datasets confirm that C4RLLaMA consistently surpasses both post hoc and just-in-time CCI detection approaches. More importantly, C4RLLaMA outperforms substantially the only known CCI rectification approach in terms of multiple performance metrics. To further examine C4RLLaMA's efficacy in rectifying inconsistencies, we conducted a manual evaluation, and the results showed that the percentage of correct comment updates by C4RLLaMA was 65.0% and 55.9% in just-in-time and post hoc, respectively, implying C4RLLaMA's real potential in practical use.
Guoping Rong, Yongda Yu, Haifeng Shen, Jidong Hu
ICSE6
2025 AUCAD: Automated Construction of Alignment Dataset from Log-Related Issues for Enhancing LLM-based Log Generation
abstract
Log statements have become an integral part of modern software systems.Prior research efforts have focused on supporting the decisions of placing log statements, such as where/what to log.With the increasing adoption of Large Language Models (LLMs) for coderelated tasks such as code completion or generation, automated approaches for generating log statements have gained much momentum.However, the performance of these approaches still has a long way to go.This paper explores enhancing the performance of LLM-based solutions for automated log statement generation by post-training LLMs with a purpose-built dataset.Thus the primary contribution is a novel approach called AUCAD, which automatically constructs such a dataset with information extracting from log-related issues.Researchers have long noticed that a significant portion of the issues in the open-source community are related to log statements.However, distilling this portion of data requires manual efforts, which is labor-intensive and costly, rendering it impractical.Utilizing our approach, we automatically extract logrelated issues from 1,537 entries of log data across 88 projects and identify 808 code snippets (i.e., methods) with retrievable source code both before and after modification of each issue (including log statements) to construct a dataset.Each entry in the dataset consists of a data pair representing high-quality and problematic log statements, respectively.With this dataset, we proceed to post-train multiple LLMs (primarily from the Llama series) for automated * Corresponding author.
Hao Zhang 0210, Dongjun Yu, Lei Zhang 0160, Guoping Rong, Yongda Yu, Haifeng Shen, He Zhang 0001, Dong Shao, Hongyu Kuang
Internetware6
2025 Automated detection of affected libraries from vulnerability reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Runfeng Mao, Lanxin Yang, Haifeng Shen
Autom. Softw. Eng.8
2025 DeepMetaIoT: A Multimodal Deep Learning Framework Harnessing Metadata for IoT Sensor Data Classification
abstract
Internet of Things (IoT) sensor data, which capture time series physical measurements such as temperature and humidity, often lack proper classification. This limits their effective understanding, integration, and reuse. While sensor metadata—textual descriptions of the measurements—is sometimes available, it is frequently incomplete or ambiguous. As a result, classification often depends solely on the time series data. Leveraging both time series sensor readings and textual metadata for automated and accurate classification remains a challenge due to the heterogeneity and inconsistency of these data sources. In this paper, we propose DeepMetaIoT, a multimodal deep learning framework that integrates time series and textual data for classification. DeepMetaIoT employs a cross-residual architecture comprising a time series encoder and a text encoder based on a pre-trained large language model, enabling effective fusion of both modalities. Experimental results on real-world IoT sensor datasets show that DeepMetaIoT consistently outperforms state-of-the-art machine learning and deep learning baselines.
Muhammad Sakib Khan Inan, Kewen Liao, Haifeng Shen, Prem Prakash Jayaraman, Federico Montori, Dimitrios Georgakopoulos 0001
IEEE Internet Things J.3
2025 DLAP: A Deep Learning Augmented Large Language Model Prompting framework for software vulnerability detection
Yanjing Yang, Xin Zhou 0016, Runfeng Mao, Jinwei Xu, Lanxin Yang, Haifeng Shen, He Zhang 0001
J. Syst. Softw.7
2025 Fine-Tuning Large Language Models to Improve Accuracy and Comprehensibility of Automated Code Review
abstract
As code review is a tedious and costly software quality practice, researchers have proposed several machine learning-based methods to automate the process. The primary focus has been on accuracy, that is, how accurately the algorithms are able to detect issues in the code under review. However, human intervention still remains inevitable since results produced by automated code review are not 100% correct. To assist human reviewers in making their final decisions on automatically generated review comments, the comprehensibility of the comments underpinned by accurate localization and relevant explanations for the detected issues with repair suggestions is paramount. However, this has largely been neglected in the existing research. Large language models (LLMs) have the potential to generate code review comments that are more readable and comprehensible by humans, thanks to their remarkable processing and reasoning capabilities. However, even mainstream LLMs perform poorly in detecting the presence of code issues because they have not been specifically trained for this binary classification task required in code review. In this article, we contribute Comprehensibility of Automated Code Review using Large Language Models ( Carllm ), a novel fine-tuned LLM that has the ability to improve not only the accuracy but, more importantly, the comprehensibility of automated code review, as compared to state-of-the-art pre-trained models and general LLMs.
Yongda Yu, Guoping Rong, Haifeng Shen, He Zhang 0001, Dong Shao, Zhao Wei, Juhong Wang
ACM Trans. Softw. Eng. Methodol.3
2024 KRIOTA: A framework for Knowledge-management of dynamic Reference Information and Optimal Task Assignment in hybrid edge-cloud environments to support situation-aware robot-assisted operations
Muhammad Aufeef Chauhan, Muhammad Ali Babar 0001, Haifeng Shen
Future Gener. Comput. Syst.3
2024 Self-Supervised Monocular Depth Estimation via Binocular Geometric Correlation Learning
abstract
Monocular depth estimation aims to infer a depth map from a single image. Although supervised learning-based methods have achieved remarkable performance, they generally rely on a large amount of labor-intensively annotated data. Self-supervised methods, on the other hand, do not require any annotation of ground-truth depth and have recently attracted increasing attention. In this work, we propose a self-supervised monocular depth estimation network via binocular geometric correlation learning. Specifically, considering the inter-view geometric correlation, a binocular cue prediction module is presented to generate the auxiliary vision cue for the self-supervised learning of monocular depth estimation. Then, to deal with the occlusion in depth estimation, an occlusion interference attenuated constraint is developed to guide the supervision of the network by inferring the occlusion region and producing paired occlusion masks. Experimental results on two popular benchmark datasets have demonstrated that the proposed network obtains competitive results compared to state-of-the-art self-supervised methods and achieves comparable results to some popular supervised methods.
Bo Peng 0007, Jianjun Lei 0001, Bingzheng Liu, Haifeng Shen, Wanqing Li 0001, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Distilling Quality Enhancing Comments From Code Reviews to Underpin Reviewer Recommendation
abstract
Code review is an important practice in software development. One of its main objectives is for the assurance of code quality. For this purpose, the efficacy of code review is subject to the credibility of reviewers, i.e., reviewers who have demonstrated strong evidence of previously making quality-enhancing comments are more credible than those who have not. Code reviewer recommendation (CRR) is designed to assist in recommending suitable reviewers for a specific objective and, in this context, assurance of code quality. Its performance is susceptible to the relevance of its training dataset to this objective, composed of all reviewers’ historical review comments, which, however, often contains a plethora of comments that are irrelevant to the enhancement of code quality. Furthermore, recommendation accuracy has been adopted as the sole metric to evaluate a recommender's performance, which is inadequate as it does not take reviewers’ relevant credibility into consideration. These two issues form the ground truth problem in CRR as they both originate from the relevance of dataset used to train and evaluate CRR algorithms. To tackle this problem, we first propose the concept of Quality-Enhancing Review Comments (QERC), which includes three types of comments - change-triggering inline comments, informative general comments, and approve-to-merge comments. We then devise a set of algorithms and procedures to obtain a distilled dataset by applyingQERCto the original dataset. We finally introduce a new metric – reviewer's credibility for quality enhancement (RCQE) – as a complementary metric to recommendation accuracy for evaluating the performance of recommenders. To validate the proposed QERC-based approach to CRR, we conduct empirical studies using real data from seven projects containing over 82K pull requests and 346K review comments. Results show that: (a)QERCcan effectively address the ground truth problem by distilling quality-enhancing comments from the dataset containing original code reviews, (b)QERCcan assist recommenders in finding highly credible reviewers at a slight cost of recommendation accuracy, and (c) even “wrong” recommendations using the distilled dataset are likely to be more credible than those using the original dataset.
Guoping Rong, Yongda Yu, He Zhang 0001, Haifeng Shen, Dong Shao, Hongyu Kuang, Zhao Wei, Juhong Wang
IEEE Trans. Software Eng.5
2023 Fed-SC: One-Shot Federated Subspace Clustering over High-Dimensional Data
abstract
Recent work has explored federated clustering and developed an efficient k-means based method. However, it is well known that k-means clustering underperforms in high-dimensional space due to the so-called "curse of dimensionality". In addition, high-dimensional data (e.g., generated from healthcare, medical, and biological sectors) are pervasive in the big data era, which poses critical challenges to federated clustering in terms of, but not limited to, clustering effectiveness and communication efficiency. To fill this significant gap in federated clustering, we propose a one-shot federated subspace clustering scheme Fed-SC that can achieve remarkable clustering effectiveness on high-dimensional data while keeping communication cost low using only one round of communication for each local device. We further establish theoretical guarantees on the clustering effectiveness of one-shot Fed-SC and exploit the benefits of statistical heterogeneity across distributed data. Extensive experiments on synthetic and real-world datasets demonstrate significant effectiveness gains of Fed-SC compared with both subspace clustering and one-shot federated clustering methods.
Songjie Xie, Youlong Wu, Kewen Liao, Lu Chen 0008, Chengfei Liu, Haifeng Shen, MingJian Tang 0001, Lu Sun 0001
ICDE6
2023 How Do Developers' Profiles and Experiences Influence their Logging Practices? An Empirical Study of Industrial Practitioners
abstract
Logs record the behavioral data of running programs and are typically generated by executing log statements. Software developers generally carry out logging practices with clear intentions and associated concerns (I&Cs). However, I&Cs may not be properly fulfilled in source code as log placement - specifically determination of a log statement's context and content - is often susceptible to an individual's profile and experience. Some industrial studies have been conducted to discern developers' main logging I&Cs and the way I&Cs are fulfilled. However, the findings are only based on the developers from a single company in each individual study and hence have limited generalizability. More importantly, there lacks a comprehensive and deep understanding of the relationships between developers' profiles and experiences and their logging practices from a wider perspective. To fill this significant gap, we conducted an empirical study using mixed methods comprising questionnaire surveys, semi-structured interviews, and code analyses with practitioners from a wide range of companies across a variety of industrial domains. Results reveal that while developers share common logging I&Cs and conduct logging practices mainly in the coding stage, their profiles and experiences profoundly influence their logging I&Cs and the way the I&Cs are fulfilled. These findings pave the way to facilitate the acceptance of important logging I&Cs and the adoption of good logging practices by developers
Guoping Rong, Shenghui Gu, Haifeng Shen, He Zhang 0001, Hongyu Kuang
ICSE3
2023 DeepHeteroIoT: Deep Local and Global Learning over Heterogeneous IoT Sensor Data
Muhammad Sakib Khan Inan, Kewen Liao, Haifeng Shen, Prem Prakash Jayaraman, Dimitrios Georgakopoulos 0001, Ming Jian Tang
MobiQuitous (1)3
2023 Revisit security in the era of DevOps: An evidence-based inquiry into DevSecOps industry
abstract
Abstract By adopting agile and lean practices, DevOps aims to achieve rapid value delivery by speeding up development and deployment cycles, which however lead to more security concerns that cannot be fully addressed by an isolated security role only in the final stage of development. DevSecOps promotes security as a shared responsibility integrated into the DevOps process that seamlessly intertwines development, operations, and security from the start throughout to the end of cycles. While some companies have already begun to embrace this new strategy, both industry and academia are still seeking a common understanding of the DevSecOps movement. The goal of this study is to report the state‐of‐the‐practice of DevSecOps, including the impact of DevOps on security, practitioners' understanding of DevSecOps, and the practices associated with DevSecOps as well as the challenges of implementing DevSecOps. The authors used a mixed‐methods approach for this research. The authors carried out a grey literature review on DevSecOps, and surveyed the practitioners of DevSecOps in industry of China. The status quo of DevSecOps in industry is summarized. Three major software security risks are identified with DevOps, where the establishment of DevOps pipeline provides opportunities for security‐related activities. The authors classify the interpretations of DevSecOps into three core aspects of DevSecOps capabilities, cultural enablers, and technological enablers. To materialise the interpretations into daily software production activities, the recommended DevSecOps practices from three perspectives—people, process, and technology. Although a preliminary consensus is that DevSecOps is regarded as an extension of DevOps, there is a debate on whether DevSecOps is a superfluous term. While DevSecOps is attracting an increasing attention by industry, it is still in its infancy and more effort needs to be invested to promote it in both research and industry communities.
Xin Zhou 0016, Runfeng Mao, He Zhang 0001, Qiming Dai, Haifeng Shen, Jingyue Li, Guoping Rong
IET Softw.6
2023 Evaluating the efficacy of using a novel gaze-based attentive user interface to extend ADHD children's attention span
Haifeng Shen, Othman Asiry, Muhammad Ali Babar 0001, Tomasz Bednarz
Int. J. Hum. Comput. Stud.1
2023 RGB-D Human Matting: A Real-World Benchmark Dataset and a Baseline Method
abstract
The last decade has witnessed an increasing exploration and development of human matting. However, existing matting works primarily focus on predicting better alpha mattes from RGB images. So far few efforts have been devoted to tackling human matting in real-world activity scenarios with RGB-D information. To this end, this paper concentrates on the RGB-D human matting task, and provides the first public RGB-D human matting benchmark dataset as well as a baseline method for deep learning-based RGB-D human matting. To support the research on RGB-D human matting, a new RGB-D human-matting dataset (HDM-2K) is collected and released, which contains 2,270 high-resolution human images in various real-world scenarios and the corresponding depth maps. Additionally, a baseline method for RGB-D human matting is further proposed, which automatically generates the alpha matte by jointly exploiting the spatial structure information in the depth map and detailed texture information in the RGB image. Finally, extensive experiments conducted on the HDM-2K dataset demonstrate that the depth maps are effective for the matting task and the proposed baseline method achieves promising performance on human matting.
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2023 Recurrent Interaction Network for Stereoscopic Image Super-Resolution
abstract
Recently, deep learning-based stereoscopic image super-resolution has attracted extensive attention and made great progress. However, existing methods have not adequately explored the inter-view dependency among two-view multi-level features. In this paper, a recurrent interaction network for stereoscopic image super-resolution (RISSRnet) is proposed to learn the inter-view dependency. To efficiently utilize the relationship between the two views, a recurrent interaction module is designed to achieve recurrent interaction among two-view multi-level features from the regrouped sequences, which are generated by a coupled queue-regroup mechanism. In addition, to recursively enhance features in the recurrent interaction module, an iterative propagation strategy is developed for sufficient interaction. Extensive experimental results demonstrate the effectiveness and superiority of the proposed RISSRnet.
Zhe Zhang 0041, Bo Peng 0007, Jianjun Lei 0001, Haifeng Shen, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Reducing Background Induced Domain Shift for Adaptive Person Re-Identification
abstract
Cross-domain person re-identification (Re-ID) is a challenging and important task in monitoring safety and procedure compliance of industrial work places. In this article, a novel method is proposed to reduce background induced domain shift for adaptive person Re-ID. Specifically, a foreground-background joint clustering module is proposed to extract discriminative foreground and background features and an attention-based feature disentanglement module is designed to reduce the interference of background with the extraction of discriminative foreground features. Experimental results on three widely used person Re-ID benchmarking datasets (Market-1501, DukeMTMC-reID, and MSMT17) have demonstrated that the proposed method achieves promising performance compared with the state-of-the-art methods.
Jianjun Lei 0001, Tianyi Qin, Bo Peng 0007, Wanqing Li 0001, Zhaoqing Pan, Haifeng Shen, Sam Kwong
IEEE Trans. Ind. Informatics6
2023 ZS-SBPRnet: A Zero-Shot Sketch-Based Point Cloud Retrieval Network Based on Feature Projection and Cross-Reconstruction
abstract
With the widespread deployment of 3D sensors, point cloud analysis has become an important topic in the field of industrial information. This article proposes a novel zero-shot sketch-based point cloud retrieval network based on feature projection and cross reconstruction, termed as ZS-SBPRnet. As far as we know, the proposed ZS-SBPRnet is the first attempt at retrieving point clouds based on sketches under the zero-shot scenario. To tackle the problem of the cross-modal differences, a structure-preserving learnable feature projection module is designed to obtain view feature representations from point cloud features containing spatial structure information through feature projection. Besides, to achieve efficient cross-modal feature alignment under the zero-shot scenario, a sketch-point cloud cross-reconstruction mechanism is presented to promote cross-modal feature alignment between sketches and point clouds in visual space. Experimental results on the benchmark datasets validate the superiority of the proposed ZS-SBPRnet.
Bo Peng 0007, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
IEEE Trans. Ind. Informatics4
2023 Modeling Long-range Dependencies and Epipolar Geometry for Multi-view Stereo
abstract
This article proposes a network, referred to as Multi-View Stereo TRansformer (MVSTR) for depth estimation from multi-view images. By modeling long-range dependencies and epipolar geometry, the proposed MVSTR is capable of extracting dense features with global context and 3D consistency, which are crucial for reliable matching in multi-view stereo (MVS). Specifically, to tackle the problem of the limited receptive field of existing CNN-based MVS methods, a global-context Transformer module is designed to establish intra-view long-range dependencies so that global contextual features of each view are obtained. In addition, to further enable features of each view to be 3D consistent, a 3D-consistency Transformer module with an epipolar feature sampler is built, where epipolar geometry is modeled to effectively facilitate cross-view interaction. Experimental results show that the proposed MVSTR achieves the best overall performance on the DTU dataset and demonstrates strong generalization on the Tanks & Temples benchmark dataset.
Bo Peng 0007, Wanqing Li 0001, Haifeng Shen, Qingming Huang, Jianjun Lei 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2023 TrinityRCL: Multi-Granular and Code-Level Root Cause Localization Using Multiple Types of Telemetry Data in Microservice Systems
abstract
The microservice architecture has been commonly adopted by large scale software systems exemplified by a wide range of online services. Service monitoring through anomaly detection and root cause analysis (RCA) is crucial for these microservice systems to provide stable and continued services. However, compared with monolithic systems, software systems based on the layered microservice architecture are inherently complex and commonly involve entities at different levels of granularity. Therefore, for effective service monitoring, these systems have a special requirement of multi-granular RCA. Furthermore, as a large proportion of anomalies in microservice systems pertain to problematic code, to timely troubleshoot these anomalies, these systems have another special requirement of RCA at the finest code-level. Microservice systems rely on telemetry data to perform service monitoring and RCA of service anomalies. The majority of existing RCA approaches are only based on a single type of telemetry data and as a result can only support uni-granular RCA at either application-level or service-level. Although there are attempts to combine metric and tracing data in RCA, their objective is to improve RCA's efficiency or accuracy rather than to support multi-granular RCA. In this article, we propose a new RCA solutionTrinityRCLthat is able to localize the root causes of anomalies at multiple levels of granularity including application-level, service-level, host-level, and metric-level, with the unique capability of code-level localization by harnessing all three types of telemetry data to construct a causal graph representing the intricate, dynamic, and nondeterministic relationships among the various entities related to the anomalies. By implementing and deployingTrinityRCLin a real production environment, we evaluateTrinityRCLagainst two baseline methods and the results show thatTrinityRCLhas a significant performance advantage in terms of accuracy at the same level of granularity with comparable efficiency and is particularly effective to support large-scale systems with massive telemetry data.
Shenghui Gu, Guoping Rong, Tian Ren, He Zhang 0001, Haifeng Shen, Yongda Yu, Jian Ouyang, Chunan Chen
IEEE Trans. Software Eng.5
2023 Logging Practices in Software Engineering: A Systematic Mapping Study
abstract
Background:Logging practices provide the ability to record valuable runtime information of software systems to support operations tasks such as service monitoring and troubleshooting. However, current logging practices face common challenges. On the one hand, although the importance of logging practices has been broadly recognized, most of them are still conducted in an arbitrary or ad-hoc manner, ending up with questionable or inadequate support to perform these tasks. On the other hand, considerable research effort has been carried out on logging practices, however, few of the proposed techniques or methods have been widely adopted in industry.Objective:This study aims to establish a comprehensive understanding of the research state of logging practices, with a focus on unveiling possible problems and gaps which further shed light on the potential future research directions.Method:We carried out a systematic mapping study on logging practices with 56 primary studies.Results:This study provides a holistic report of the existing research on logging practices by systematically synthesizing and analyzing the focus and inter-relationship of the existing research in terms of issues, research topics and solution approaches. Using3W1H—Why to log,Where to log,What to logandHow well is the logging—as the categorization standard, we find that: (1) the best known issues in logging practices have been repeatedly investigated; (2) the issues are often studied separately without considering their intricate relationships; (3) theWhere and Whatquestions have attracted the majority of research attention while little research effort has been made on theWhyandHow wellquestions; and (4) the relationships between issues, research topics, and approaches regarding logging practices appear many-to-many, which indicates a lack of profound understanding of the issues in practice and how they should be appropriately tackled.Conclusions:This study indicates a need to advance the state of research on logging practices. For example, more research effort should be invested onwhy to logto set the anchor of logging practices as well as onhow well is the loggingto close the loop. In addition, a holistic process perspective should be taken into account in both the research and the adoption related to logging practices.
Shenghui Gu, Guoping Rong, He Zhang 0001, Haifeng Shen
IEEE Trans. Software Eng.4
2023 The Why, When, What, and How About Predictive Continuous Integration: A Simulation-Based Investigation
abstract
Continuous Integration (CI) enables developers to detect defects early and thus reduce lead time. However, the high frequency and long duration of executing CI have a detrimental effect on this practice. Existing studies have focused on using CI outcome predictors to reduce frequency. Since there is no reported project using predictive CI, it is difficult to evaluate its economic impact. This research aims to investigate predictive CI from a process perspective, including why and when to adopt predictors, what predictors to be used, and how to practice predictive CI in real projects. We innovatively employ Software Process Simulation to simulate a predictive CI process with a Discrete-Event Simulation (DES) model and conduct simulation-based experiments. We develop the Rollback-based Identification of Defective Commits (RIDEC) method to account for the negative effects of false predictions in simulations. Experimental results show that: 1) using predictive CI generally improves the effectiveness of CI, reducing time costs by up to 36.8% and the average waiting time before executing CI by 90.5%; 2) the time-saving varies across projects, with higher commit frequency projects benefiting more; and 3) predictor performance does not strongly correlate with time savings, but the precision of both failed and passed predictions should be paid more attention. Simulation-based evaluation helps identify overlooked aspects in existing research. Predictive CI saves time and resources, but improved prediction performance has limited cost-saving benefits. The primary value of predictive CI lies in providing accurate and quick feedback to developers, aligning with the goal of CI.
Bohan Liu 0003, He Zhang 0001, Weigang Ma, Gongyuan Li, Shanshan Li 0002, Haifeng Shen
IEEE Trans. Software Eng.6
2022 Challenges and solutions when adopting DevSecOps: A systematic review
Roshan Namal Rajapakse, Mansooreh Zahedi, Muhammad Ali Babar 0001, Haifeng Shen
Inf. Softw. Technol.4
2022 SIEV-Net: A Structure-Information Enhanced Voxel Network for 3D Object Detection From LiDAR Point Clouds
abstract
As one of the fundamental tasks in scene understanding, 3D object detection from LiDAR point clouds has drawn extensive attention in the past few years. Although the existing voxel-based methods have achieved remarkable performance, how to effectively exploit geometric structure information of the point clouds to boost the detection performance remains to be explored. In this paper, we propose a novel structure-information enhanced voxel network (SIEV-Net) for 3D object detection from LiDAR point clouds. The proposed SIEV-Net learns feature representations of 3D objects by jointly considering uneven spatial distribution and height information of the point clouds. Specifically, considering the uneven spatial distribution characteristics of point clouds, a hierarchical-voxel feature encoding module is proposed to effectively extract features of voxels in both sparse and dense regions. Besides, by utilizing the Bird’s Eye View (BEV) map of point clouds, a height information complement module is designed to minimize the height information lost in the process of point feature aggregation in a voxel network. Experimental results on the widely used KITTI benchmark dataset have demonstrated the efficacy of the proposed SIEV-Net.
Chuanbo Yu, Jianjun Lei 0001, Bo Peng 0007, Haifeng Shen, Qingming Huang
IEEE Trans. Geosci. Remote. Sens.4
2022 CNN Attention Guidance for Improved Orthopedics Radiographic Fracture Classification
abstract
Convolutional neural networks (CNNs) have gained significant popularity in orthopedic imaging in recent years due to their ability to solve fracture classification problems. A common criticism of CNNs is their opaque learning and reasoning process, making it difficult to trust machine diagnosis and the subsequent adoption of such algorithms in clinical setting. This is especially true when the CNN is trained with limited amount of medical data, which is a common issue as curating sufficiently large amount of annotated medical imaging data is a long and costly process. While interest has been devoted to explaining CNN learnt knowledge by visualizing network attention, the utilization of the visualized attention to improve network learning has been rarely investigated. This paper explores the effectiveness of regularizing CNN network with human-provided attention guidance on where in the image the network should look for answering clues. On two orthopedics radiographic fracture classification datasets, through extensive experiments we demonstrate that explicit human-guided attention indeed can direct correct network attention and consequently significantly improve classification performance. The development code for the proposed attention guidance is publicly available on https://github.com/zhibinliao89/fracture_attention_guidance.
Zhibin Liao, Kewen Liao, Haifeng Shen, Marouska F. van Boxel, Jasper Prijs, Ruurd L. Jaarsma, Job N. Doornberg, Anton van den Hengel, Johan Verjans
IEEE J. Biomed. Health Informatics3
2021 Understanding the effects of real-time sentiment analysis and morale visualisation in backchannel systems: A case study
Theodor Wyeld, Peerumporn Jiranantanagorn, Haifeng Shen, Kewen Liao, Tomasz Bednarz
Int. J. Hum. Comput. Stud.3
2021 Quality Assessment in Systematic Literature Reviews: A Software Engineering Perspective
Lanxin Yang, He Zhang 0001, Haifeng Shen, Xin Huang 0019, Xin Zhou 0016, Guoping Rong, Dong Shao
Inf. Softw. Technol.3
2021 Processes, challenges and recommendations of Gray Literature Review: An experience report
He Zhang 0001, Runfeng Mao, Qiming Dai, Xin Zhou 0016, Haifeng Shen, Guoping Rong
Inf. Softw. Technol.6
2021 SRGAT: Single Image Super-Resolution With Graph Attention Network
abstract
Deep neural networks have demonstrated remarkable reconstruction for single-image super-resolution (SISR). However, most existing CNN-based SISR methods directly learn the relation between low-resolution (LR) and high-resolution (HR) images, neglecting to explore the recurrence of internal patches, hence hindering the representational power of CNNs. In this paper, we propose a novel single image Super-Resolution network based on Graph ATtention network (SRGAT) to make full use of the internal patch-recurrence in a natural image. The proposed model employs a feature mapping block with a recurrent structure to refine low-level representations with high-level information. Especifically, the feature mapping block contains a parallel graph similarity branch and a content branch, where the graph similarity branch aims at exploiting the similarity and symmetry across different image patches in low-resolution feature space and provides additional priors for the content branch to enhance texture details. Specifically, we consider the internal patch-recurrence of an image by constructing a graph network on image feature patches. In this way, the information from neighboring patches can be interacted using graph attention network (GAT) to help it recover additional textures, which complements the textures learned from the content branch. Extensive quantitative and qualitative evaluations on five benchmark datasets demonstrate that the proposed algorithm performs favorably against the state-of-the-art super-resolution methods.
Yanyang Yan, Wenqi Ren, Xiaobin Hu, Kun Li 0029, Haifeng Shen, Xiaochun Cao
IEEE Trans. Image Process.5
2020 CFM: A Consistency Filtering Mechanism for Road Damage Detection
abstract
This article presents the solution that we use in the Global Road Damage Detection Challenge 2020, which is designed to recognize the road damages present in an image captured from three countries: India, Japan, and Czech. In this challenge, Cascade R-CNN is selected as a baseline model to detect objects in images. It is commonly known that making a precise annotation in a large dataset is crucial to the performance of object detection and placing bounding boxes for every object in each image is time-consuming and costs a lot. To make full use of available unlabeled data, the consistency filtering mechanism (CFM) with self-supervised methods is proposed to utilize high-confident samples with pseudo-labels for training. And we also apply a series of data augmentation techniques (road segmentation, flip, mixup, CLAHE) to labeled data in training phase. Moreover, we ensemble models with different tricks by weighted boxes fusion to produce the final prediction. Finally, our proposed method can achieve a great mean f1-score of 0.6290 on the test1 dataset and 0.6219 on the test2 dataset respectively, which wins the Bronze Prize (ranks 3rd place). Code and trained models are available at the following link: https://pan.baidu.com/s/1VjLuNBVJGS34mMMpDkDRGQ, password: xzc6.
Zixiang Pei, Rongheng Lin, Xiubao Zhang, Haifeng Shen, Jian Tang 0008
IEEE BigData4
2020 PropagationNet: Propagate Points to Curve to Learn Structure Information
abstract
Deep learning technique has dramatically boosted the performance of face alignment algorithms. However, due to large variability and lack of samples, the alignment problem in unconstrained situations, e.g. large head poses, exaggerated expression, and uneven illumination, is still largely unsolved. In this paper, we explore the instincts and reasons behind our two proposals, i.e. Propagation Module and Focal Wing Loss, to tackle the problem. Concretely, we present a novel structure-infused face alignment algorithm based on heatmap regression via propagating landmark heatmaps to boundary heatmaps, which provide structure information for further attention map generation. Moreover, we propose a Focal Wing Loss for mining and emphasizing the difficult samples under in-the-wild condition. In addition, we adopt methods like CoordConv and Anti-aliased CNN from other fields that address the shift variance problem of CNN for face alignment. When implementing extensive experiments on different benchmarks, i.e. WFLW, 300W, and COFW, our method outperforms the state-of-the-arts by a significant margin. Our proposed approach achieves 4.05% mean error on WFLW, 2.93% mean error on 300W full-set, and 3.71% mean error on COFW.
Xiehe Huang, Weihong Deng, Haifeng Shen, Xiubao Zhang, Jieping Ye
CVPR3
2020 An Experimental Evaluation of Imbalanced Learning and Time-Series Validation in the Context of CI/CD Prediction
abstract
Background: Machine Learning (ML) has been widely used as a powerful tool to support Software Engineering (SE). The fundamental assumptions of data characteristics required for specific ML methods have to be carefully considered prior to their applications in SE. Within the context of Continuous Integration (CI) and Continuous Deployment (CD) practices, there are two vital characteristics of data prone to be violated in SE research. First, the logs generated during CI/CD for training are imbalanced data, which is contrary to the principles of common balanced classifiers; second, these logs are also time-series data, which violates the assumption of cross-validation. Objective: We aim to systematically study the two data characteristics and further provide a comprehensive evaluation for predictive CI/CD with the data from real projects. Method: We conduct an experimental study that evaluates 67 CI/CD predictive models using both cross-validation and time-series-validation. Results: Our evaluation shows that cross-validation makes the evaluation of the models optimistic in most cases, there are a few counter-examples as well. The performance of the top 10 imbalanced models are better than the balanced models in the predictions of failed builds, even for balanced data. The degree of data imbalance has a negative impact on prediction performance. Conclusion: In research and practice, the assumptions of the various ML methods should be seriously considered for the validity of research. Even if it is used to compare the relative performance of models, cross-validation may not be applicable to the problems with time-series features. The research community need to revisit the evaluation results reported in some existing research.
Bohan Liu 0003, He Zhang 0001, Lanxin Yang, Liming Dong 0001, Haifeng Shen, Kaiwen Song
EASE5
2020 Single Image Super-Resolution via a Holistic Attention Network
Weilei Wen, Wenqi Ren, Xiangde Zhang, Lianping Yang, Shuzhen Wang, Kaihao Zhang, Xiaochun Cao, Haifeng Shen
ECCV (12)9
2020 Adaptive Object Detection with Dual Multi-label Prediction
Yuhong Guo, Haifeng Shen, Jieping Ye
ECCV (28)3
2020 Multi-modal Fusion Using Spatio-temporal and Static Features for Group Emotion Recognition
abstract
This paper presents our approach for Audio-video Group Emotion Recognition sub-challenge in the EmotiW 2020. The task is to classify a video into one of the group emotions such as positive, neutral, and negative. Our approach exploits two different feature levels for this task, spatio-temporal feature and static feature level. In spatio-temporal feature level, we adopt multiple input modalities (RGB, RGB difference, optical flow, warped optical flow) into multiple video classification network to train the spatio-temporal model. In static feature level, we crop all faces and bodies in an image with the state-of the-art human pose estimation method and train kinds of CNNs with the image-level labels of group emotions. Finally, we fuse all 14 models result together, and achieve the third place in this sub-challenge with classification accuracies of 71.93% and 70.77% on the validation set and test set, respectively.
Wei Gou, Haifeng Shen, Jian Tang 0008, Jieping Ye
ICMI5
2020 A Multi-Modal Approach for Driver Gaze Prediction to Remove Identity Bias
abstract
Driver gaze prediction is an important task in Advanced Driver Assistance System (ADAS). Although the Convolutional Neural Network (CNN) can greatly improve the recognition ability, there are still several unsolved problems due to the challenge of illumination, pose and camera placement. To solve these difficulties, we propose an effective multi-model fusion method for driver gaze estimation. Rich appearance representations, i.e. holistic and eyes regions, and geometric representations, i.e. landmarks and Delaunay angles, are separately learned to predict the gaze, followed by a score-level fusion system. Moreover, pseudo-3D appearance supervision and identity-adaptive geometric normalization are proposed to further enhance the prediction accuracy. Finally, the proposed method achieves state-of-the-art accuracy of 82.5288% on the test data, which ranks 1st at the EmotiW2020 driver gaze prediction sub-challenge.
Zehui Yu, Xiehe Huang, Xiubao Zhang, Haifeng Shen, Qun (Tracy) Li, Weihong Deng, Jian Tang 0008, Jieping Ye
ICMI4
2020 $P^{2}$ Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation
abstract
The target of human pose estimation is to determine the body parts and joint locations of persons in the image. Angular changes, motion blur and occlusion in the natural scenes make this task challenging, while some joints are more difficult to be detected than others. In this paper, we propose an augmented Parallel-Pyramid Net ( P2Net) with feature refinement by dilated bottleneck and attention module. During data preprocessing, we proposed a differentiable auto data augmentation ( DA2) method. We formulate the problem of searching data augmentaion policy in a differentiable form, so that the optimal policy setting can be easily updated by back propagation during training. DA2improves the training efficiency. A parallel-pyramid structure is followed to compensate the information loss introduced by the network. We innovate two fusion structures, i.e. Parallel Fusion and Progressive Fusion, to process pyramid features from backbone network. Both fusion structures leverage the advantages of spatial information affluence at high resolution and semantic comprehension at low resolution effectively. We propose a refinement stage for the pyramid features to further boost the accuracy of our network. By introducing dilated bottleneck and attention module, we increase the receptive field for the features with limited complexity and tune the importance to different feature channels. To further refine the feature maps after completion of feature extraction stage, an Attention Module ( AM) is defined to extract weighted features from different scale feature maps generated by the parallel-pyramid structure. Compared with the traditional up-sampling refining, AM can better capture the relationship between channels. Experiments corroborate the effectiveness of our proposed method. Notably, our method achieves the best performance on the challenging MSCOCO and MPII datasets.
Luanxuan Hou, Jie Cao 0002, Haifeng Shen, Jian Tang 0008, Ran He 0001
ICPR4
2020 Preliminary Findings about DevSecOps from Grey Literature
abstract
Context: Emerging from the agile culture, DevOps particularly emphasizes development and deployment speed to achieve rapid value delivery, which however brings some security risks to the software development process. DevSecOps is an extension of DevOps, which is considered as a means to intertwine development, operation and security. Some companies with security concerns begin to take DevSecOps into consideration when it comes to the application of DevOps. Objective: The goal of this study is to report the state-of-the-practice of DevSecOps as well as calling for academia to pay more attention to DevSecOps. Method: Using Google search engine to collect articles on DevSecOps, we conducted a Grey Literature Review (GLR) on the selected articles. Results: Whilst there exists three major software security risks in DevOps, the establishment of DevOps pipeline provides opportunities for software security activities. Based on the preliminary consensus that DevSecOps is an extension of DevOps, it is observed that the interpretations of DevSecOps can be classified into three core aspects, which are: DevSecOps capabilities, cultural enablers, and technological enablers. Furthermore, to materialize the interpretations into daily software production activities, the recommended DevSecOps practices we obtain from Grey Literature (GL) can be categorized in terms of process, infrastructure and collaboration. Conclusion: Although DevSecOps is getting increasing attention by industry, it is still in its infancy and needs to be promoted by both academia and industry.
Runfeng Mao, He Zhang 0001, Qiming Dai, Guoping Rong, Haifeng Shen, Lianping Chen, Kaixiang Lu
QRS6
2019 An adaptive differential evolution algorithm to optimal multi-level thresholding for MRI brain image segmentation
Omid Tarkhaneh, Haifeng Shen
Expert Syst. Appl.2
2018 Cross-Generating GAN for Facial Identity Preserving
abstract
The large variations of pose and illumination have been the great challenges to face recognition for many years. Because of these variations, many classical recognition methods fail to work. The key to solve this problem is to extract identity feature from face images. In recent years, people have been concentrating on synthesizing rotated faces, however, neglected the form of facial identity representation. In this paper, we propose Cross-generating Generative Adversarial Network (CG-GAN) to generate rotated faces while extracting discriminative identity. CG-GAN is allowed to learn a network to exchange poses and illuminations of two different subjects' picture. Within the network, each input image is resolved into a variation code and a identity code at the representation layer; then these codes are randomly combined for generating corresponding pictures. Not only does CG-GAN synthesis vivid face under desired pose from one picture, but also the represention layer is very suitable for face recognition task. We train and test CG-GAN on the Multi-PIE dataset and achieve state-of-the-art results.
Weilong Chai, Weihong Deng, Haifeng Shen
FG3
2018 Deep Unsupervised Domain Adaptation for Face Recognition
abstract
Face recognition is challenge task which involves determining the identity of facial images. With availability of a massive amount of labeled facial images gathered from Internet, deep convolution neural networks(DCNNs) have achieved great success in face recognition tasks. Those images are gathered from unconstrain environment, which contain people with different ethnicity, age, gender and so on. However, in the actual application scenario, the target face database may be gathered under different conditions compared with source training dataset, e.g. different ethnicity, different age distribution, disparate shooting environment. These factors increase domain discrepancy between source training database and target application database which makes the learnt model degenerate in target database. Meanwhile, for the target database where labeled data are lacking or unavailable, directly using target data to fine-tune pre-learnt model becomes intractable and impractical. In this paper, we adopt unsupervised transfer learning methods to address this issue. To alleviate the discrepancy between source and target face database and ensure the generalization ability of the model, we constrain the maximum mean discrepancy (MMD) between source database and target database and utilize the massive amount of labeled facial images of source database to training the deep neural network at the same time. We evaluate our method on two face recognition benchmarks and significantly enhance the performance without utilizing the target label.
Zimeng Luo, Jiani Hu, Weihong Deng, Haifeng Shen
FG4
2018 Extending Attention Span for Children ADHD Using an Attentive Visual Interface
abstract
Attention Deficit Hyperactivity Disorder (ADHD) is a common developmental disorder usually accompanying other developmental disorders including speech, language and reading. Children with ADHD tend to lose attention after a short period of time. Extending the attention span for those children could help them do better in school and in life. The aim of this work is to assess the role of text color (highlighting, contrast, sharpening) on the attention of children with ADHD while they are reading. Attention is tracked via the two modalities of webcam and mouse as some ADHD children have difficulties in maintaining the calibration of webcam. Visual color schemes are modified to evaluate different ways of maintaining attention. The results reveal that: (a) all color schemes have a significant effect on attention span, (b) highlighting has the greatest effect regardless of tracking modality, and (c) the degree of effect is subject to the tracking modality.
Othman Asiry, Haifeng Shen, Theodor Wyeld, Soher Balkhy
IV2
2018 Virtual Class Enhanced Discriminative Embedding Learning
abstract
Recently, learning discriminative features to improve the recognition performances gradually becomes the primary goal of deep learning, and numerous remarkable works have emerged. In this paper, we propose a novel yet extremely simple method Virtual Softmax to enhance the discriminative property of learned features by injecting a dynamic virtual negative class into the original softmax. Injecting virtual class aims to enlarge inter-class margin and compress intra-class distribution by strengthening the decision boundary constraint. Although it seems weird to optimize with this additional virtual class, we show that our method derives from an intuitive and clear motivation, and it indeed encourages the features to be more compact and separable. This paper empirically and experimentally demonstrates the superiority of Virtual Softmax, improving the performances on a variety of object classification and face verification tasks.
Binghui Chen, Weihong Deng, Haifeng Shen
NeurIPS3
2018 A smartphone-based point-of-care quantitative urinalysis device for chronic kidney disease patients
Shaymaa Akraa, Anh Pham Tran Tam, Haifeng Shen, Youhong Tang, Ben Zhong Tang, Sandy Walker
J. Netw. Comput. Appl.3
2018 Integrating Localization and Energy-Awareness: A Novel Geographic Routing Protocol for Underwater Wireless Sensor Networks
Kun Hao, Haifeng Shen, Yonglei Liu, Xiujuan Du
Mob. Networks Appl.2
2017 iLSE: An Intelligent Web-Based System for Log Structuring and Extraction
abstract
Analysing software log files has become a challenging task due to the diversity in file structure and the nonstandardisation of log syntax. During the process of extracting log data, it is required to manually decode the log syntax and interpret data semantics, which can become tedious and is often error-prone if not performed carefully. Contemporary log analysis software tools do exist in the market and most of them offer numerous options to analyse log files, however, their sheer focus is on providing log management solutions instead of log analysis capabilities. In particular, none of them offers a generic parsing and extracting solution that can discover hidden data structures, a critical and effortful task in log analysis. We thereby devise such a solution that is able to automatically identify hidden patterns in a given log file and extract useful information by generalising the patterns. The solution is implemented as an intelligent Web-based system known as iLSE (intelligent Log Structuring and Extraction) whose users are not required to possess fluent programming skills. This paper presents a reference architecture for the system as well as a comparison study on how the system performs against contemporary log analysis systems.
Sahan Serasinghe, Haifeng Shen, David Chen 0002
APSEC2
2017 On the Feasibility of a Smartphone-based Solution to Rapid Qantitative Urinalysis using Nanomaterial Bioprobes
abstract
The main objective of this research is to design and develop a smartphone-based urinalysis device known as uTest that has the ability for patients themselves to conduct rapid quantitative diagnosis of human serum albumin (HSA) in urine using the aggregation-induced emission (AIE) nanomaterial bioprobes anytime, anywhere and with any mobile device. Work reported in this study has confirmed the feasibility of such a solution, which can achieve accurate urinalysis for HSA concentrations in the range of 0-100 mg/dL.
Shaymaa Akraa, Haifeng Shen, Youhong Tang, Gobert N. Lee, Ben Zhong Tang
MobiQuitous3
2017 Are you a human or a humanoid: Predictive user modelling through behavioural analysis of online gameplay data
Kaiqi Jin, Haifeng Shen, Muhammad Ali Babar 0001
Adv. Eng. Informatics3
2016 Concealing jitter in Multi-Player Online Games through predictive behaviour modeling
abstract
Network latency imposes a major hinderance on the responsiveness and consistency of a Multi-Player Online Game (MPOG). In past decades, several network topologies and latency handling solutions have been proposed and adopted in a variety of MPOGs. Another issue closely related to latency is jitter, which is caused by the variation of latency. Most of the existing MPOGs adopt a simplistic approach to tackling jitter: when one's varying latency exceeds a threshold, one will be forced to leave the game; otherwise latency is treated constant or estimated from historical data. However, forcing a player to quit in the middle of a game simply because of a spike of unusual lengthy latency has a significant negative impact on both the fairness of the game and the player's gameplay quality of experience (QoE). In this paper, we propose an alternative approach that instead conceals jitter by seamlessly and transparently switching between a remote human player and their intelligent agent that resembles them. To model such an intelligent agent, we further contribute a novel technique referred to as predictive modeling of user behaviour (PREMUB), which predicts an object's future state based on how the remote player interacts with the object in the past. We have also developed an online table tennis game to demonstrate this idea and compare the prediction accuracy between using PREMUB and using the existing technique of dead reckoning that does not consider a user's playing pattern.
Haifeng Shen, Muhammad Ali Babar 0001
CSCWD2
2016 NSSSD: A new semantic hierarchical storage for sensor data
abstract
Sensor networks usually generate mass of data, which if not structured for future applications, will require much effort on analytical processing and interpretations. Thus, storing sensor data in an effective and structured format is a key issue in the area of sensor networks. In the meantime, even a little improvement on data storing structure may lead to a significant effect on the lifetime and performance of the sensor network. This paper describes a new method for sensor storage that combines semantic web concepts, a data aggregation method along with aligning sensors in hierarchical form. This solution is able to reduce the amount of data stored at the sink nodes significantly. At the same time, the method structures sensed data in a way that we can respond to semantic web-based queries with less consumption of energy compared to previous conventional methods. Results show that, in some situations especially when the diversity of query responses and life of network are vital, the efficiency of our new solution is much better.
Mehdi Gheisari, Ali Akbar Movassagh, Yongrui Qin, Jianming Yong, Xiaohui Tao 0001, Ji Zhang 0001, Haifeng Shen
CSCWD7
2016 Web of Credit: Adaptive Personalized Trust Network Inference From Online Rating Data
abstract
Trust is a pivotal element of any information system that allows users to share, communicate, interact, or collaborate with one another. Trust inference is particularly crucial for online social networks where interaction with acquaintances or even anonymous strangers is widely a norm. In the past decade, a number of trust inference algorithms have been proposed to address this issue, which are primarily based either on the “reputation” or the “Web of trust (WoT)” model. The reputation-based model supports objective inference of a universal reputation for each user by analyzing the interaction histories among the users; however, it does not allow individual users to specify personalized trust measures for the same other users. In contrast, the WoT-based model allows each individual user to specify a trust value for their direct neighbors within a trust network. However, the accuracy of such a subjective trust value is questionable and further subject to loss in the course of propagating trust measures to nonneighboring users in the network. In this paper, we propose a new trust model referred to as “Web of credit (WoC),” where one gives credit to those others one has interacted with based on the quality of the information one's peers have provided. Credit flows from one user to another within a trust network, forming trust relationships. This new model combines the objectivism from the reputation-based model for credit assignment by exploiting the actual interaction histories among users in the form of online rating data and the individualism from the WoT-based model for personalized trust measures. We further contribute a WoC-based trust inference algorithm that is adaptive to the change of user profiles by automatically redistributing credit and reinferring trust measures within the network. Experiments with two real-world data sets have shown that the WoC-based trust inference algorithm is not only able to infer more accurate trust measures than both reputation-based and WoT-based algorithms do but also fast enough to be a viable solution for real-time trust inference in large-scale trust networks.
Haifeng Shen
IEEE Trans. Comput. Soc. Syst.2
2014 SORC: Service-Oriented Distributed Revision Control for collaborative web programming
abstract
Web applications become ubiquitous and more complex, often requiring team developmental work. Software configuration management (SCM) systems have long been used for managing collaborative development. However, the centralised architecture - the most common architecture of today's SCM systems - which requires developers to replicate all project source files, is not suitable for Web applications, especially those that are service-oriented and distributed by nature. In this paper, we present Service-Oriented Revision Control (SORC), a distributed SCM model specifically for effectively supporting collaborative programming of Web applications, which does not rely on the centralised architecture or replicate project source files across developers. SORC allows a developer's code to be exposed to their peers as Web services, while revision control of the project is at the service rather than the file level. We have further developed a prototype SCM system SORCER to test the feasibility of the new model.
Ahmad Sholehin Bin Sarib, Haifeng Shen
CSCWD2
2013 Online silk road: nurturing social search through knowledge bartering
abstract
Social search empowers seekers to help each other find the information they need by sharing their domain knowledge and search efforts. Current social search activities are primarily voluntary, acted on the goodwill to help others or the purpose for self-promotion and the contributed content is mostly retrievable free of charge to the public. However, the voluntary nature of social search compromises its long-term sustainability as participants are not offered intrinsic incentives to contribute and share information, and free information presents intricate ramifications on the its quality. In this paper, we present the idea of knowledge bartering, where one can barter a knowledge item they have for another item they wish to have. To make the idea viable, we propose the online silk road solution to automate a knowledge bartering process that can maximise the social welfare within a community.
Haifeng Shen, Chengzheng Sun
CSCW2
2013 Achieving critical consistency through progressive slowdown in highly interactive Multi-Player Online Games
abstract
Multi-Player Online Games (MPOGs) have attracted enormous users in recent years. In contrast to the large number of low interactive MPOGs, the choices of highly interactive MPOGs are limited and most of them do not accept players whose network connections are slow due to the complexity involved in consistency maintenance. Most MPOGs adopt a data-centric consistency maintenance mechanism that only ensures the consistency of the final states, which however does not work well for highly interactive MPOGs because consistent views of incremental state changes are also important in those games. Instead of striving for a general solution for all highly interactive MPOGs, we advocate tailored solutions for different kinds of games by exploiting their semantics. To showcase this idea, we present the critical consistency model for highly interactive two-player online ball games, which only requires consistent views to be reached at critical states. We further contribute the progressive slowdown latency compensation technique to achieve critical consistency, which is demonstrated by an online table tennis game P2T2 (Peer-to- Peer Table Tennis).
Haifeng Shen, Suiping Zhou
CSCWD1
2012 Creative conflict resolution in realtime collaborative editing systems
abstract
Conflict is common in collaboration, and may have both negative and positive effects on collaborative work. Past research has focused on controlling negative aspects of conflict by preventing, eliminating or isolating conflicts, but done little on exploring positive aspects of conflict. In this paper, we contribute a novel creative conflict resolution (CCR) approach to address these issues in real-time collaborative editing systems. In addition to maintaining consistency, the CCR approach is able to create new results from conflicts, generate alternative solutions based on collective effects of conflict operations, and support users to choose suitable conflict solutions and conflict resolution policies according to their needs. The CCR approach provides not only a new way of resolving conflicts in real-time collaborative editing systems, but also a framework for supporting a range of existing conflict resolution strategies. Techniques and user interface issues related to the CCR approach and a prototype implementation are discussed in this paper.
David Sun, Chengzheng Sun, Steven Xia, Haifeng Shen
CSCW4
2012 ATCoPE: any-time collaborative programming environment for seamless integration of real-time and non-real-time teamwork in software development
abstract
Real-time collaborative programming and non-real-time collaborative programming are two classes of methods and techniques for supporting programmers to jointly conduct complex programming work in software development. They are complementary to each other, and both are useful and effective under different programming circumstances. However, most existing programming tools and environments have been designed for supporting only one of them, and little has been done to provide integrated support for both. In this paper, we contribute a novel Any-Time Collaborative Programming Environment (ATCoPE) to seamlessly integrate conventional non-real-time collaborative programming tools and environments with emerging real-time collaborative programming techniques and support collaborating programmers to work in and flexibly switch among different collaboration modes according to their needs. We present the general design objectives for ATCoPE, the system architecture, functional design and specifications, rationales beyond design decisions, and major technical issues and solutions in detail, as well as a proof-of-concept implementation of the ATCoEclipse prototype system.
Hongfei Fan, Chengzheng Sun, Haifeng Shen
GROUP3
2012 From credit and risk to trust: towards a credit flow based trust model for social networks
abstract
Trust management is a paramount issue in social networks. Existing models based on global reputation are simplistic as they do not support personalised measures for individual users. Models based on local trust propagation tend to be too subjective to be reliable as they do not consider a social network in its entirety. More importantly, neither model has taken the risk factor into the consideration of trust management. In this paper, we contribute a novel trust model that allows personalised measures to be naturally established on objective grounds through tracing credit flows within a social network, where the trust between a pair of users can be derived from the credit flowing from one into the other and the relative risk disparity between them. This model uses power flows in an electrical grid as a metaphor for the credit flows in a social network and is based on the hypothesis that the credit flows in a social network are similar in nature to the power flows in an electrical grid. Experiments with a real-world dataset have proved the hypothesis and the results have shown that the credit flow based trust model can derive not only personalised but also more accurate trust measures than existing models do.
Haifeng Shen, Chengzheng Sun
GROUP2
2012 Personalized multi-user view and content synchronization and retrieval in real-time mobile social software applications
Haifeng Shen, Mark D. Reilly
J. Comput. Syst. Sci.1
2011 A Probabilistic Topic Model with Social Tags for Query Reformulation in Informational Search
Haifeng Shen, Chengzheng Sun
ADMA (1)2
2011 Collaborative design: Improving efficiency by concurrent execution of Boolean tasks
Haifeng Shen, Chengzheng Sun
Expert Syst. Appl.2
2011 Achieving Data Consistency by Contextualization in Web-Based Collaborative Applications
abstract
Recent years have witnessed the emergence and rapid development of collaborative Web-based applications exemplified by Web-based office productivity applications. One major challenge in building these applications is maintaining data consistency while meeting the requirements of fast local response, total work preservation, unconstrained interaction, and customizable collaboration mode. These requirements are important in determining users’ experiences in interaction and collaboration, and in meeting users’ diverse needs under complex and dynamic collaboration and networking environments; but none of existing solutions is able to meet all of them. In this article, we present a data consistency maintenance solution capable of meeting these requirements for collaborative Web-based applications. Major technical contributions include an efficient sequence-based operation transformation control algorithm based on the concept of contextualization, an operation broadcast protocol for supporting a variety of collaboration modes, an operation replaying algorithm for ensuring fast local response and efficient remote operation replay, and a set of communication protocols for managing the integrity of collaborative Web-based sessions. The proposed solution has been implemented in a prototype collaborative Web-based editor WRACE and the correctness of the solution is formally verified in the article.
Haifeng Shen, Chengzheng Sun
ACM Trans. Internet Techn.1
2010 Supporting exploratory information seeking by epistemology-based social search
abstract
Formulating proper keywords and evaluating search results are common difficulties in exploratory information seeking. Reusing and refining others' successful searches are pragmatic directions to tackle these difficulties. In this paper, we present a novel epistemology-based social search solution, where search epistemologies are effectively shared, reused, and refined by others with the same or similar search interests through novel user interfaces. We have developed a prototype system Baijia and experimental results show that an epistemology-based social search system outperforms a conventional search engine in supporting exploratory information seeking.
Haifeng Shen, Chengzheng Sun
IUI2
2010 Inspiring Innovative Design Integration by Collaborative Exploration of Boolean Operations
abstract
Computer-Aided Design (CAD) applications provide industrial design communities with various computer-based tools to perform design activities. Innovation is an important requirement in most design tasks. This paper presents a novel approach to inspiring innovative design integration in a horizontal collaboration process. The approach allows individual designers' works to be integrated and multiple design choices to be explored. The technique is currently applied to design work involving Boolean operations-widely available in most CAD operations, where a major technical challenge is to integrate a group of concurrent Boolean operations for design exploration. The approach is particularly helpful to inspire innovation in the integration phases of industrial design tasks such as accessory design, art design, and digital media design.
Haifeng Shen, Chengzheng Sun
IEEE Trans. Ind. Informatics2
2009 Leveraging single-user AutoCAD for collaboration by transparent adaptation
abstract
It has been years since Computer-Aided Design (CAD) technology was used to assist engineers, architects and other design professionals in their design activities. Collaboration has been increasingly needed in the CAD community and various CSCW technologies have thus been adopted to support the development of collaborative design systems. Existing collaborative CAD applications suffer from several problems such as not allowing concurrent work and poor local response. In this paper, we report our approaches to converting single-user CAD applications into collaborative ones, based on the Transparent Adaptation (TA) technique. The approach is able to significantly improve productivity of collaborative designers as well as provide a human-centered collaborative design environment. To prove the feasibility of our approach, we have successfully converted the single-user AutoCAD application into a collaborative one, named as CoAutoCAD. Key issues on the design of CoAutoCAD are reported in this paper.
Haifeng Shen, Chengzheng Sun
CSCWD2
2009 Maintaining constraints of UML models in distributed collaborative environments
Haifeng Shen
J. Syst. Archit.1
2008 CoMaya: incorporating advanced collaboration capabilities into 3d digital media design tools
abstract
Complex 3D digital media creation demands anytime and anywhere collaboration support. The CoMaya project aims to incorporate such advanced collaboration capabilities into Autodesk Maya. This paper reports some research findings and lessons we learned from extending the transparent adaptation approach from 2D office applications to 3D digital media design tools.
Agustina, Steven Xia, Haifeng Shen, Chengzheng Sun
CSCW4
2008 Distributed Constraints Maintenance in Collaborative UML Modeling Environments
abstract
Constraints maintenance plays an important role in keeping the integrity and validity of models in UML software modeling. Constraints maintenance capabilities are reasonably adequate in UML modeling applications, but few work has been done to address the distributed constraints maintenance issue in collaborative UML modeling environments. In this paper, we propose a novel solution to the issue, which can retain the effects of all concurrent modeling operations even though they may cause constraints violations. We further contribute a distributed constraints maintenance framework, in which the solution is encapsulated as a generic engine to be mounted in a variety of single-user UML modeling applications for supporting distributed collaborative UML modeling and distributed constraints maintenance.
Haifeng Shen, Steven Xia, Chengzheng Sun
ASE1
2008 Dynamic Self-Healing for Service Flows with Semantic Web Services
abstract
With an increasing complexity of business processes, self-healing capability is becoming an important issue in order to support robust service flow execution. In this paper, a dynamic self-healing mechanism is proposed, which can dynamically identify suitable alternatives and replace faulty services such that a service flow can be performed successfully despite of unexpected exceptions. This mechanism explicitly utilizes semantic Web services for service matching and selection of a composite service in business service flow, and Semantic Web services are equipped with rich business rules in a domain-dependent manner. We explore the self-healing mechanism for supporting self-healable service flow execution which is modeled in BPEL4WS. A demo system of self-healing capable service flow execution is built to validate its effectiveness by a concrete scenario, PC manufacturing application.
Gang Chen 0002, Haifeng Shen, Jing-Bing Zhang, Chor Ping Low, David Chen 0002, Chengzheng Sun
Web Intelligence3
2008 Ant Colony Inspired Self-Healing for Resource Allocation in Service-Oriented Environment Considering Resource Breakdown
abstract
The ant colony optimization (ACO) algorithm is a metaheuristic inspired from the behavior of foraging ants. Instead of exploring its ability in finding optimal solutions, the current study investigates another unique property - self-healing mechanism for resource allocation in a service-oriented environment where unexpected resource breakdown can occur. A system architecture is first proposed to detect, diagnose and react to disturbances. Then the performance of the ACO self-healing mechanism is tested and compared based on a modified benchmark problem. The experimental results show that the self-healing mechanism can promptly recover an obsolete schedule with high quality solutions.
Rong Zhou 0004, Ren Wei, Gang Chen 0002, Haifeng Shen, Jing-Bing Zhang, Ming Luo 0003
Web Intelligence5
2007 Integrating Advanced Collaborative Capabilities into Web-Based Word Processors
Haifeng Shen, Steven Xia, Chengzheng Sun
CDVE1
2006 Real-Time Collaborative Software Modeling Using UML with Rational Software Architect
abstract
Modeling is commonly used in the process of software development. UML (Unified Modeling Language) is a standard software modeling language and has been widely adopted for software analysis and design. As software systems are getting larger and more complex nowadays, software modeling using UML often requires collective and collaborative efforts from multiple software designers. In contrast, most of today's software modeling applications are still single-user-oriented and do not offer much help to coordinate interaction and collaboration among team members. In this paper, we will present the technical challenges and solutions in providing advanced collaboration capabilities and transparently integrating them into mainstream software modeling applications to effectively facilitate collaboration among geographically dispersed software designers. The work has been tested and demonstrated by the design of CoRSA (Collaborative Rational Software Architect) - an experimental collaborative software modeling prototype based on RSA, one of the most widely used commercial software modeling applications in the market
Haifeng Shen, Steven Xia, Chengzheng Sun
CollaborateCom3
2006 A Generic WebDAV-Based Document Repository Manager for Collaborative Systems
abstract
Collaborative document repository manager (CDRM) is one of the three key components in a collaborative editing system. It provides a repository for storing shared documents and a portal for managing shared documents, initiating collaborative editing sessions, and viewing session information. While most collaborative systems have been using ad hoc and tightly-coupled CDRMs, our work is to provide a generic WebDAV-based CDRM to take advantage of WebDAV's ubiquitous accessibility, directly writable medium, and access transparency. This WebDAV-based CDRM has been deployed into the CoOffice (collaborative office) and the CoStarOffice (collaborative staroffice) systems for evaluation, demonstration, and usability study
Haifeng Shen, Chengzheng Sun, Suiping Zhou, Zaw Wai Phyo
Web Intelligence1
2006 Model-Based Feature Compensation for Robust Speech Recognition
Haifeng Shen, Qunxia Li, Jun Guo 0002, Gang Liu 0008
Fundam. Informaticae1
2006 Transparent adaptation of single-user applications for multi-user real-time collaboration
abstract
Single-user interactive computer applications are pervasive in our daily lives and work. Leveraging single-user applications for supporting multi-user collaboration has the potential to significantly increase the availability and improve the usability of collaborative applications. In this article, we report an innovative Transparent Adaptation (TA) approach and associated supporting techniques that can be used to convert existing and new single-user applications into collaborative ones, without changing the source code of the original application. The cornerstone of the TA approach is the operational transformation (OT) technique and the method of adapting the single-user application programming interface to the data and operation models of OT. This approach and supporting techniques were developed and tested in the process of transparently converting two commercial off-the-shelf single-user applications (Microsoft Word and PowerPoint) into real-time collaborative applications, called CoWord and CoPowerPoint, respectively. CoWord and CoPowerPoint not only retain the functionalities and “look-and-feel” of their single-user counterparts, but also provide advanced multi-user collaboration capabilities for supporting multiple interaction paradigms, ranging from concurrent and free interaction to sequential and synchronized interaction, and for supporting detailed workspace awareness, including multi-user telepointers and radar views. The TA approach and generic collaboration engine software component developed from this work are potentially applicable and reusable in adapting a wide range of single-user applications.
Chengzheng Sun, Steven Xia, David Sun, David Chen 0002, Haifeng Shen, Wentong Cai 0001
ACM Trans. Comput. Hum. Interact.5
2005 Syntax-based reconciliation for asynchronous collaborative writing
abstract
With the rapid popularity of computer networks, collaborative document writing becomes increasingly desirable in recent years. In practice, collaborative writing is more likely to be done in an asynchronous manner, where collaborators usually work in parallel with different time schedules and are not present at the same time. Merging is the key technique to support concurrent writing and textual merging remains the primary and the only successful merging function to date. However, most of existing systems support constrained textual merging for the simplicity of the underling merging algorithms. In this paper, we propose a flexible operation-based syntactic textual merging algorithm that is capable of reconciling changes made in parallel by different users according to the syntax of the files to be merged or user-specified merging policies. Moreover, this syntax-based reconciliation algorithm is able to preserve the intentions of individual changes.
Haifeng Shen, Chengzheng Sun
CollaborateCom1
2005 Two-Domain Feature Compensation for Robust Speech Recognition
Haifeng Shen, Gang Liu 0008, Jun Guo 0002, Qunxia Li
ISNN (2)1
2005 Non-stationary Environment Compensation Using Sequential EM Algorithm for Robust Speech Recognition
Haifeng Shen, Jun Guo 0002, Gang Liu 0008, Qunxia Li
PKDD1
2004 A Complete Textual Merging Algorithm for Software Configuration Management Systems
abstract
Software configuration management (SCM) systems are very important for coordinating group efforts in developing large and complex software systems. The ability to support concurrent software development is the key to deliver high quality software with low time-to-market, where merging is the core enabling technique. Textual merging is the primary and the only successful merging function available in today's SCM systems. However none of them supports complete textual merging, which is not only very useful itself but also the foundation for syntactic and semantic textual merging. We propose a novel operation-based textual merging algorithm, which has the capability of supporting complete textual merging while still preserving the intentions of individual editing operations.
Haifeng Shen, Chengzheng Sun
COMPSAC1
2004 Leveraging single-user applications for multi-user collaboration: the coword approach
abstract
Single-user interactive computer applications are pervasive in our daily lives and work. Leveraging single-user applications for multi-user collaboration has the potential to significantly increase the availability and improve the usability of collaborative applications. In this paper, we report an innovative transparent adaptation approach for this purpose. The basic idea is to adapt the single-user application programming interface to the data and operational models of the underlying collaboration supporting technique, namely Operational Transformation. Distinctive features of this approach include: (1) Application transparency: it does not require access to the source code of the single-user application; (2) Unconstrained collaboration: it supports concurrent and free interaction and collaboration among multiple users; and (3) Reusable collaborative software components: collaborative software components developed with this approach can be reused in adapting a wide range of single-user applications. This approach has been applied to transparently convert MS Word into a real-time collaborative word processor, called CoWord, which supports multiple users to view and edit any objects in the same Word document at the same time over the Internet. The generality of this approach has been tested by re-applying it to convert MS PowerPoint into CoPowerPoint.
Steven Xia, David Sun, Chengzheng Sun, David Chen 0002, Haifeng Shen
CSCW5
2004 Improving real-time collaboration with highlighting
Haifeng Shen, Chengzheng Sun
Future Gener. Comput. Syst.1
2002 A Log Compression Algorithm for Operation-based Version Control Systems
abstract
Version control systems are widely used to support distributed concurrent software development, where document merging is a key function. Most existing systems adopt state-based merging, which relies on the derivation of deltas among documents. The derivation of deltas involves transferring documents over the network and executing time-consuming text differentiation algorithms, which may result in a poor system response. Operation-based merging saves executed operations in logs as deltas, thus eliminating the need for deriving deltas. However, for the operation-based merging to be adopted in version control systems, a major technical challenge is how to keep the size of logs small so that it requires less time to transfer the log over the network and to re-execute operations in the log. In this paper we contribute a novel compression algorithm, which is able to minimize the size of a log as well as the number of operations within it. It has been proven both correct and complete in the sense that the compressed log has the same effect as the original one and operations that can be merged have already been merged.
Haifeng Shen, Chengzheng Sun
COMPSAC1
2002 Flexible notification for collaborative systems
abstract
Notification is an essential feature in collaborative systems, which determines a system's capability and flexibility in supporting different kinds of collaborative work. In the past years, various notification strategies have been designed for different systems. However, the design of notification components has been ad hoc, and the techniques used for supporting notification have been application-dependent. In this paper, we contribute a flexible notification framework that can be used to describe and compare a range of notification strategies used in existing collaborative systems, and to guide the design of notification components for new collaborative systems. The framework has been applied to the design of a notification component for a group editor, which uses a single notification mechanism to support various notification policies for meeting both real-time and non-real-time collaboration needs. In addition, a new operational transformation control algorithm has been devised in combination with the notification component, which is significantly simpler and more efficient than existing algorithms.
Haifeng Shen, Chengzheng Sun
CSCW1