VLDB 2026 Research / reviewers in the wild / expert
Haibo Hu 0002
dblp:90/5236-2
· DBLP profile ↗
17ranked-venue papers
2as first author
8since 2021 · last 2024
0000-0001-8442-5222ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Computer networks · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DFEPT: Data Flow Embedding for Enhancing Pre-Trained Model Based Vulnerability DetectionabstractSoftware vulnerabilities represent one of the most pressing threats to computing systems. Identifying vulnerabilities in source code is crucial for protecting user privacy and reducing economic losses. Traditional static analysis tools rely on experts with knowledge in security to manually build rules for operation, a process that requires substantial time and manpower costs and also faces challenges in adapting to new vulnerabilities. The emergence of pre-trained code language models has provided a new solution for automated vulnerability detection. However, code pre-training models are typically based on token-level large-scale pre-training, which hampers their ability to effectively capture the structural and dependency relationships among code segments. In the context of software vulnerabilities, certain types of vulnerabilities are related to the dependency relationships within the code. Consequently, identifying and analyzing these vulnerability samples presents a significant challenge for pre-trained models. Zhonghao Jiang, Weifeng Sun 0004, Tao Wen 0012, Haibo Hu 0002, Meng Yan 0001 |
Internetware | 6 |
| 2024 | VisRepo: A Visual Retrieval Tool for Large-Scale Open-Source ProjectsabstractTo improve software development productivity, developers frequently search for projects on open-source communities such as GitHub. However, it is challenging for users to quickly find suitable projects from numerous results due to the overload of project information. Although many tools have been proposed to rank the relevancy of searched results, manually inspecting them one by one is irreplaceable and time-consuming. To fill this gap, we propose a visual retrieval tool named VisRepo for open-source software projects. Firstly, it mines software project data from four perspectives including topic, technology, usability, and comprehensibility, and connects projects based on the same owners/contributors and similar topics. Then, visualization technique is employed to present complex software data intuitively. VisRepo provides users an interactive retrieval paradigm of Search-Explore-Check-Recommend with in-depth insights and better exploration experience. We evaluate VisRepo on 7w+ open-source JavaScript projects. Experimental results showed that VisRepo outperforms GitHub search engine in terms of time consumption and accuracy, meanwhile enabling a more interactive and useful user experience. Xiaoqi Yue, Chao Liu 0014, Neng Zhang 0001, Haibo Hu 0002, Xiaohong Zhang 0002 |
Internetware | 4 |
| 2024 | Guiding ChatGPT for Better Code Generation: An Empirical StudyabstractAutomated code generation is a powerful technique for software development, which can significantly reduce developers' effort and time for writing code. Recently, OpenAI's large language model ChatGPT has emerged as a powerful tool for generating human-like responses to a wide range of textual inputs (i.e., prompts), including those related to code generation. However, the effectiveness of ChatGPT in code generation is still not well understood. The code generation performance could also be heavily influenced by the choice of prompts, which should be further explored. In this paper, we report an empirical study on ChatGPT's capabilities for two types of code generation tasks, namely text-to-code and code-to-code generation. We investigate different types of prompts by leveraging the chain-of-thought strategy with multi-step optimizations. Our empirical results show that by carefully designing prompts to guide ChatGPT, the code generation performance can be improved substantially. We also analyze the factors that influence the prompt design and provide insights that could guide future research. Chao Liu 0014, Xuanlin Bao, Hongyu Zhang 0002, Neng Zhang 0001, Haibo Hu 0002, Xiaohong Zhang 0002, Meng Yan 0001 |
SANER | 5 |
| 2024 | Understanding the implementation issues when using deep learning frameworks
Chao Liu 0014, Runfeng Cai, Yiqun Zhou, Haibo Hu 0002, Meng Yan 0001 |
Inf. Softw. Technol. | 5 |
| 2024 | TransforLearn: Interactive Visual Tutorial for the Transformer ModelabstractThe widespread adoption of Transformers in deep learning, serving as the core framework for numerous large-scale language models, has sparked significant interest in understanding their underlying mechanisms. However, beginners face difficulties in comprehending and learning Transformers due to its complex structure and abstract data representation. We present TransforLearn, the first interactive visual tutorial designed for deep learning beginners and non-experts to comprehensively learn about Transformers. TransforLearn supports interactions for architecture-driven exploration and task-driven exploration, providing insight into different levels of model details and their working processes. It accommodates interactive views of each layer's operation and mathematical formula, helping users to understand the data flow of long text sequences. By altering the current decoder-based recursive prediction results and combining the downstream task abstractions, users can deeply explore model processes. Our user study revealed that the interactions of TransforLearn are positively received. We observe that TransforLearn facilitates users' accomplishment of study tasks and a grasp of key concepts in Transformer effectively. Zekai Shao 0001, Ziqin Luo, Haibo Hu 0002, Cagatay Turkay, Siming Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Unified abstract syntax tree representation learning for cross-language program classificationabstractProgram classification can be regarded as a high-level abstraction of code, laying a foundation for various tasks related to source code comprehension, and has a very wide range of applications in the field of software engineering, such as code clone detection, code smell classification, defects classification, etc. The cross-language program classification can realize code transfer in different programming languages, and can also promote cross-language code reuse, thereby helping developers to write code quickly and reduce the development time of code transfer. Most of the existing studies focus on the semantic learning of the code, whilst few studies are devoted to cross-language tasks. The main challenge of cross-language program classification is how to extract semantic features of different programming languages. In order to cope with this difficulty, we propose a Unified Abstract Syntax Tree (namely UAST in this paper) neural network. In detail, the core idea of UAST consists of two unified mechanisms. First, UAST learns an AST representation by unifying the AST traversal sequence and graph-like AST structure for capturing semantic code features. Second, we construct a mechanism called unified vocabulary, which can reduce the feature gap between different programming languages, so it can achieve the role of cross-language program classification. Besides, we collect a dataset containing 20,000 files of five programming languages, which can be used as a benchmark dataset for the cross-language program classification task. We have done experiments on two datasets, and the results show that our proposed approach outperforms the state-of-the-art baselines in terms of four evaluation metrics (Precision, Recall, F1-score, and Accuracy). Kesu Wang, Meng Yan 0001, Haibo Hu 0002 |
ICPC | 4 |
| 2022 | A Secure Task Matching Scheme in Crowdsourcing Based on Blockchain
Jiajun Chen 0003, Chunqiang Hu, Haibo Hu 0002 |
WASA (2) | 5 |
| 2022 | CrossNet: Detecting Objects as CrossesabstractWith the use of deep learning, object detection has achieved great breakthroughs. However, existing object detection methods still can not cope with challenging environments, such as dense objects, small objects, and object scale variations. To address these issues, this paper proposes a novel keypoint-based detection framework, called CrossNet, which significantly improves detection performance with minimal costs. In our approach, an object is modeled as a cross that consists of a center keypoint and a specific size, which eliminates the need of hand-craft anchor design. The proposed CrossNet outputs three types of maps: the center map, size map, and offset map, where both center map and offset map are to predict the center keypoints of objects and the size map is to estimate the sizes (width and height) of objects. Specifically, we first design a cascaded center prediction method that introduces a coarse-to-fine idea to improve center prediction. Furthermore, since center prediction considered as a classification task is easier than size regression relatively, we design a center-attention size regression module that uses the detection results of centers to assist the size prediction. In addition, a slightly modified hourglass network is designed to enhance the quality of feature maps for center and size prediction. Extensive experiments are conducted to demonstrate the effectiveness of CrossNet on the challenging PASCAL VOC, COCO, KITTI, and WiderFace datasets. Empirical studies show that CrossNet achieves competitive results with top-ranked one-stage and two-stage detectors while being time-efficient. Jiaxu Leng, Ying Liu 0039, Zhihui Wang 0003, Haibo Hu 0002, Xinbo Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | SIOT-RIMM: Towards Secure IOT-Requirement Implementation Maturity ModelabstractIt is very crucial for an organization to encapsulate the requirements in its early stage when they are intending to build a novel system such as the internet of things (IoT), particularly when it comes to capturing privacy and security requirements to gain the public confidence. The proposed research is focused to develop a secure IoT-requirement implementation maturity model (SIOT-RIMM). The proposed model will assist the software development organizations to improve and modify their requirement engineering processes in terms of security and privacy of IoT. The SIOT-RIMM model will be developed based on the existing IoT literature pertaining to security and privacy, industrial empirical study and understanding of the challenges that could negatively influence the implementation of security and privacy in IoT. To develop the maturity levels of SIOT-RIMM, we will consider the concepts of existing maturity models of other software engineering domains. In this preliminary study, 19 challenges were identified using the SLR approach that might have a negative impact on the IoT requirements engineering process. The identified challenges will contribute to the development of SIOT-RIMM maturity levels. Haibo Hu 0002, Muhammad Azeem Akbar, Yasir Hussain, Ali Mahmoud Baddour |
EASE | 2 |
| 2019 | Forward Engineering Completeness for Software by Using Requirements Validation Framework (S)abstractIn software development environment, software companies usually ignore the user requirements validation process in requirement gathering phase, which results in large number of modifications being required in the software maintenance phase to fulfill the customer requirements.Identification of accurate requirements from user stories and determining the effectiveness of work deliverable of software industry has always been a challenging task.In this paper, a new measurement approach for forward engineering completeness for software was introduced by using requirements validation framework.The forward engineering completeness for software was measured in two steps.In the first step, software component structure was developed in order to find the functional and nonfunctional requirements rejected by the customers in the requirement validation framework.In the second step, completeness of software from component-based development was determined in which the following parameters, such as functional, non-functional completeness attributes, were considered in the measurement process, and the unadopted attributes of the reuse code were also considered.Quality level for the attributes were assigned based upon the valuation of interior quality of the source code.Therefore, it resulted in the reduction of development time required for the software and the cost required for the software development was also reduced.A case study was incorporated in this research to explain the measurement process of forward engineering completeness.If the forward engineering code is satisfying the quality standards, then the code is in the completeness form.The attributes of code that negates to be used were considered as unadopted attributes. Nayyar Iqbal, Jun Sang, Haibo Hu 0002, Hong Xiang |
SEKE | 4 |
| 2019 | An Efficient and Recoverable Data Sharing Mechanism for Edge Storage
Yuwen Pu, Feihong Yang, Chunqiang Hu, Haibo Hu 0002 |
WASA | 6 |
| 2019 | Background Subtraction Based on Integration of Alternative Cues in Freely Moving CameraabstractPrevious approaches to background subtraction in freely moving camera typically focus on improving the accuracy of motion estimation. In this paper, we propose that the accurate background subtraction is possible with the integration of alternative cues about foreground and background. We also put forward a novel background subtraction framework called the integration of foreground and background cues. Here, the foreground cues are extracted by the Gaussian mixture model compensated with image alignment, while the background cues are obtained from the spatiotemporal features filtered by the homography transformation. Subsequently, the integration is devised as a hierarchical competition procedure based on super-pixels under multiple levels with the underlying motivation to utilize the exclusiveness between these cues for the compensation of their corresponding defects. The result of competition between foreground and background cues in a particular super-pixel is used as the proximity, and the foreground is segmented by combining super-pixels with proximity under multiple levels. Comprehensive evaluations using standard benchmarks demonstrate the superiority of our work compared with the state-of-the-art. Chenqiu Zhao, Aneeshan Sain, Ying Qu 0007, Yongxin Ge, Haibo Hu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Revisiting the Correlation Between Alerts and Software Defects: A Case Study on MyFaces, Camel, and CXFabstractStatic analysis tools (e.g., FindBugs) are widely used to detect potential defects in software development. A recent study suggests that there is a moderate correlation between the alerts reported by static analysis tools and software defects [1]. However, despite the actionable alerts reported by static analysis tools, they may report too many meaningless unactionable alerts. Actionable alert refers to the alert which is meaningful and fixable. Unactionable alert (i.e., false positive alert) refers to the alert which is regarded as unimportant to developers, inessential to source code, or will not be fixed by developers. Are all alerts (including both actionable and unactionable alerts) suitable for indicating software defects? To address this question, we classify all the alerts into two categories, namely actionable alerts and unactionable alerts. By the following, we conduct an empirical study to evaluate the degree of correlation between defects and alerts on the evolution of three open source projects with totally 40 releases. The objective of the study is to explore two kinds of correlation analysis: one is the correlation between all the alerts reported by FindBugs and defects among the release history of a project, the other is the correlation between the actionable alerts and defects. As a result, we find that not all the alerts but the actionable alerts are suitable to be an early predictor of defects. Meng Yan 0001, Xiaohong Zhang 0002, Haibo Hu 0002, Xin Xia 0001 |
COMPSAC (1) | 4 |
| 2017 | An approach to translating OCL invariants into OWL 2 DL axioms for checking inconsistency
Chunlei Fu, Dan Yang 0001, Xiaohong Zhang 0002, Haibo Hu 0002 |
Autom. Softw. Eng. | 4 |
| 2013 | Detection of local invariant features using contourabstractThis study proposes a new method for the detection of local invariant features with contour. This method differs from traditional methods that use image intensity. Image contours can be extracted stably with changes in viewpoint, scale, illumination and other factors. The proposed algorithm first extracts the stable corner from the contour, then it fits the supporting region of the contour near the corner to an angle, and uses its bisector as the direction of the feature. Next, it searches the contour for the tangent point in the direction of the angle bisector. Finally, with the corner as the centre, and in combination with the tangent point and the feature direction, an elliptic invariant region is constructed. The feasibility of the algorithm was verified experimentally by comparing its repetition rate. Test images obtained from actual scenes include several types of transformations, such as rotation, scaling, affinity, illumination and noise. The results of the experiment show the feasibility of the proposed method for use in local invariant features detection. Haibo Hu 0002, Xiaoze Lin, Xiaohong Zhang 0002 |
IET Image Process. | 1 |
| 2013 | Scalable RDF Graph Querying Using Cloud Computing
Dan Yang 0001, Haibo Hu 0002, Juan Xie |
J. Web Eng. | 3 |
| 2011 | Semantic Web-based policy interaction detection method with rules in smart home for detecting interactions among user policiesabstractThe emerging technologies such as the Internet-of-Things, sensors, communication networks, have been or will be introduced to conventional domotics to provide a wide variety of smart home services to facilitate the household appliances or home cares and improve the lifestyles of people. Currently, smart home system are integrated with different features from product line and equipped with various sensors and actuators to meet the requirements of house occupants by specifying their customised user policies. However, the introduction of features and policies may result in undesired behaviours, and this effect is known as feature interactions. In this study, the authors proposed a Semantic Web-based policy interaction detection method with rules to model smart home services and policies with the aids of ontological analysis in the smart home domain, so as to construct a semantic context for inferring the interaction of policies. The authors focus their work on user policies interaction, which are detected by using the Semantic Web rule language in semantic context. The approach is successfully applied to the smart home system and is able to detect 90 interactions among 32 user policies by automated reasoning with tools support as Protégé and Jess. Haibo Hu 0002, Dan Yang 0001, Hong Xiang, Chunlei Fu, Jun Sang, Chunxiao Ye |
IET Commun. | 1 |