Bo Cai 0003

dblp:32/3878-3 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-5261-0191ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Modeling Structural Features with Structure Position-Aware Attention for Code Summarization
abstract
The automatic generation of code comments is a crucial aspect in software engineering as it allows the description of source code functions in natural language. In recent years, researchers have utilized advanced machine translation models, particularly transformer-based models, to achieve remarkable results in the task of code summarization. Despite these advancements, current methods still face some challenges. First, existing models fail to incorporate the rich structural information of the code well, resulting in inadequate summary statements. Second, most generative models are too sensitive to the editing of code text, which results in insufficient generalization ability of the model. To address these issues, we proposed a new model named SPA-Trans(Structure Position-aware Attention Transformer-based model). SPA-Trans uses a distance matrix based on the relative distance of nodes in the AST (Abstract Syntax Tree) to represent the structural correlation between each token, enhancing the model’s ability to capture structural information of the source code. Additionally, in order to reduce the sensitivity of the model to code editing, we used the adversarial training method in the embedding layer to simulate code editing, which improves the generalization of the model. Our experiments on real-world datasets in Java and Python validate the effectiveness of our proposed method.
Zhiheng Qu, Bo Cai 0003
Int. J. Softw. Eng. Knowl. Eng.3
2026 Semantic image synthesis for unknown object detection in autonomous driving scenes
Aihua Ke, Bo Cai 0003
Inf. Sci.3
2025 Enhancing Fault Localization via Causal Measurement Features
abstract
Fault localization is the task of automatically identifying the program components (e.g., files, functions, or lines of code) associated with specific defects. Existing studies on fault localization utilize statistical methods or neural networks to establish correlations between program components and defects. However, the results obtained by these methods may be influenced by confounding factors, as they do not explicitly measure the causal relationships between fault locations and defects. Considering the lack of causal correlation measurement in existing learning-based fault localization methods, we propose Causal Measurement based Fault Localization (CMFL) method. CMFL is the first learning-based approach that focuses on identifying the root causes triggering defects and measuring the causal associations between fault locations and defects, thus introducing causal inference theory into learning-based fault localization approaches. We apply counterfactual theory and causal intervention to compute the causality between program components and defects, thereby extracting causal measurement features of the program. Additionally, we extract semantic and structural features of the code using pre-trained language models. By fusing these features, we achieve a comprehensive approach that considers causality, code semantics, and structure, enabling more effective fault localization. Our extensive experiments on the widely used benchmark Defects4J show that CMFL outperforms existing traditional and learning-based fault localization baselines. Furthermore, by analyzing the rankings of the localized faults, we show that the root cause of defects can be more accurately and directly identified through causal measurement.
Bo Cai 0003, Chengqing Cai, Yangbo Qiu
IJCNN2
2025 Learning hierarchical scene graph and contrastive learning for object goal navigation
Bo Cai 0003, Yaoxiang Yu, Aihua Ke
Knowl. Based Syst.3
2025 Learning Bird's Eye View scene graph and knowledge-inspired policy for embodied visual navigation
Jie Yang 0076, Siwei Huang, Bo Cai 0003
Knowl. Based Syst.5
2024 Learning multimodal adaptive relation graph and action boost memory for visual navigation
Bo Cai 0003, Yaoxiang Yu, Aihua Ke
Adv. Eng. Informatics2
2024 Text-guided image-to-sketch diffusion models
Aihua Ke, Jie Yang 0076, Bo Cai 0003
Knowl. Based Syst.4
2023 Prompt Learning for Developing Software Exploits
abstract
A software exploit is a sequence of commands that exploits software vulnerabilities or security flaws, written either by security researchers as a Proof-Of-Concept (POC) threat or by malicious attackers for use in their operations. Writing exploits is a challenging task since it is time-consuming and costly. Pre-trained Language Models (PLMs) for code can benefit automatic exploit generation, achieving state-of-the-art performance. However, the typical paradigm for using these PLMs is fine-tuning, which leads to a significant gap between the language model pre-training and the target task fine-tuning process since they are in different forms. In this paper, we propose a prompt learning approach PT4Exploits based on the PLM, i.e., CodeT5 to automatically generate desired exploits by modifying the original English description inputs. More specifically, PT4Exploits can better elicit knowledge from the PLM by inserting trainable prompt tokens into the original input to construct the same form as the pre-training tasks. Experimental results show that our approach can significantly outperform baseline models on both automatic evaluation and human evaluation. Additionally, we conduct extensive experiments to investigate the performance of prompt learning on few-shot settings and the scenario of different prompt templates, which also show the competitive effectiveness of PT4Exploits.
Xiaoming Ruan, Yaoxiang Yu, Bo Cai 0003
Internetware4
2023 Adapting Hierarchical Transformer for Scene-Level Sketch-Based Image Retrieval
abstract
Sketch-based image retrieval (SBIR) is an essential application of sketches. Research on object-level SBIR is relatively mature, but the study of more complex scene-level SBIR is still in its early stages. In order to advance this research, we investigate previous works and identify two main shortcomings: (1) insufficient utilization of multi-scale features from sketches and images, and (2) lack of effective modules to eliminate the substantial domain gap between them. To address these issues, we propose SketchRetriever, a hierarchical Transformer-based scene-level SBIR model. In our model, the hierarchical Transformer and compressors are capable of efficiently capturing feature maps at various granularities and compressing them into corresponding feature vectors, and the modality-specific Adapters can project the feature embeddings of sketches and images into the same feature space, thereby closing the domain gap between them. We adopt the adapter-tuning strategy, which not only considerably reduces the number of tunable parameters but also effectively avoids overfitting. Extensive experiments demonstrate that SketchRetriever significantly outperforms state-of-the-art methods on two benchmark datasets with lower fine-tuning overhead.
Jie Yang 0076, Aihua Ke, Bo Cai 0003
MMAsia3
2023 Scene sketch semantic segmentation with hierarchical Transformer
Jie Yang 0076, Aihua Ke, Yaoxiang Yu, Bo Cai 0003
Knowl. Based Syst.4
2022 Multi-Perspective Alignment Mechanism for Code Search
abstract
Developers often tend to search and reuse code snippets from large code repositories to improve their programming skills. To support code reuse, early code search models used information retrieval (IR) techniques to index a large corpus of code and then returned relevant code based on search statements. However, IR-based models can cause information loss due to keyword matching. To solve this problem, developers applied deep learning (DL) techniques to encode search models. However, these models either learn isolated representations of codes and queries or learn interdependent representations of codes and queries in a single way, which both limit the effectiveness of the models.In this study, we propose a code search model MpaCS, which can interact from multiple perspectives to enable cross-language retrieval between code and query statements. MpaCS extracts code and query information from the code textual features (i.e., method name, and tokens), the code structural feature (i.e., abstract syntax tree), and the query feature (i.e., tokens). We first implement the alignment between code features and query features at the fragment level. Then we propose a context matching module to achieve a deeper interaction between the code and the query. Finally, we propose a new attention network to explore the maximum alignment relationship between code and query from another perspective. We evaluate the performance of MpaCS on two existing large-scale datasets with 69k and 105k code snippets, respectively. Experimental results show that MpaCS outperforms four state-of-the-art models DeepCS [1],UNIF [2], MPCAT [3] and CARLCS-CNN [4].
Bo Cai 0003
APSEC2
2022 Method Name Generation Based on Code Structure Guidance
abstract
The proper names of software engineering functions and methods can greatly assist developers in understanding and maintaining the code. Most researchers convert the method name generation task into the text summarization task. They take the token sequence and the abstract syntax tree (AST) of source code as input, and generate method names with a decoder. However, most proposed models learn semantic and structural features of the source code separately, resulting in poor performance in the method name generation task. Actually, each token in source code must have a corresponding node in its AST. Inspired by this observation, we propose SGMNG, a structure-guided method name generation model that learns the representation of two combined features. Additionally, we build a code graph called code relation graph (CRG) to describe the code structure clearly. CRG retains the structure of the AST of source code and contains data flows and control flows. SGMNG captures the semantic features of the code by encoding the token sequence and captures the structural features of the code by encoding the CRG. Then, SGMNG matches tokens in the sequence and nodes in the CRG to construct the combination of two features. We demonstrate the effectiveness of the proposed approach on the public dataset Java-Small with 700K samples, which indicates that our approach achieves significant improvement over the state-of-the-art baseline models in the ROUGE metric.
Zhiheng Qu, Jianhui Zeng, Bo Cai 0003
SANER4
2022 Novel spatial and temporal interpolation algorithms based on extended field intensity model with applications for sparse AQI
Bo Cai 0003, Zeyuan Shi, Jianhui Zhao 0001
Multim. Tools Appl.1
2022 A computer-aided diagnostic system for mammograms based on YOLOv3
Jianhui Zhao 0001, Tianquan Chen, Bo Cai 0003
Multim. Tools Appl.3
2021 GUIS2Code: A Computer Vision Tool to Generate Code Automatically from Graphical User Interface Sketches
Jiaqi Fang, Bo Cai 0003, Yingtao Zhang
ICANN (3)3
2019 A dynamic texture based segmentation method for ultrasound images with Surfacelet, HMT and parallel computing
Bo Cai 0003, Wei Ye 0007, Jianhui Zhao 0001
Multim. Tools Appl.1
2019 A methodology for 3D geological mapping and implementation
Bo Cai 0003, Jianhui Zhao 0001, Xiangyu Yu
Multim. Tools Appl.1
2018 APs Deployment Optimization for Indoor Fingerprint Positioning with Adaptive Particle Swarm Algorithm
Jianhui Zhao 0001, Haojun Ai, Bo Cai 0003
ICA3PP (3)4
2018 Deployment Optimization of Indoor Positioning Signal Sources with Fireworks Algorithm
Jianhui Zhao 0001, Shiqi Wen, Haojun Ai, Bo Cai 0003
ICA3PP (3)4