VLDB 2026 Research / reviewers in the wild / expert
Zhixiang Cao
dblp:356/8697
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PDRNet: Pinwheel-Guided Dynamic Representation Learning for Visible-Infrared Person Re-IdentificationabstractABSTRACT Visible‐infrared person re‐identification (VI‐ReID) is a cross‐modal retrieval task characterized by significant challenges, with the objective of precisely identifying and matching pedestrian instances across different spectral modalities, namely visible‐light and infrared imagery. The primary difficulties stem from substantial inter‐modal discrepancies and intra‐modal feature variations, which complicate effective cross‐modal matching. While existing approaches generally focus on embedding heterogeneous modal data into a unified feature space to extract shared representations, they often overlook the discriminative identity information embedded within modality‐specific features. To overcome this inherent limitation, we propose a novel pinwheel‐guided dynamic representation network (PDRNet), designed to mine and enhance the directional structural cues and scale‐sensitive discriminative features inherent in modality‐specific representations. Specifically, we integrate direction‐aware pinwheel convolution (PConv) into a two‐stream architecture to strengthen local structural representation and guide the learning of shared semantic features, thereby improving both the discriminability and structural modeling of modality‐specific information. Furthermore, to accommodate scale disparities across modalities and individuals, we incorporate a scale‐based dynamic loss (SD loss), which adaptively adjusts the loss weights related to scale and positional information. This mechanism mitigates the error amplification often observed in small‐scale samples and enhances both the discriminative power and robustness of cross‐modal matching across varying scales. We perform extensive experiments on multiple well‐established public benchmarks. The results consistently show that the proposed PDRNet achieves superior performance compared to existing methods in both recognition accuracy and cross‐modal matching effectiveness. Fengshan Lai, Zhixiang Cao, Rongyu Jia, Daoxun Xia |
Concurr. Comput. Pract. Exp. | 2 |
| 2026 | Deep Learning Framework Testing via Model Mutation: How Far Are We?abstractDeep Learning (DL) frameworks are fundamental components of DL systems in their development, deployment, and execution, while defects in DL frameworks can cause severe consequences. Ensuring the quality of DL frameworks has therefore become a pressing challenge. Among the various testing techniques, model mutation has emerged as a widely adopted approach. Such methods generate mutants by applying mutation operators to DL models (e.g., structural changes or parameter edits) and then analyzing inconsistencies, crashes, or abnormal behaviors across different frameworks or hardware. Despite its effectiveness, existing methods suffer from the following limitations. First, they mainly reuse operators designed for model testing, raising doubts about their ability to expose framework-level defects. Besides, they insufficiently consider mutation constraints, such as mutation type, position, and order, which directly affect the defect detection ability of generated mutants. Finally, they rely on the limited detection range and narrow test oracles, focusing on functional correctness in model inference while overlooking defects in efficiency, resource usage, and other defects that developers care about in other stages, such as model training or deployment. These limitations result in a weak alignment with the critical defects that developers are most concerned about in practice. Motivated by these observations, this study conducts a comprehensive investigation into the effectiveness of existing mutation-based testing methods. We first collect and classify defect reports from PyTorch and MindSpore according to developers’ priority tags, building a taxonomy of seven categories and 19 sub-categories of HP defects. We then map the defects reported by five state-of-the-art methods into this taxonomy to evaluate their detection abilities. To explain these limitations, we further analyze how three key factors, mutation type, mutation position, and mutation order, affect the generated mutants. Based on the experiment results, we summarize ten findings ranging from revealing the priority of developers on fixing framework defects, evaluating the defect detection ability of existing methods, to how mutation factors affect the generated mutants. Furthermore, we reveal four limitations and their root causes of existing methods and propose four targeted optimization strategies. We further apply these strategies to COMET and successfully uncover six new defects spanning four types, including two previously unreported categories. Overall, our study identifies 38 unique framework defects, of which 30 are confirmed by developers and 12 have been fixed, demonstrating the practical value of our findings. Yanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen 0005, Peiran Yang, Zhixiang Cao, Ruixiang Qian, Shaoyu Yang 0002, Zhenyu Chen 0001 |
IEEE Trans. Software Eng. | 7 |
| 2025 | ARTransformer: An Architecture of Resolution Representation Learning for Cross-Resolution Person Re-IdentificationabstractABSTRACT Cross‐resolution person re‐identification (CR‐ReID) seeks to overcome the challenge of retrieving and matching specific person images across cameras with varying resolutions. Numerous existing studies utilize established CNNs and ViTs to resize captured low‐resolution (LR) images and align them with high‐resolution (HR) image features or construct common feature spaces to match between images of different resolutions. However, these methods ignore the potential feature connection between the LR and HR images of the same pedestrian identity. Besides, the CNNs or ViTs usually obtain outliers within the attention maps of LR images; this inclination to excessively concentrate on anomalous information may obscure the genuine and anticipated characteristics between images, which makes it challenging to extract meaningful information from the images. In this work, we propose the abnormal feature elimination and reconfiguration Transformer (ARTransformer), a novel network architecture for robust cross‐resolution person re‐identification tasks. This method uses a resolution feature discriminator to learn resolution‐invariant features and output feature matrices of images with different resolutions. It then calculates the potential feature relationships between images of pedestrians with the same identity but different resolutions through a new cross‐resolution landmark agent attention (CR‐LAA) mechanism. Conclusively, it utilizes output feature matrices to model LR and HR image interactions by mitigating abnormal image features and prioritizing attention on the target person by learning representations from input images of various resolutions. Experimental results show that ARTransformer performs well in matching images with different resolutions, even with unseen resolution, and extensive evaluations on four real‐world datasets confirm the excellent results of our approach. Fengshan Lai, Zhixiang Cao, Daoxun Xia |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | DevMuT: Testing Deep Learning Framework via Developer Expertise-Based MutationabstractDeep learning (DL) frameworks are the fundamental infrastructure for various DL applications. Framework defects can profoundly cause disastrous accidents, thus requiring sufficient detection. In previous studies, researchers adopt DL models as test inputs combined with mutation to generate more diverse models. Though these studies demonstrate promising results, most detected defects are considered trivial (i.e., either treated as edge cases or ignored by the developers). To identify important bugs that matter to developers, we propose a novel DL framework testing method DevMuT, which generates models by adopting mutation operators and constraints derived from developer expertise. DevMuT simulates developers' common operations in development and detects more diverse defects within more stages of the DL model lifecycle (e.g., model training and inference). We evaluate the performance of DevMuT on three widely used DL frameworks (i.e., PyTorch, JAX, and Mind-Spore) with 29 DL models from nine types of industry tasks. The experiment results show that DevMuT outperforms state-of-the-art baselines: it can achieve at least 71.68% improvement on average in the diversity of generated models and 28.20% improvement on average in the legal rates of generated models. Moreover, DevMuT detects 117 defects, 63 of which are confirmed, 24 are fixed, and eight are of high value confirmed by developers. Finally, DevMuT has been deployed in the MindSpore community since December 2023. These demonstrate the effectiveness of DevMuT in detecting defects that are close to the real scenes and are of concern to developers. Yanzhou Mu, Juan Zhai, Chunrong Fang, Xiang Chen 0005, Zhixiang Cao, Peiran Yang, Yinglong Zou, Tao Zheng 0005, Zhenyu Chen 0001 |
ASE | 5 |
| 2023 | APICom: Automatic API Completion via Prompt Learning and Adversarial Training-based Data AugmentationabstractBased on developer needs and usage scenarios, API (Application Programming Interface) recommendation is the process of assisting developers in finding the required API among numerous candidate APIs. Previous studies mainly modeled API recommendation as the recommendation task, which can recommend multiple candidate APIs for the given query, and developers may not yet be able to find what they need. Motivated by the neural machine translation research domain, we can model this problem as the generation task, which aims to directly generate the required API for the developer query. After our preliminary investigation, we find the performance of this intuitive approach is not promising. The reason is that there exists an error when generating the prefixes of the API. However, developers may know certain API prefix information during actual development in most cases. Therefore, we model this problem as the automatic completion task and propose a novel approach APICom based on prompt learning, which can generate API related to the query according to the prompts (i.e., API prefix information). Moreover, the effectiveness of APICom highly depends on the quality of the training dataset. In this study, we further design a novel gradient-based adversarial training method ATCom for data augmentation, which can improve the normalized stability when generating adversarial examples. To evaluate the effectiveness of APICom, we consider a corpus of 33k developer queries and corresponding APIs. Compared with the state-of-the-art baselines, our experimental results show that APICom can outperform all baselines by at least 40.02%, 13.20%, and 16.31% in terms of the performance measures EM@1, MRR, and MAP. Finally, our ablation studies confirm the effectiveness of our component setting (such as our designed adversarial training method, our used pre-trained model, and prompt learning) in APICom. Yafeng Gu, Yiheng Shen 0002, Xiang Chen 0005, Shaoyu Yang 0002, Zhixiang Cao |
Internetware | 6 |