VLDB 2026 Research / reviewers in the wild / expert
Yihong Cao
dblp:305/7406
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0003-1751-5505ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAF: A Structure-Aware Framework for Radial Ice Thickness Detection on Overhead Transmission LinesabstractIce thickness estimation on overhead transmission lines (OHTL) is essential for mitigating icing-induced mechanical failures and ensuring safe grid operation. To address the challenges of detecting radial ice thickness in complex power line corridors, particularly geometric fragmentation of slender conductors and semantic ambiguity near occluded boundaries, this work proposes a structure-aware framework (SAF) based on 3-D point cloud segmentation and geometry-guided modeling. SAF introduces a structure-aware segmentation network, which integrates a cross-level spatial encoding module to preserve geometric continuity and a partition-aware loss to improve boundary localization under vegetation or tower occlusion. Building on accurate segmentation, a geometry-guided module performs centerline fitting and cross-sectional reconstruction to infer slice-level ice thickness. To support evaluation, a large-scale uncrewed aerial vehicle (UAV)-based point cloud dataset covering 32 OHTL is constructed, including six lines with ground-truth ice labels. Experimental results demonstrate that SAF achieves robust and accurate ice estimation across varied voltage levels and terrains, supporting its practical application in intelligent transmission line inspection and icing risk prevention. Hui Zhang 0023, Youyuan Tang, Yihong Cao, Kaining Zhang, Yunkang Cao, Tongzhi Niu, Jianxu Mao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | NRSeg: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World ModelsabstractBirds' Eye View (BEV) semantic segmentation is an indispensable perception task in end-to-end autonomous driving systems. Unsupervised and semi-supervised learning for BEV tasks, as pivotal for real-world applications, underperform due to the homogeneous distribution of the labeled data. In this work, we explore the potential of synthetic data from driving world models to enhance the diversity of labeled data for robustifying BEV segmentation. Yet, our preliminary findings reveal that generation noise in synthetic data compromises efficient BEV model learning. To fully harness the potential of synthetic data from world models, this article proposes NRSeg, a noise-resilient learning framework for BEV semantic segmentation. Specifically, a Perspective-Geometry Consistency Metric (PGCM) is proposed to quantitatively evaluate the guidance capability of generated data for model learning. This metric originates from the alignment measure between the perspective road mask of generated data and the mask projected from the BEV labels. Moreover, a Bi-Distribution Parallel Prediction (BiDPP) is designed to enhance the inherent robustness of the model, where the learning process is constrained through parallel prediction of multinomial and Dirichlet distributions. The former efficiently predicts semantic probabilities, whereas the latter adopts evidential deep learning to realize uncertainty quantification. Furthermore, a Hierarchical Local Semantic Exclusion (HLSE) module is designed to address the non-mutual exclusivity inherent in BEV semantic segmentation tasks. The proposed framework is evaluated on BEV semantic segmentation using data generated by multiple world models, with comprehensive testing conducted on the public nuScenes dataset under unsupervised and semi-supervised settings. Experimental results demonstrate that NRSeg achieves state-of-the-art performance, yielding the highest improvements in mIoU of 13.8% and 11.4% in unsupervised and semi-supervised BEV segmentation tasks, respectively. The source code will be made publicly available at https://github.com/lynn-yu/NRSeg. Siyu Li 0002, Yihong Cao, Kailun Yang 0001, Zhiyong Li 0001, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation
Yihong Cao, Jiaming Zhang 0001, Xu Zheng 0002, Hao Shi 0004, Kunyu Peng, Kailun Yang 0001, Hui Zhang 0023 |
ICCV | 1 |
| 2025 | VSLNet: Multimodal Data Fusion Network for Tree Species Classification in Overhead Transmission Line CorridorsabstractThe classification of tree species for overhead transmission lines (OHTL) is of great significance, facilitating the the timely removal of safety hazards posed by trees on power lines. Addressing the challenges in classifying OHTL line tree species, including subtle differences in target shape appearance, densely distributed targets, and limited representation in single-modal data, this article proposes a tree species classification network, VSLNet, based on multimodal data fusion. VSLNet constructs three asymmetric branches, which automatically select more discriminative features among spectra during spectral information processing, and jointly guide the extracted visible light information, ensuring global and local consistency for accurate multispectral classification. Furthermore, in LiDAR processing, the segmentation of individual trees contributes data such as tree height and crown diameter, and seamlessly integrates GPS data with multispectral classification results. Experimental results demonstrate that VSLNet is a feasible and reliable solution for tree classification, with potential applicability to other multimodal tasks. Hui Zhang 0023, Hang Zhong, Yihong Cao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Multimodal Fusion Network for Power Tower Semantic Segmentation and Inclination DetectionabstractAs the most fundamental supporting infrastructure of the distribution network, power towers require regular checks of their tilting status to ensure the system’s smooth operation. To overcome distinguishing objects in large-scale scenes using solely images or point clouds poses significant difficulties, we propose a multimodal fusion semantic segmentation network (MFSS) for power tower semantic segmentation and inclination detection. First, effective near-ground filtering and fixed-area slicing algorithms are proposed to address the issues of sample imbalance and insufficient data. Second, MFSS integrates RGB information and point cloud features to enhance the descriptive ability of the tower. Finally, a novel inclination detection method for distribution towers is proposed, estimating tower tilt from the axis between top and bottom centroids to improve accuracy and stability. Experimental results on our constructed dataset show that the proposed method outperforms existing algorithms, achieving 77.3% IoU and 96.69% per-class accuracy in tower segmentation. The mean angle deviation for tilt detection is 0.78$^{\circ }$, with a state judgment false rate of just 2.4%. Hui Zhang 0023, Hang Zhong, Youyuan Tang, Yihong Cao, Yaonan Wang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | Reliable Wind Turbine Blade Performance Monitoring System Using Aerodynamic Audio Signals and Deep Learning ApproachesabstractWind turbines have emerged as a prominent and environmentally friendly energy generation solution. However, with the widespread use of new materials, ensuring the reliability of these devices has become as a critical issue. Developing efficient and cost-effective monitoring methods for the wind turbine's blades (WTBs), the most expensive components of wind turbine, has become a focal point of research. In this article, we present a novel monitoring system for WTBs that employs a deep convolutional neural network approach based on the medical auscultatory method. The system is designed to balance economic efficiency and engineering reliability. First, we proposed a lightweight WTBs monitoring framework based on edge computing that leverages the signals from the programmable logic controller output of wind turbine to enable efficient collection of relevant aerodynamic audio signals while filtering out irrelevant data. Second, we present a set of audio enhancement algorithms that employ multiscale feature extraction, self-adaptive mask targeting, and deep neural networks to reduce noise in the audio signals generated by WTBs. Third, we introduce a new approach for compressing deep convolution neural networks that makes them suitable for resource-constrained edge computing devices and efficiently utilizes audio-generated spectrograms to diagnose faults in WTBs. Baheti Biekezat, Hui Zhang 0023, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Reliab. | 3 |
| 2024 | Occlusion-Aware Seamless Segmentation
Yihong Cao, Jiaming Zhang 0001, Hao Shi 0004, Kunyu Peng, Yuhongxuan Zhang, Hui Zhang 0023, Rainer Stiefelhagen, Kailun Yang 0001 |
ECCV (19) | 1 |
| 2024 | Deep Correspondence Matching-Based Robust Point Cloud Registration of Profiled PartsabstractDue to ability to estimate the spatial transformation of coordinate frames, point cloud registration is a fundamental technique in manufacturing. Previous methods prone to converge to wrong local minima, in the cases of large initialization, noise, outliers, and partiality. This article presents a new learning-based robust point cloud registration approach to predict a rigid transformation in a one-shot way. Our network aims to determine a matchability matrix to yield an accurate registration result. Each element of the matchability matrix refers to similarity of learned per-point embeddings and represents the probability of a potential correspondence. The following two major blocks are developed to guide the matchability matrix to represent correct correspondences: an attention block is introduced to enhance the discriminativeness of learned per-point embeddings, and a zero-mean Gaussian-based annealing layer and a differentiable Sinkhorn normalization layer are designed to enforce a permutation matchability matrix. With the matchability matrix, an intuitive solution is integrated to obtain the relative transformation of the source and target point clouds. Different from the existing work, our network can handle partially overlapped point-cloud pairs effectively. Experimental results demonstrate the superiority of the proposed approach over the state-of-the-art registration approaches in terms of accuracy and robustness. Weixing Peng, Yaonan Wang 0001, Hui Zhang 0023, Yihong Cao, Jiawen Zhao, Yiming Jiang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Adaptive Refining-Aggregation-Separation Framework for Unsupervised Domain Adaptation Semantic SegmentationabstractUnsupervised domain adaptation has attracted widespread attention as a promising method to solve the labeling difficulties of semantic segmentation tasks. It trains a segmentation network for unlabeled real target images using easily available labeled virtual source images. To improve performance, clustering is used to obtain domain-invariant feature representations. However, most clustering-based methods indiscriminately cluster all features mapped by category from both domains, causing the centroid shift and affecting the generation of discriminative features. We propose a novel clustering-based method that uses an adaptive refining-aggregation-separation framework, which learns the discriminative features by designing different adaptive schemes for different domains and features. The clustering does not require any tunable thresholds. To estimate more accurate domain-invariant centroids, we design different ways to guide the adaptive refinement of different domain features. A critic is proposed to directly evaluate the confidence of target features to solve the absence of target labels. We introduce a domain-balanced aggregation loss and two adaptive separation losses for distance and similarity respectively, which can discriminate clustering features by combining the refinement strategy to improve segmentation performance. Experimental results on GTA$5\rightarrow $Cityscapes and SYNTHIA$\rightarrow $Cityscapes benchmarks show that our method outperforms existing state-of-the-art methods. Yihong Cao, Hui Zhang 0023, Xiao Lu 0002, Yurong Chen 0003, Yaonan Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Cycle Consistency Based Pseudo Label and Fine Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to transfer knowledge from a well-labeled source domain to an unlabeled target domain with a correlative distribution. Numerous existing approaches process this hard nut by directly matching the marginal distribution between two domains, which confront the obstacle of rough alignment and blurred decision boundary. Recent advances in UDA introduce target pseudo-label and subdomain adaptation to reduce misalignment and distribution discrepancy. Whereas, they frequently ignore that the production of target pseudo-label is so dependent on the source-trained classifier, which without reasonable restriction to discriminate generated pseudo-label is whether confident. Meanwhile, many methods in the subdomain alignment metric ignore exploring the potential distribution discrepancy between same-class samples of the intra-domain. To address these two issues simultaneously, this paper proposes a Cycle Consistency based Pseudo Label and Fine Alignment (CCPLFA) approach for UDA. In particular, firstly, a novel cycle-consistency based pseudo label module is designed, which is a simple yet effective way to alleviate the noise of pseudo labels and improve their semantic correctness. Secondly, we develop a Fine-Alignment distribution matching metric. Which can maximize the feature distribution density of intra-class cross-domains and not overlook the distribution structure of the global aspect. Comprehensive experiment results on four benchmarks demonstrate the capability of plug and play and the well generalization performance of our proposed method. Hui Zhang 0023, Junkun Tang, Yihong Cao, Yurong Chen 0003, Yaonan Wang 0001, Q. M. Jonathan Wu |
IEEE Trans. Multim. | 3 |
| 2022 | Video Shadow Detection via Spatio-Temporal Interpolation Consistency TrainingabstractIt is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results. In this paper, we address these challenges by proposing a Spatio-Temporal Interpolation Consistency Training (STICT) framework to rationally feed the unlabeled video frames together with the labeled images into an image shadow detection network training. Specifically, we propose the Spatial and Temporal ICT, in which we define two new interpolation schemes, i.e., the spatial interpolation and the temporal interpolation. We then derive the spatial and temporal interpolation consistency constraints accordingly for enhancing generalization in the pixel-wise classification task and for encouraging temporal consistent predictions, respectively. In addition, we design a Scale- Aware Network for multi-scale shadow knowledge learning in images, and propose a scale-consistency constraint to minimize the discrepancy among the predictions at different scales. Our proposed approach is extensively validated on the ViSha dataset and a self-annotated dataset. Experimental results show that, even without video labels, our approach is better than most state of the art supervised, semi-supervised or unsupervised image/video shadow detection methods and other methods in related tasks. Code and dataset are available at https://github.com/yihong-97/STICT. Xiao Lu 0002, Yihong Cao, Chengjiang Long, Zipei Chen, Xuanyu Zhou, Yimin Yang 0001, Chunxia Xiao |
CVPR | 2 |
| 2021 | Real-time stage-wise object tracking in traffic scenes: an online tracker selection method via deep reinforcement learning
Xiao Lu 0002, Yihong Cao, Xuanyu Zhou, Yimin Yang 0001 |
Neural Comput. Appl. | 2 |