Hui Yin 0002

dblp:47/3229-2 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-4226-4368ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 14 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RobusTor3D: Robust Multimodal 3D Object Detector for Autonomous Driving by Vision-Language Knowledge Blending
abstract
Multimodal 3D object detection for autonomous driving, a task for real-world applications, poses substantial challenges in maintaining robust performance under various perturbations and complex environmental conditions. However, most existing approaches primarily focus on performance optimization under relatively ideal scenarios or focus on one or few disturbing conditions (or adverse conditions), lacking systematic exploration of robustness against real-world factors, including high class imbalance, adverse weather conditions, sensor jitter and failures, and significant scene variations. To address this issue, we propose a robust multimodal 3D detector, termed RobusTor3D, which integrates robustness at both the structural and supervisory levels by blending the knowledge from Vision-Language Models (VLMs). Structurally, textual descriptions are incorporated to enhance the semantic richness and diversity of rare classes. This novel semantic injection operation compensates for the inherent class imbalance and modality weakness in conventional visual features. Furthermore, semantic alignment capability and robust representation by Vision-Language Knowledge Extraction (V-LKE) serve as semantic priors to complement modality-specific representations, significantly improving model adaptability. At the supervisory level, we propose a Scene-level Multimodal Consistency Learning (SMCL) strategy, which jointly enforces global semantic constraints across modalities, encouraging the learning of stable and abundant semantic representations. This special design specifically reduces the impact of spatial alignment, while notably enabling semantic compensation under modality-loss conditions. Extensive robustness experiments conducted on KITTI, KITTI-C, and CADC benchmarks evaluate five robustness aspects, including long-tail problem, adverse weather (rain, snow, fog, strong sunlight), sensor spatial misalignment and motion blur, modality loss, and cross-domain scenarios. The results show that RobusTor3D demonstrates superior robustness across all five evaluated aspects. It consistently outperforms the state-of-the-art methods under various challenging conditions.
Hui Yin 0002, Ai-Xin Chong, Hui Wang 0001, Zhengyin Liang
AAAI2
2026 Dual-model joint prompting for image dehazing
Hui Yin 0002, Qianqian Du 0002
Knowl. Based Syst.2
2026 MGCFI-Net: Multi-scale globally aware feature learning with cross-view feature interaction for multi-view stereo
Hui Yin 0002, Ai-Xin Chong
Neural Networks2
2026 Time-series forecasting based on fuzzy cognitive maps and GRU-autoencoder
Jiahu Qin, Hui Yin 0002, Yanyan Yang 0001
Soft Comput.5
2026 Shadow vanishing point detection via combined human/shadow adaptive modulation
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Signal Process. Image Commun.3
2026 xDLG: Cross Domain Label-Guided Learning for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic Segmentation
abstract
Unsupervised domain adaptation (UDA) is critical for cross-modal 3D semantic segmentation in real scenarios, as it significantly reduces the annotation costs in new domains. Existing methods primarily focus on various alignment across domains or modalities, however, domain shift and modal differences are inbuilt, and learning under the guidance of source semantic knowledge instead of forced alignment may be more favorable for UDA tasks. Motivated by this insight, we propose cross Domain Label-Guided learning (xDLG) for cross-modal UDA 3D semantic segmentation, which utilizes multimodal pseudo-labels generated online as a bridge for cross-domain label knowledge transfer. Specifically, we propose two types of label-guided learning: Label-Guided Semantic Prototype Learning (LG-SPL) and Label-Guided Pseudo-Label Learning (LG-PLL). First, in LG-SPL, we utilize massive class instances from the source domain to correct the semantic prototypes in the target domain, embedding genuine semantic similarity to tackle the issue of class confusion. Next in LG-PLL, we conceptualize the refinement of pseudo-labels as a reverse diffusion process, leveraging label information from the source domain as a guiding signal to adjust the semantic distribution of pseudo-labels on a global scale. Furthermore, we incorporate domain-specific cues from target 2D images as supplementary conditions throughout the denoising stages, ensuring semantic consistency at finer details. We evaluate our method on 3D semantic segmentation tasks. The results demonstrate that the proposed method achieves state-of-the-art performance across multiple cross-modal UDA scenarios.
Zhengyin Liang, Hui Yin 0002, Qianqian Du 0002
IEEE Trans. Multim.2
2025 UniDxMD: Towards Unified Representation for Cross-Modal Unsupervised Domain Adaptation in 3D Semantic Segmentation
Zhengyin Liang, Hui Yin 0002
ICCV2
2025 Gradual interaction network for stereo matching
Ai-Xin Chong, Hui Yin 0002, Qianqian Du 0002
Pattern Recognit.2
2025 EnIter: Enhancing Iterative Multi-View Depth Estimation with Universal Contextual Hints
abstract
Iterative inference approaches have shown promising success in the task of multi-view depth estimation. However, these methods put excessive emphasis on the universal inter-view correspondences while neglecting the correspondence ambiguity in regions of low texture and depth discontinuous areas. Thus, they are prone to produce inaccurate or even erroneous depth estimations, which is further exacerbated due to cumulative errors especially in the iterative pipeline, providing unreliable information in many real-world scenarios. In this article, we revisit this issue from the intra-view contextual hints and introduce a novel enhancing iterative approach, named EnIter. Concretely, at the beginning of each iteration, we present a Depth Intercept (DI) modulator to provide more accurate depth by aggregating neighbor uncertainty, correlation volume of reference and normal. This plug and play modulator is effective at intercepting the erroneous depth estimations with implicit guidance from the universal correlation contextual hints, especially for the challenging regions. Furthermore, at the end of each iteration, we refine the depth map with another plug and play modulator termed as Depth Refine (DR). It mines the latent structure knowledge of reference contextual hints and establishes one-way dependency using local attention from reference features to depth, yielding delicate depth in detail. Extensive experiment demonstrates that our method not only achieves state-of-the-art performance over existing models but also exhibits remarkable universality in popular iterative pipelines, e.g., CasMVS, UCSNet, TransMVS, and UniMVS.
Qianqian Du 0002, Hui Yin 0002, Lang Nie, Jin Wan
ACM Trans. Multim. Comput. Commun. Appl.2
2024 Enhanced feature pyramid for multi-view stereo with adaptive correlation cost volume
Hui Yin 0002, Ai-Xin Chong, Qianqian Du 0002
Appl. Intell.2
2024 Estimating intrinsic characteristics of images for shadow removal
Hui Yin 0002, Jin Wan, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Comput. Graph.3
2024 CRFormer: A cross-region transformer for shadow removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Image Vis. Comput.2
2024 JoCaD: a joint training method by combining consistency and diversity
Heyan Yang, Hui Yin 0002, Zhengze Yang
Multim. Tools Appl.2
2024 Reference-Based Image Dehazing With Internal and External Contrastive Learning
abstract
Collecting paired pixel-aligned hazy/haze-free image pairs in real-world is arduous for full-supervised image dehazing. Alternatively, methods employing unpaired hazy/clear images have been developed, yet their learning ability about content information of the hazy images is easily disturbed by content-independent clear images, causing artifact problems, particularly for thick hazy images. To address the above issues, we propose a new reference-based image dehazing paradigm with hazy/reference images, where the reference image is clear and taken at the same scene as the hazy image. Therefore, how to maximize the reference value from the hazy/reference images with similar content but unaligned pixels becomes a key issue. Here, we construct a reference-based contrastive learning framework to realize the effective utilization of hazy/reference image pairs. Specifically, internal contrastive learning is designed to preserve the local content invariance between the dehazed images and hazy images in a patch-wise contrastive manner, while the other external contrastive learning learns the global content consistency between the dehazed images and reference images in an overall contrastive manner. Additionally, we design a style consistency loss committee consisting of a regular adversarial loss and a style loss. The former aims to ensure each dehazed image consistent with the overall style distribution of the entire reference set, while the latter is intended to make each dehazed image have an exclusive style with the corresponding reference image. Extensive experiments corroborate that the reference-based dehazing paradigm is recommendable and reliable, and the proposed method performs admirably against other state-of-the-art methods.
Hui Yin 0002, Ai-Xin Chong, Jin Wan
IEEE Trans. Circuits Syst. Video Technol.2
2023 Joint Memory Propagation and Rectification for Video Object Segmentation
Hui Yin 0002, Jin Wan, Jianhuan Chen
ICIG (4)3
2023 SA-Net: Scene-Aware Network for Cross-domain Stereo Matching
Ai-Xin Chong, Hui Yin 0002, Jin Wan, Qianqian Du 0002
Appl. Intell.2
2023 Enhanced spatial-temporal freedom for video frame interpolation
Hao-Dong Li, Hui Yin 0002
Appl. Intell.2
2023 DAGCRN: Graph convolutional recurrent network for traffic forecasting with dynamic adjacency matrix
Zheng Shi 0004, Jiahu Qin, Hui Yin 0002
Expert Syst. Appl.6
2022 Style-Guided Shadow Removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
ECCV (19)2
2022 Multi-hierarchy feature extraction and multi-step cost aggregation for stereo matching
Ai-Xin Chong, Hui Yin 0002, Jin Wan
Neurocomputing2
2022 Edge Aware Network for Image Dehazing
abstract
The learning-based methods have recently shown their advantages in the image dehazing task. However, most existing learning-based methods do not pay much attention to the restoration in the edges of the hazy image, resulting in the edge blur of the dehazing results. To mitigate this issue, in this letter, we propose a novel Edge Aware Network (EA-Net) for image dehazing, which can simultaneously model edge features and contextual features into a single network for restoring haze-free image with sharp edges. Firstly, we extract the multi-scales contextual features of hazy image by a progressive fusion way. Furthmore, the abundant edge features are inferred by a edge subnetwork with Edge Feature Extraction Module(EFEM). Finally, we present an Edge Attention (EA) mechanism to couple the edge features with contextual features at various resolutions for sufficiently leveraging these complementary features. Due to the rich edge information and feature fusion strategy, the fused features can make the haze-free image to be clearer, especially at the edges, which is very important for the high-level vision tasks. Extensive experiments demonstrate that the proposed method achieves significant improvements over the state-of-the-art methods.
Hui Yin 0002, Jin Wan, Ai-Xin Chong
IEEE Signal Process. Lett.2
2021 From Shadow Generation To Shadow Removal
abstract
Shadow removal is a computer-vision task that aims to restore the image content in shadow regions. While almost all recent shadow-removal methods require shadow-free images for training, in ECCV 2020 Le and Samaras introduces an innovative approach without this requirement by cropping patches with and without shadows from shadow images as training samples. However, it is still laborious and time-consuming to construct a large amount of such unpaired patches. In this paper, we propose a new G2R-ShadowNet which leverages shadow generation for weakly-supervised shadow removal by only using a set of shadow images and their corresponding shadow masks for training. The proposed G2R-ShadowNet consists of three sub-networks for shadow generation, shadow removal and refinement, respectively and they are jointly trained in an end-to-end fashion. In particular, the shadow generation sub-net stylises non-shadow regions to be shadow ones, leading to paired data for training the shadow-removal sub-net. Extensive experiments on the ISTD dataset and the Video Shadow Removal dataset show that the proposed G2R-ShadowNet achieves competitive performances against the current state of the arts and outperforms Le and Samaras’ patch-based shadow-removal method.
Hui Yin 0002, Xinyi Wu 0002, Zhenyao Wu, Yang Mi, Song Wang 0002
CVPR2
2021 Pop-net: A self-growth network for popping out the salient object in videos
abstract
Abstract It is a big challenge for unsupervised video segmentation without any object annotation or prior knowledge. In this article, we formulate a completely unsupervised video object segmentation network which can pop out the most salient object in an input video by self‐growth, called Pop‐Net. Specifically, in this article, a novel self‐growth strategy which helps a base segmentation network to gradually grow to stick out the salient object as the video goes on, is introduced. To solve the sample generation problem for the unsupervised method, the sample generation module which fuses the appearance and motion saliency is proposed. Furthermore, the proposed sample optimization module improves the samples by using contour constrains for each self‐growth step. Experimental results on several datasets (DAVIS, DAVSOD, VideoSD, Segtrack‐v2) show the effectiveness of the proposed method. In particular, the state‐of‐the‐art methods on completely unfamiliar datasets (no fine‐tuned datasets) are performed.
Hui Yin 0002, Jin Wan
IET Comput. Vis.1
2021 ADSCN: Adaptive dense skip connection network for railway infrastructure displacement monitoring images super-resolution
Hui Yin 0002, Jin Wan, Shi-Jie Zhang
Multim. Tools Appl.1
2021 Shadow Removal by a Lightness-Guided Network With Training on Unpaired Data
abstract
Shadow removal can significantly improve the image visual quality and has many applications in computer vision. Deep learning methods based on CNNs have become the most effective approach for shadow removal by training on either paired data, where both the shadow and underlying shadow-free versions of an image are known, or unpaired data, where shadow and shadow-free training images are totally different with no correspondence. In practice, CNN training on unpaired data is more preferred given the easiness of training data collection. In this paper, we present a new Lightness-Guided Shadow Removal Network (LG-ShadowNet) for shadow removal by training on unpaired data. In this method, we first train a CNN module to compensate for the lightness and then train a second CNN module with the guidance of lightness information from the first CNN module for final shadow removal. We also introduce a loss function to further utilise the colour prior of existing data. Extensive experiments on widely used ISTD, adjusted ISTD and USR datasets demonstrate that the proposed method outperforms the state-of-the-art methods with training on unpaired data.
Hui Yin 0002, Yang Mi, Mengyang Pu, Song Wang 0002
IEEE Trans. Image Process.2
2020 Progressive residual networks for image super-resolution
Jin Wan, Hui Yin 0002, Ai-Xin Chong
Appl. Intell.2
2019 Adaptive convolutional neural network for large change in video object segmentation
abstract
This study tackles the semi‐supervised segmentation task for the objects that have large motion or appearance change in a video sequence, which is very challenging to the existing methods of video object segmentation (VOS). In this study, a novel adaptive approach is presented, named adaptive convolutional neural network for large change VOS, which determines when and how to fine‐tune the convolutional neural network through the motion metric and the appearance metric among consecutive video frames. Additionally, a lightweight optimisation algorithm for the predictive binary mask is introduced which is effective for pixel prediction by eliminating the discrete points cluster. To illustrate the advantages of this approach, experiments have been performed on four VOS datasets, which demonstrate that the proposed method is highly effective and could achieve the state‐of‐the‐art on these datasets.
Hui Yin 0002, Jin Wan
IET Comput. Vis.1
2008 A Global Contour-Grouping Algorithm Based on Spectral Clustering
Hui Yin 0002, Siwei Luo
ISNN (2)1