Zhihao Chen 0004

dblp:50/505-4 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-1686-9854ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Video Shadow Detection with Intra-and Inter-video Cooperation
Zhihao Chen 0004, Junting Zhao, Lei Zhu 0003, Huazhu Fu, Wei Feng 0005
Int. J. Comput. Vis.2
2025 A Dense Convolutional Bi-Mamba Framework for EEG-Based Emotion Recognition
Mingya Zhang, Yuqian Zhuang, Zhihao Chen 0004, Yiyuan Ge, XianPing Tao
CogSci4
2025 Uncertainty-Aware Medical Diagnostic Phrase Identification and Grounding
abstract
Medical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels.
Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound Devices
abstract
Handheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA.
Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu
IEEE Trans. Medical Imaging3
2025 Topicwise Separable Sentence Retrieval for Medical Report Generation
abstract
Automated radiology reporting holds immense clinical potential in alleviating the burdensome workload of radiologists and mitigating diagnostic bias. Recently, retrieval-based report generation methods have garnered increasing attention. These methods predefine a set of candidate queries and compose reports by searching for sentences in an off-the-shelf sentence gallery that best match these candidate queries. However, due to the long-tail distribution of the training data, these models tend to learn frequently occurring sentences and topics, overlooking the rare topics. Regrettably, in many cases, the descriptions of rare topics often indicate critical findings that should be mentioned in the report. To address this problem, we introduce a Topicwise Separable Sentence Retrieval (Teaser) for medical report generation. To ensure comprehensive learning of both common and rare topics, we categorize queries into common and rare types to learn differentiated topics, and then propose Topic Contrastive Loss to effectively align topics and queries in the latent space. Moreover, we integrate an Abstractor module following the extraction of visual features, which aids the topic decoder in gaining a deeper understanding of the visual observational intent. Experiments on the MIMIC-CXR and IU X-ray datasets demonstrate that Teaser surpasses state-of-the-art models, while also validating its capability to effectively represent rare topics and establish more dependable correspondences between queries and topics. The code is available at https://github.com/CindyZJT/Teaser.git.
Junting Zhao, Yang Zhou 0017, Zhihao Chen 0004, Huazhu Fu
IEEE Trans. Medical Imaging3
2024 Reliable Source Approximation: Source-Free Unsupervised Domain Adaptation for Vestibular Schwannoma MRI Segmentation
Hongye Zeng, Ke Zou, Zhihao Chen 0004, Huazhu Fu
MICCAI (10)3
2024 Learning Physical-Spatio-Temporal Features for Video Shadow Removal
abstract
Shadow removal in a single image has received increasing attention in recent years. However, removing shadows over dynamic scenes remains largely under-explored. In this paper, we propose the first data-driven video shadow removal model, termed PSTNet, by exploiting three essential characteristics of video shadows, i.e., physical property, spatio relation, and temporal coherence. Specifically, a dedicated physical branch was established to conduct local illumination estimation, which is more applicable for scenes with complex lighting and textures, and then enhance the physical features via a mask-guided attention strategy. Then, we develop a progressive aggregation module to enhance the spatio and temporal characteristics of features maps, and effectively integrate the three kinds of features. Furthermore, to tackle the lack of datasets of paired shadow videos, we synthesize a dataset (SVSRD-85) with aid of the popular game GTAV by controlling the switch of the shadow renderer. Experiments against 9 state-of-the-art models, including image shadow removers and image/video restoration methods, show that our method improves the best SOTA in terms of RMSE error for the shadow area by 14.7%. In addition, we develop a lightweight model adaptation strategy to make our synthetic-driven model effective in real world scenes. The visual comparison on the public SBU-TimeLapse dataset verifies the generalization ability of our model in real scenes.
Zhihao Chen 0004, Yefan Xiao, Lei Zhu 0003, Huazhu Fu
IEEE Trans. Circuits Syst. Video Technol.1
2023 Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment
Zhihao Chen 0004, Yang Zhou 0017, Junting Zhao, Gideon Ooi, Lionel Tim-Ee Cheng, Choon Hua Thng, Xinxing Xu, Yong Liu 0026, Huazhu Fu
MICCAI (7)1
2023 Uncertainty-Aware Multi-Dimensional Mutual Learning for Brain and Brain Tumor Segmentation
abstract
Existing segmentation methods for brain MRI data usually leverage 3D CNNs on 3D volumes or employ 2D CNNs on 2D image slices. We discovered that while volume-based approaches well respect spatial relationships across slices, slice-based methods typically excel at capturing fine local features. Furthermore, there is a wealth of complementary information between their segmentation predictions. Inspired by this observation, we develop an Uncertainty-aware Multi-dimensional Mutual learning framework to learn different dimensional networks simultaneously, each of which provides useful soft labels as supervision to the others, thus effectively improving the generalization ability. Specifically, our framework builds upon a 2D-CNN, a 2.5D-CNN, and a 3D-CNN, while an uncertainty gating mechanism is leveraged to facilitate the selection of qualified soft labels, so as to ensure the reliability of shared information. The proposed method is a general framework and can be applied to varying backbones. The experimental results on three datasets demonstrate that our method can significantly enhance the performance of the backbone network by notable margins, achieving a Dice metric improvement of 2.8% on MeniSeg, 1.4% on IBSR, and 1.3% on BraTS2020.
Junting Zhao, Zhaohu Xing, Zhihao Chen 0004, Tong Han, Huazhu Fu, Lei Zhu 0003
IEEE J. Biomed. Health Informatics3
2022 ICBNet: Iterative Context-Boundary Feedback Network for Polyp Segmentation
abstract
Accurate polyp segmentation from colonoscopy images, which is critical to automatic colorectal cancer diagnosis, attracts increasing attentions in recent years. Most existing deep learning-based methods adopt the one-stage processing pipeline, by usually fusing features from different levels or employing boundary-related attention. In this paper, we propose an novel Iterative Context-Boundary feedback Network, namely ICBNet, for robust and accurate polyp segmentation. By mimicking the “from-Preliminary-to-Refined” working paradigm of doctors, ICBNet adopts an iterative feedback learning strategy. Differently from other feedback methods which only use the prediction mask as a guide for foreground features, ICBNet refines encoder features with contextual and boundary-aware details from the preliminary segmentation and boundary predictions, and conducts such strategy in an iterative manner to achieve progressive improvement. Moreover, a dual-branch iterative feedback unit (IFU) is developed to enhance features under the guidance of segmentation and boundary predictions to enable the iterative learning. Extensive experiments on five widely-used polyp segmentation datasets demonstrate that the proposed ICBNet can utilize progressive refinement to effectively address the challenges of large appearance variations and obscure boundaries, and hence achieves more accurate and robust results against the state-of-the-arts methods.
Yefan Xiao, Zhihao Chen 0004, Lequan Yu, Lei Zhu 0003
BIBM2
2022 DSU-Net: Distraction-Sensitive U-Net for 3D lung tumor segmentation
Junting Zhao, Meng Dang, Zhihao Chen 0004
Eng. Appl. Artif. Intell.3
2021 Triple-Cooperative Video Shadow Detection
abstract
Shadow detection in a single image has received significant research interests in recent years. However, much fewer works have been explored in shadow detection over dynamic scenes. The bottleneck is the lack of a well-established dataset with high-quality annotations for video shadow detection. In this work, we collect a new video shadow detection dataset (ViSha), which contains 120 videos with 11,685 frames, covering 60 object categories, varying lengths, and different motion/lighting conditions. All the frames are annotated with a high-quality pixel-level shadow mask. To the best of our knowledge, this is the first learning-oriented dataset for video shadow detection. Furthermore, we develop a new baseline model, named triple-cooperative video shadow detection network (TVSD-Net). It utilizes triple parallel networks in a cooperative manner to learn discriminative representations at intra-video and inter-video levels. Within the network, a dual gated co-attention module is proposed to constrain features from neighboring frames in the same video, while an auxiliary similarity loss is introduced to mine semantic information between different videos. Finally, we conduct a comprehensive study on ViSha, evaluating 12 state-of-the-art models (including single image shadow detectors, video object segmentation, and saliency detection methods). Experiments demonstrate that our model outperforms SOTA competitors.
Zhihao Chen 0004, Lei Zhu 0003, Huazhu Fu, Wennan Liu, Harry Qin
CVPR1
2020 A Multi-Task Mean Teacher for Semi-Supervised Shadow Detection
abstract
Existing shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow detection by leveraging unlabeled data and exploring the learning of multiple information of shadows simultaneously. To be specific, we first build a multi-task baseline model to simultaneously detect shadow regions, shadow edges, and shadow count by leveraging their complementary information and assign this baseline model to the student and teacher network. After that, we encourage the predictions of the three tasks from the student and teacher networks to be consistent for computing a consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from the predictions of the multi-task baseline model. Experimental results on three widely-used benchmark datasets show that our method consistently outperforms all the compared state-of- the-art methods, which verifies that the proposed network can effectively leverage additional unlabeled data to boost the shadow detection performance.
Zhihao Chen 0004, Lei Zhu 0003, Song Wang 0002, Wei Feng 0005, Pheng-Ann Heng
CVPR1
2020 Selective Spatial Regularization by Reinforcement Learned Decision Making for Object Tracking
abstract
Spatial regularization (SR) is known as an effective tool to alleviate the boundary effect of correlation filter (CF), a successful visual object tracking scheme, from which a number of state-of-the-art visual object trackers can be stemmed. Nevertheless, SR highly increases the optimization complexity of CF and its target-driven nature makes spatially-regularized CF trackers may easily lose the occluded targets or the targets surrounded by other similar objects. In this paper, we propose selective spatial regularization (SSR) for CF-tracking scheme. It can achieve not only higher accuracy and robustness, but also higher speed compared with spatially-regularized CF trackers. Specifically, rather than simply relying on foreground information, we extend the objective function of CF tracking scheme to learn the target-context-regularized filters using target-context-driven weight maps. We then formulate the online selection of these weight maps as a decision making problem by a Markov Decision Process (MDP), where the learning of weight map selection is equivalent to policy learning of the MDP that is solved by a reinforcement learning strategy. Moreover, by adding a special state, representing not-updating filters, in the MDP, we can learn when to skip unnecessary or erroneous filter updating, thus accelerating the online tracking. Finally, the proposed SSR is used to equip three popular spatially-regularized CF trackers to significantly boost their tracking accuracy, while achieving much faster online tracking speed. Besides, extensive experiments on five benchmarks validate the effectiveness of SSR.
Qing Guo 0005, Rui-Ze Han, Wei Feng 0005, Zhihao Chen 0004
IEEE Trans. Image Process.4
2018 Background-Suppressed Correlation Filters for Visual Tracking
abstract
Correlation filters (CF) visual object tracking is a powerful framework, with excellent tracking accuracy and beyond real-time frame rate. Its performance, however, can be severely degraded in cluttered background images. In this paper, we propose background-suppressed correlation filters (BSCF), a better CF tracking scheme, which can significantly improve the reliability and accuracy of CF trackers, without harming their beyond real-time speed. Specifically, we present a unified BSCF object function. We show that both the correlation filters and BS weight map can be efficiently and jointly solved in frequency domain. Extensive experiments on OTB-100 benchmark validate the effectiveness and generality of BS in improving multiple CF trackers with higher accuracy and robustness while maintaining their fast tracking speed. We also show BS boosted CF tracker can achieve comparable accuracy of the state-of-the-art spatially-regularized CF tracker but is 14 times faster.
Zhihao Chen 0004, Qing Guo 0005, Wei Feng 0005
ICME1