Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Yifei Han

dblp:262/9610 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Segmentation and scene understanding · 55% Vision and language · 34% Efficient and distributed learning · 7%
Computer graphics and multimedia
1 paper
Rendering · 50% Geometric modeling and processing · 50%

Topics — the 15 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Segmentation and scene understanding
medical image segmentation
1.012026
Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation · AAAI 2026
Computer vision › Segmentation and scene understanding › semantic segmentation
transformer-based segmentation
1.012026
Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation · AAAI 2026
Computer vision › Segmentation and scene understanding › medical image segmentation
vessel segmentation
1.012026
Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation · AAAI 2026
Computer vision › Vision and language › vision-language model
CLIP
0.912025
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation · IEEE Trans. Image Process. 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation
0.912025
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation · IEEE Trans. Image Process. 2025
Computer vision › Vision and language
vision-language model
0.912025
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation · IEEE Trans. Image Process. 2025
Rendering › gaussian splatting
3d gaussian splatting
0.912025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Geometric modeling and processing › 3d reconstruction
3d scene reconstruction
0.912025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Geometric modeling and processing › 3d reconstruction › 3d scene reconstruction
large-scale scene reconstruction
0.912025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Rendering
neural rendering
0.912025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Computer vision › Vision and language
image captioning
0.712023
Infrared Image Captioning with Wearable Device · ICRA 2023
Medical and health informatics
cardiac image analysis
0.312026
Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.312025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation
0.312025
Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction · ICCV 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.312025
Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation · IEEE Trans. Image Process. 2025

Methods — techniques the papers use, named apart from their topics

semantic clustering attention · 2.0adaptive morph-patch partitioning · 2.0momentum-based self-distillation · 1.7block-wise parallel training · 1.7block weighting · 1.7training-free adaptation · 0.9multi-level feature fusion · 0.9attention calibration · 0.9deep learning · 0.7
YearPublicationVenuePosition
2026 Adaptive Morph-Patch Transformer for Aortic Vessel Segmentation
abstract
Accurate segmentation of aortic vascular structures is critical for diagnosing and treating cardiovascular diseases. Traditional Transformer-based models have shown promise in this domain by capturing long-range dependencies between vascular features. However, their reliance on fixed-size rectangular patches often influences the integrity of complex vascular structures, leading to suboptimal segmentation accuracy. To address this challenge, we propose the adaptive Morph-Patch Transformer (MPT), a novel architecture specifically designed for aortic vascular segmentation. Specifically, MPT introduces an adaptive patch partitioning strategy that dynamically generates morphology-aware patches aligned with complex vascular structures. This strategy can preserve semantic integrity of complex vascular structures within individual patches. Moreover, a Semantic Clustering Attention (SCA) method is proposed to dynamically aggregate features from various patches with similar semantic characteristics. This method enhances the model's capability to segment vessels of varying sizes, preserving the integrity of vascular structures. Extensive experiments on three open-source datasets (AVT, AortaSeg24 and TBAD) demonstrate that MPT achieves state-of-the-art performance, with improvements in segmenting intricate vascular structures.
Fuchen Zheng, Adnan Iltaf, Yifei Han, Zhenyu Chen 0001, Yue Du, Bin Li 0083, Tianyong Liu, Shoujun Zhou
AAAI4
2026 A novel clustering algorithm for categorical data with MGR based reference set selection method
Keqi Cheng, Xiuqin Ma, Hongwu Qin, Yifei Han
Neurocomputing5
2026 Real-Time Multi-Modal Social Event Detection: A New Dataset and a Key Instance-Driven, Quality-Aware Graph Neural Network
abstract
Social event detection (SED) involves identifying and analyzing significant real-world events using data generated on social media platforms. With the rapid growth of platforms like Weibo and Twitter, users are sharing not just text but also images. However, most existing SED methods remain text-focused, limiting their ability to fully capture the complexity of real-world social dynamics. Moreover, the lack of multi-modal datasets specifically designed for SED has blocked the development of models that can effectively exploit these rich content types. To address these limitations, we introduced WEIBO2022, an extensive multi-modal SED dataset that includes both text and image data. The dataset is available in two versions: WEIBO2022-Medium, containing 25,435 entries and WEIBO2022-Large, containing 79,825 entries. In addition, we presented a novel network called the Key Instance-driven, Quality-aware Graph Neural Network (KQGNN), which features a key instance-driven library, a quality-aware learning process, and a multi-modal fusion module, enhancing its ability to detect events accurately in both offline and real-time settings. Extensive experiments showcase the exceptional performance and superiority of the proposed model, showing improvements in detection accuracy and effective prevention of catastrophic forgetting during continuous training.
Yifei Han, Congbo Ma, Shan Xue 0001, Jian Yang 0001, Jia Wu 0001
IEEE Trans. Big Data1
2025 Momentum-Gs: Momentum Gaussian Self-Distillation for High-Quality Large Scene Reconstruction
abstract
3D Gaussian Splatting has demonstrated notable success in large-scale scene reconstruction, but challenges persist due to high training memory consumption and storage overhead. Hybrid representations that integrate implicit and explicit features offer a way to mitigate these limitations. However, when applied in parallelized block-wise training, two critical issues arise since reconstruction accuracy deteriorates due to reduced data diversity when training each block independently, and parallel training restricts the number of divided blocks to the available number of GPUs. To address these issues, we propose Momentum-GS, a novel approach that leverages momentum-based self-distillation to promote consistency and accuracy across the blocks while decoupling the number of blocks from the physical GPU count. Our method maintains a teacher Gaussian decoder updated with momentum, ensuring a stable reference during training. This teacher provides each block with global guidance in a self-distillation manner, promoting spatial consistency in reconstruction. To further ensure consistency across the blocks, we incorporate block weighting, dynamically adjusting each block's weight according to its reconstruction accuracy. Extensive experiments on large-scale scenes show that our method consistently outperforms existing techniques, achieving a 12.8% improvement in LPIPS over CityGaussian with much fewer divided blocks and establishing a new state of the art. Project page: https://jixuan-fan.github.io/Momentum-GS_Page/
Jixuan Fan, Wanhua Li 0001, Yifei Han, Tianru Dai, Yansong Tang
ICCV3
2025 GS-3Det: Elevating 3D Gaussian Splatting for Real-Time Multi-View 3D Object Detection
abstract
3D Gaussian Splatting (3DGS) has emerged as a high-quality and efficient alternative to Neural Radiance Fields (NeRF), offering distinct advantages in scene representation and suitability for multi-view 3D object detection (MV-3DOD) tasks. However, conventional 3DGS-based detection methods typically employ a reconstruction-then-detection pipeline, which is time-consuming and unsuitable for real-time applications. This approach arises from 3DGS’s original design for reconstruction tasks, which lacks network-based training. In this paper, we propose GS-3Det, an online Gaussian detection framework for MV-3DOD, achieving real-time performance through a single forward pass for scene reconstruction and detection. Specifically, we introduce a detection-aware Gaussian grid that enables directly prediction of explicit 3D Gaussians from multi-view images, enabling efficient and robust 3D scene understanding. Additionally, we propose a Dual-Path Consistency module that leverages 3D constraints to improve the accuracy of Gaussian grid representation and detection. Experiments on the ScanNet V2 dataset demonstrate that GS-3Det surpasses state-of-the-art methods by 3.5% in [email protected] and 2.7% in [email protected], underscoring its generalization capability and real-time performance.
Yifei Han, Jixuan Fan, Sule Bai, Yuji Wang, Yansong Tang
VCIP1
2025 Self-Calibrated CLIP for Training-Free Open-Vocabulary Segmentation
abstract
Recent advancements in pre-trained vision-language models like CLIP, have enabled the task of open-vocabulary segmentation. CLIP demonstrates impressive zero-shot capabilities in various downstream tasks that require holistic image understanding. However, due to the image-level contrastive learning and fully global feature interaction, ViT-based CLIP struggles to capture local details, resulting in poor performance in segmentation tasks. Our analysis of ViT-based CLIP reveals that anomaly tokens emerge during the forward process, attracting disproportionate attention from normal patch tokens and thereby diminishing spatial awareness. To address this issue, we propose Self-Calibrated CLIP (SC-CLIP), a training-free method that calibrates CLIP to generate finer representations while preserving its original generalization ability-without introducing new parameters or relying on additional backbones. Specifically, we mitigate the negative impact of anomaly tokens from two complementary perspectives. First, we explicitly identify the anomaly tokens and replace them based on local context. Second, we reduce their influence on normal tokens by enhancing feature discriminability and attention correlation, leveraging the inherent semantic consistency within CLIP's mid-level features. In addition, we introduce a two-pass strategy that effectively integrates multi-level features to enrich local details under the training-free setting. Together, these strategies enhance CLIP's feature representations with improved granularity and semantic coherence. Experimental results demonstrate the effectiveness of SC-CLIP, achieving state-of-the-art results across all datasets and surpassing previous methods by 9.5%. Notably, SC-CLIP boosts the performance of vanilla CLIP ViT-L/14 by 6.8 times. Furthermore, we discuss our method's applicability to other vision-language models and tasks for a comprehensive evaluation. Our source code is available at https://github.com/SuleBai/SC-CLIP.
Sule Bai, Yong Liu 0033, Yifei Han, Haoji Zhang 0001, Yansong Tang, Jie Zhou 0001, Jiwen Lu
IEEE Trans. Image Process.3
2024 RDNRnet: A Reconstruction Solution of NDVI Based on SAR and Optical Images by Residual-in-Residual Dense Blocks
abstract
The reconstruction of the Normalized Difference Vegetation Index (NDVI) is a crucial prerequisite for numerous spatiotemporal continuous studies. To address the limitations posed by satellite temporal resolution and challenging atmospheric conditions, the combination of Synthetic Aperture Radar (SAR) and optical images from diverse sources has proven to be effective and widely employed. In this study, we employ the Spatial-Temporal Savitzky-Golay algorithm to rectify MODIS NDVI maps and eliminate interruptions caused by noise. Random Forest and Gradient Boosting Decision Trees serve as a dual filter to select SAR indices with the highest impact on NDVI reconstruction, ensuring that the chosen indices encapsulate the most valuable information. Subsequently, we conduct a series of ablation experiments and develop a deep learning network named Residual-in-Residual Dense Block NDVI Reconstruction net (RDNRnet). This network effectively mitigates the impacts of MODIS coarse resolution and speckle noises in SAR data. We also evaluate the network performance in reconstructing NDVI across all seasons and land cover types. Our findings highlight that the modified dual-polarimetric SAR vegetation index and the standard deviation of the VV band are the most crucial SAR indices. The predictions for summer exhibit the highest performance, with a coefficient of determination (R2) reaching 0.9757. Optimal performances by land cover type are observed in forests, paddy fields, and dry farming fields, all with R2values exceeding 0.9580. Our adaptive NDVI reconstruction solution demonstrates robust performance across different data availability scenarios, effectively catering to all seasons and land cover types.
Yifei Han, Jinliang Huang, Feng Ling 0003, Hong Chi
IEEE Trans. Geosci. Remote. Sens.1
2023 Infrared Image Captioning with Wearable Device
abstract
Wearable devices have garnered widespread attention as a mobile solution, and various intelligent modules based on wearable devices are increasingly being integrated. Additionally, image captioning is an important task in computer vision that maps images to text. Existing image captioning achievements are based on high-quality visible images. However, higher target complexity and insufficient light can lead to reduced captioning performance and mistakes. In this paper, we present an infrared image captioning framework designed to solve the problem of invalid visible image captioning in special conditions. Remarkably, we integrate the infrared image captioning model into the wearable device. Volunteers perform offline and real-time environmental analysis tasks in the real world to evaluate the framework's effectiveness in multiple scenarios. The results indicate that both the accuracy of infrared image captioning and the feedback from wearable device users are promising.
Chenjun Gao, Yanzhi Dong, Xiaohu Yuan, Yifei Han
ICRA4