Zhongyi Han

dblp:181/7439 · DBLP profile ↗
← Back
49ranked-venue papers
13as first author
41since 2021 · last 2026
0000-0003-2851-193XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Retriever Encoder Selection Matters for In-Context Learning-based Medical Segmentation
abstract
In-context learning-based medical segmentation (ICLM) enables foundation models to generalize to unseen cases without retraining. To enhance performance on test queries, existing methods typically follow a two-stage process: (1) using a retrieval encoder (RE) to map both queries and training samples into a shared feature space, and (2) retrieving and utilizing the top-k most similar training samples. While current methods fix the RE and focus on optimizing stage (2), we show that the choice of RE in stage (1) alone can account for over 70% of the performance variation, highlighting RE selection as a critical yet often overlooked factor in ICLM. In this paper, we conduct an analysis of the RE selection and make two main findings: (1) dynamically selecting the RE for each query outperforms selecting a fixed RE for the entire task; and (2) feature-space heuristics (e.g., intra-class compactness and inter-class separability) fail to predict RE quality. To this end, we propose the instance-adaptive retrieval encoder selection (IRES) method that can select the optimal RE for each query based on output predictions. IRES is based on the intuition that a good RE retrieves relevant demonstrations, helping the ICL model generate more accurate and stable segmentation masks. Thus, we introduce the shape stability score (S³), which evaluates the morphological stability of predicted masks under iterative erosion. Experiments show S³ correlates strongly with true RE quality (Pearson > 0.8), serving as a reliable selection proxy. To reduce S³’s per-query cost, we propose parallel prediction with reciprocal neighbor reuse (P2R), which accelerates inference by parallelizing encoding and reusing encoder selections across reciprocal neighbors, avoiding redundant computation. Built on S³ and P2R, IRES improves ICLM performance across FUNDUS, Brain MRI, and Chest X-ray datasets, with up to 10.6% gain on fundus segmentation.
Zhongyi Han, Yongshun Gong, Yilong Yin
AAAI2
2026 From Small to Large: In-Context Learning as a New Paradigm for Domain Generalization
Guanglin Zhou, Zhongyi Han, Shaoan Xie, Shiming Chen 0002, Biwei Huang, Liming Zhu 0001, Xinbo Gao 0001, Lina Yao 0001, Salman Khan 0001
Int. J. Comput. Vis.2
2026 SA-Diff: Semantic-Aware graph outlier generation via diffusion models for graph out-of-Distribution detection
Yicong Dong, Rundong He, Zhongyi Han, Jieming Shi 0001, Yilong Yin
Knowl. Based Syst.3
2026 Adaptive feature unlearning for trustworthy medical imaging privacy
Zhongyi Han, Bin Wang 0045, Shenjing Wu, Juexiao Zhou, Gongning Luo, Benzheng Wei, Xin Gao 0001
Medical Image Anal.1
2026 Class-mismatched semi-supervised learning from a new perspective
Rundong He, Zhongyi Han, Xiushan Nie, Qi Wei 0004, Yilong Yin
Pattern Recognit.2
2026 From Contrast-Driven Segmentation to Central Lumbar Spinal Stenosis Grading: A Comprehensive Multi-View Spinal MRI Image Analysis
abstract
Central lumbar spinal stenosis, a prevalent degenerative spinal disorder, severely impacts the quality of life for those affected. Axial and sagittal MRI images offer diverse information on tissue structure and lesions, which is crucial for accurate diagnosis. However, MRI-based diagnostic approaches still have poor lesion localization, insufficient cross-view alignment, underutilization of multi-view MRI information, and limited generalization across patient variability. To address these problems, we proposed an Encompassing Lumbar Central Spinal Stenosis Grading Model via Multi-view MRI Image Fusion called ELSG-MF. ELSG-MF consists of three stages: the first stage utilizes the extraction of robust pseudo-labels through a contrast-driven consistency reinforcement technique to guide Med-SAM in localizing and segmenting spinal tissue components. The Sagittal-Axial Pairing (SAP) Algorithm was developed by stage2 to integrate the spatial anatomical relationship between the vertebral body and the intervertebral disc, facilitating the correlation pairing between sagittal and axial images. Stage3 subsequently innovated the multi-view Adaptive Fusion (M²AF) module, which enables adaptive dynamic fusion of anatomical features across views. M²AF enhances the extraction of contextual complementary information, and significantly improves the model’s capacity to detect subtle variations in the degree of narrowness. A series of studies show that our model achieves an overall accuracy of 0.8631, AUC of 0.96, and F1-score of 0.8614. These results indicate that our model substantially outperforms mainstream approaches, attaining superior segmentation and grading accuracy, exhibiting robust generalization and clinical application potential.
Zhengchao Zhou, Xinggui Ji, Wanbo Xu, Zhongyi Han, Benzheng Wei
IEEE Trans. Medical Imaging5
2025 Empowering Multimodal Models via Active in-Context Learning for Test-Time Medical Imaging
abstract
In-context learning (ICL) has enabled large multimodal models (LMMs) to achieve effective medical image classification through the strategic utilization of relevant examples from pre-existing training sets. However, building high-quality training sets for ICL in medical imaging remains challenging due to costly manual annotations and strict privacy constraints. In this paper, we propose Active In-Context Learning (AICL), a novel paradigm that eliminates the need for pre-existing training sets. AICL dynamically selects and annotates a small, informative set of medical samples at dynamic test time during the query phase, by continuously retrieving relevant ICL examples to optimize LMM performance without relying on traditional datasets. To construct an optimal active set, we introduce Neighbor-relaxed Representative Sampling, which applies spectral clustering within each batch to select class-balanced and representative samples. By incorporating neighbor relaxation across batches, this module ensures sample diversity and better captures the overall data distribution. To fully utilize the active set, we propose Similarity-enhanced TopK Prompt Construction, which retrieves the most relevant multimodal examples using a TopK similarity strategy and embeds their visual similarities with the query samples into the text prompts. This enhances LMMs' understanding of relationships, enabling more accurate and context-aware predictions. Experiments on nine specialized medical datasets across four LMMs show the effectiveness of our method.
Zhongyi Han, Yilong Yin
BIBM2
2025 From Pretraining to Pathology: How Noise Leads to Catastrophic Inheritance in Medical Models
abstract
Foundation models pretrained on web-scale data drive contemporary transfer learning in vision, language, and multimodal tasks. Recent work shows that mild label noise in these corpora may lift in-distribution accuracy yet sharply reduce out-of-distribution generalization, an effect known as catastrophic inheritance. Medical data is especially sensitive because annotations are scarce, domain shifts are large, and pretraining sources are noisy. We present the first systematic analysis of catastrophic inheritance in medical models. Controlled label-corruption experiments expose a clear structural collapse: as noise rises, the skewness and kurtosis of feature and logit distributions decline, signaling a flattened representation space and diminished discriminative detail. These higher-order statistics form a compact, interpretable marker of degradation in fine-grained tasks such as histopathology. Guided by this finding, we introduce a fine-tuning objective that restores skewness and kurtosis through two scalar regularizers added to the task loss. The method leaves the backbone unchanged and incurs negligible overhead. Tests on PLIP models trained with Twitter pathology images, as well as other large-scale vision and language backbones, show consistent gains in robustness and cross-domain accuracy under varied noise levels.
Hao Sun 0002, Zhongyi Han, Hao Chen 0102, Jindong Wang 0001, Xin Gao 0001, Yilong Yin
NeurIPS2
2025 Streamline automated biomedical discoveries with agentic bioinformatics
abstract
The emergence of artificial intelligence agents powered by large language models marks a transformative shift in computational biology. In this new paradigm, autonomous, adaptive, and intelligent agents are deployed to tackle complex biological challenges, leading to a new research field named agentic bioinformatics. Here, we explore the core principles, evolving methodologies, and diverse applications of agentic bioinformatics. We examine how agentic bioinformatics systems work synergistically to facilitate data-driven decision-making and enable self-directed exploration of biological datasets. Furthermore, we highlight the integration of agentic frameworks in key areas such as personalized medicine, drug discovery, and synthetic biology, illustrating their potential to revolutionize healthcare and biotechnology. In addition, we address the ethical, technical, and scalability challenges associated with agentic bioinformatics, identifying key opportunities for future advancements. By emphasizing the importance of interdisciplinary collaboration and innovation, we envision agentic bioinformatics as a major force in overcoming the grand challenges of modern biology, ultimately advancing both research and clinical applications.
Juexiao Zhou, Jindong Jiang, Zhongyi Han, Xin Gao 0001
Briefings Bioinform.3
2025 Active source-free open-set domain adaptation
Zhongyi Han, Hao Sun 0002, Yilong Yin
Knowl. Based Syst.2
2025 Alignclip: navigating the misalignments for robust vision-language generalization
abstract
In the realm of Vision-Language Pretraining models, achieving robust and adaptive representations is a cornerstone for successfully handling the unpredictability of real-world scenarios. This paper delves into two pivotal misalignment challenges inherent to Contrastive Language-Image Pre-training (CLIP) models: attention misalignment, which leads to an overemphasis on background elements rather than salient objects, and predictive category misalignment, characterized by the model’s struggle to discern between classes based on similarity. These misalignments undermine the representational stability essential for dynamic, real-world applications. To address these challenges, we propose AlignCLIP, an advanced fine-tuning methodology distinguished by its attention alignment loss, designed to calibrate the distribution of attention across multi-head attention layers. Furthermore, AlignCLIP introduces semantic label smoothing, a technique that leverages textual class similarities to refine prediction hierarchies. Through comprehensive experimentation on a variety of datasets and in scenarios involving distribution shifts and unseen classes, we demonstrate that AlignCLIP significantly enhances the stability of representations and shows superior generalization capabilities.
Zhongyi Han, Gongxu Luo, Hao Sun 0002, Bo Han 0003, Mingming Gong, Kun Zhang 0001, Tongliang Liu
Mach. Learn.1
2025 CTPT: Continual Test-time Prompt Tuning for vision-language models
Zhongyi Han, Xingbo Liu, Yilong Yin, Xin Gao 0001
Pattern Recognit.2
2025 Sampling Correction Approach With Interpolation Sliding Window for FY-3D/MERSI-II On-Orbit Calibration
abstract
In order to ensure the accuracy and reliability of the observational data from the remote sensor during its on-orbit operation, a sampling correction approach with the interpolation sliding window (ISWSCA) is proposed on the basis of the quadratic fitting and interpolation correction. The ISWSCA mitigates the effects of the inhomogeneous distribution of the reflective properties on the lunar surface, significantly enhancing the sampling correction accuracy of the lunar observation data by incorporating a subpixel correction term. Based on the lunar observation data from the Fengyun-3D (FY-3D)/Medium Resolution Spectral Imager (MERSI)-II collected between 2018 and 2023, the feasibility of the ISWSCA is verified, and the ISWSCA improves the sampling correction accuracy by 2.13% and 5.25% during low and high lunar phases in comparison to the scaling method of the lunar full-disk irradiance (LFDISM), together with the correction accuracy improved by 1.37% and 6% in low and high spatial resolutions. The ISWSCA gives the long time series of the normalized calibration coefficient, together with the calibration uncertainty of the effective data being analyzed, and the results show that the on-orbit stability of the FY-3D satellite is excellent in the visible (VIS) and near-infrared (NIF) bands, with the attenuation rate below 1.29% and the calibration uncertainty within 2.11%. This study has an important significance for the on-orbit radiation calibration of the spatial remote sensor.
Hanlin Xiao, Jingjing Ai, Jiaheng Yang, Zhongyi Han, Chengli Qi, Xiuqing Hu, Hanbo Zhen, Mingkun Wang
IEEE Trans. Geosci. Remote. Sens.7
2025 Improving Representation of High-Frequency Components for Medical Visual Foundation Models
abstract
Foundation models have attracted significant attention for their impressive generalizability across diverse downstream tasks. However, they are demonstrated to exhibit great limitations in representing high-frequency components and fine-grained details. In many medical imaging tasks, precise representation of such information is crucial due to the inherently intricate anatomical structures, sub-visual features, and complex boundaries involved. Consequently, the limited representation of prevalent foundation models can result in considerable performance degradation or even failure in these tasks. To address these challenges, we propose a novel pretraining strategy for both 2D images and 3D volumes, named Frequency-advanced Representation Autoencoder (Frepa). Through high-frequency masking and low-frequency perturbation combined with embedding consistency learning, Frepa encourages the encoder to effectively represent and preserve high-frequency components in the image embeddings. Additionally, we introduce an innovative histogram-equalized image masking strategy, extending the Masked Autoencoder approach beyond ViT to other architectures such as Swin-Transformer and convolutional networks. We develop Frepa across nine medical modalities and validate it on 32 downstream tasks for both 2D images and 3D volumes. Without fine-tuning, Frepa can outperform other self-supervised pretraining methods and, in some cases, even surpasses task-specific foundation models. This improvement is particularly significant for tasks involving fine-grained details, such as achieving up to a +15% increase in dice score for retina vessel segmentation and a +8% increase in IoU for lung tumor detection. Further experiment quantitatively reveals that Frepa enables superior high-frequency representations and preservation in the embeddings, underscoring its potential for developing more generalized and universal medical image foundation models.
Yuetan Chu, Yilan Zhang, Zhongyi Han, Longxi Zhou, Gongning Luo, Xin Gao 0001
IEEE Trans. Medical Imaging3
2025 HCVP: Leveraging Hierarchical Contrastive Visual Prompt for Domain Generalization
abstract
Domain Generalization (DG) endeavors to create machine learning models that excel in unseen scenarios by learning invariant features. In DG, the prevalent practice of constraining models to a fixed structure or uniform parameterization to encapsulate invariant features can inadvertently blend specific aspects. Such an approach struggles with nuanced differentiation of inter-domain variations and may exhibit bias towards certain domains, hindering the precise learning of domain-invariant features. Recognizing this, we introduce a novel method designed to supplement the model with domain-level and task-specific characteristics. This approach aims to guide the model in more effectively separating invariant features from specific characteristics, thereby boosting the generalization. Building on the emerging trend of visual prompts in the DG paradigm, our work introduces the novelHierarchicalContrastiveVisualPrompt (HCVP) methodology. This represents a significant advancement in the field, setting itself apart with a unique generative approach to prompts, alongside an explicit model structure and specialized loss functions. Differing from traditional visual prompts that are often shared across entire datasets, HCVP utilizes a hierarchical prompt generation network enhanced by prompt contrastive learning. These generative prompts are instance-dependent, catering to the unique characteristics inherent to different domains and tasks. Additionally, we devise a prompt modulation network that serves as a bridge, effectively incorporating the generated visual prompts into the vision transformer backbone. Experiments conducted on five DG datasets demonstrate the effectiveness of HCVP, outperforming both established DG algorithms and adaptation protocols.
Guanglin Zhou, Zhongyi Han, Shiming Chen 0002, Biwei Huang, Liming Zhu 0001, Tongliang Liu, Lina Yao 0001, Kun Zhang 0001
IEEE Trans. Multim.2
2024 Exploring Channel-Aware Typical Features for Out-of-Distribution Detection
abstract
Detecting out-of-distribution (OOD) data is essential to ensure the reliability of machine learning models when deployed in real-world scenarios. Different from most previous test-time OOD detection methods that focus on designing OOD scores, we delve into the challenges in OOD detection from the perspective of typicality and regard the feature’s high-probability region as the feature’s typical set. However, the existing typical-feature-based OOD detection method implies an assumption: the proportion of typical feature sets for each channel is fixed. According to our experimental analysis, each channel contributes differently to OOD detection. Adopting a fixed proportion for all channels results in several channels losing too many typical features or incorporating too many abnormal features, resulting in low performance. Therefore, exploring the channel-aware typical features is crucial to better-separating ID and OOD data. Driven by this insight, we propose expLoring channel-Aware tyPical featureS (LAPS). Firstly, LAPS obtains the channel-aware typical set by calibrating the channel-level typical set with the global typical set from the mean and standard deviation. Then, LAPS rectifies the features into channel-aware typical sets to obtain channel-aware typical features. Finally, LAPS leverages the channel-aware typical features to calculate the energy score for OOD detection. Theoretical and visual analyses verify that LAPS achieves a better bias-variance trade-off. Experiments verify the effectiveness and generalization of LAPS under different architectures and OOD scores.
Rundong He, Zhongyi Han, Wan Su, Yilong Yin, Tongliang Liu, Yongshun Gong
AAAI3
2024 Navigating the Unknown: A Novel MGUAN Framework for Medical Image Recognition Across Dynamic Domains
abstract
Machine learning has significantly advanced medical image recognition, enhancing diagnostic accuracy in various applications. However, these advancements primarily apply to scenarios with consistent data distributions, a condition rarely met in real-world clinical settings. In real-world clinical environments, variations in device specifications and patient demographics introduce distribution shifts, class imbalance and unknown class challenges, undermining model robustness. Addressing this, we present the Medical image recognition under Generalized Universal Domain Adaptation (MGUDA) concept, targeting distribution shifts, class imbalance and unknown class detection. Our innovative Medical Dual-Prototype Adaptation Network (MDPAN) framework, integrating dual prototype learning, dual prototype employment, and weighted multi-class adversarial alignment, adeptly confronts these issues. Extensive evaluations on diverse medical image datasets validate MDPAN’s superiority in managing class imbalances and enhancing target domain classification, marking a pivotal step in robust medical image recognition across variable domains.
Wan Su, Rundong He, Zhongyi Han, Yilong Yin
BIBM4
2024 Discriminability-Driven Channel Selection for Out-of-Distribution Detection
abstract
Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world environments. Activation-based methods are a key approach in OOD detection, working to mitigate overconfident predictions of OOD data. These techniques rectifying anomalous activations, enhancing the distinguishability between in-distribution (ID) data and OOD data. However, they assume by default that every channel is necessary for OOD detection, and rectify anomalous activations in each channel. Empirical evidence has shown that there is a significant difference among various channels in OOD detection, and discarding some channels can greatly enhance the performance of OOD detection. Based on this insight, we propose Discriminability-Driven Channel Selection (DDCS), which leverages an adaptive channel selection by estimating the discriminative score of each channel to boost OOD detection. The discriminative score takes inter-class similarity and inter-class variance of training data into account. However, the estimation of discriminative score itself is susceptible to anomalous activations. To better estimate score, we pre-rectify anomalous activations for each channel mildly. The experimental results show that DDCS achieves state-of-the-art performance on CIFAR and ImageNet-1K benchmarks. Moreover, DDCS can generalize to different backbones and OOD scores.
Rundong He, Yicong Dong, Zhongyi Han, Yilong Yin
CVPR4
2024 Visual Out-of-Distribution Detection in Open-Set Noisy Environments
Rundong He, Zhongyi Han, Xiushan Nie, Yilong Yin, Xiaojun Chang
Int. J. Comput. Vis.2
2024 Generalized Universal Domain Adaptation
Wan Su, Zhongyi Han, Xingbo Liu, Yilong Yin
Knowl. Based Syst.2
2024 BIAS: Bridging Inactive and Active Samples for active source free domain adaptation
Zhongyi Han, Yilong Yin
Knowl. Based Syst.2
2024 SAFER-STUDENT for Safe Deep Semi-Supervised Learning With Unseen-Class Unlabeled Data
abstract
Deep semi-supervised learning (SSL) methods aim to utilize abundant unlabeled data to improve the seen-class classification. However, in the open-world scenario, collected unlabeled data tend to contain unseen-class data, which would degrade the generalization to seen-class classification. Formally, we define the problem as safe deep semi-supervised learning with unseen-class unlabeled data. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the lack of suitable score function, the uncalibrated model, and the small number of labeled data. To this end, we propose a safe SSL method called SAFER-STUDENT from the teacher-student view. First, to enhance the ability of teacher model to identify seen and unseen classes, we propose a general scoring framework calledDiscrepancy withRaw (DR). Second, based on unseen-class data mined by teacher model from unlabeled data, we calibrate student model by newly proposedUnseen-classEnergy-boundedCalibration (UEC) loss. Third, based on seen-class data mined by teacher model from unlabeled data, we proposeWeightedConfirmationBiasElimination (WCBE) loss to boost seen-class classification of student model. Extensive studies show that SAFER-STUDENT remarkably outperforms the state-of-the-art, verifying the effectiveness of our method in the under-explored problem.
Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin
IEEE Trans. Knowl. Data Eng.2
2023 Discriminability and Transferability Estimation: A Bayesian Source Importance Estimation Approach for Multi-Source-Free Domain Adaptation
abstract
Source free domain adaptation (SFDA) transfers a single-source model to the unlabeled target domain without accessing the source data. With the intelligence development of various fields, a zoo of source models is more commonly available, arising in a new setting called multi-source-free domain adaptation (MSFDA). We find that the critical inborn challenge of MSFDA is how to estimate the importance (contribution) of each source model. In this paper, we shed new Bayesian light on the fact that the posterior probability of source importance connects to discriminability and transferability. We propose Discriminability And Transferability Estimation (DATE), a universal solution for source importance estimation. Specifically, a proxy discriminability perception module equips with habitat uncertainty and density to evaluate each sample's surrounding environment. A source-similarity transferability perception module quantifies the data distribution similarity and encourages the transferability to be reasonably distributed with a domain diversity loss. Extensive experiments show that DATE can precisely and objectively estimate the source importance and outperform prior arts by non-trivial margins. Moreover, experiments demonstrate that DATE can take the most popular SFDA networks as backbones and make them become advanced MSFDA solutions.
Zhongyi Han, Zhiyan Zhang, Rundong He, Wan Su, Xiaoming Xi, Yilong Yin
AAAI1
2023 MHPL: Minimum Happy Points Learning for Active Source Free Domain Adaptation
abstract
Source free domain adaptation (SFDA) aims to transfer a trained source model to the unlabeled target domain without accessing the source data. However, the SFDA setting faces a performance bottleneck due to the absence of source data and target supervised information, as evidenced by the limited performance gains of the newest SFDA methods. Active source free domain adaptation (ASFDA) can break through the problem by exploring and exploiting a small set of informative samples via active learning. In this paper, we first find that those satisfying the proper-ties of neighbor-chaotic, individual-different, and source-dissimilar are the best points to select. We define them as the minimum happy (MH) points challenging to explore with existing methods. We propose minimum happy points learning (MHPL) to explore and exploit MH points actively. We design three unique strategies: neighbor environment uncertainty, neighbor diversity relaxation, and one-shot querying, to explore the MH points. Further, to fully exploit MH points in the learning process, we design a neighbor focal loss that assigns the weighted neighbor purity to the cross entropy loss of MH points to make the model focus more on them. Extensive experiments verify that MHPL remarkably exceeds the various types of baselines and achieves significant performance gains at a small cost of labeling.
Zhongyi Han, Zhiyan Zhang, Rundong He, Yilong Yin
CVPR2
2023 Topological Structure Learning for Weakly-Supervised Out-of-Distribution Detection
abstract
Out-of-distribution~(OOD) detection is the key to deploying models safely in the open world. For OOD detection, collecting sufficient in-distribution~(ID) labeled data is usually more time-consuming and costly than unlabeled data. When ID labeled data is limited, the previous OOD detection methods are no longer superior due to their high dependence on the amount of ID labeled data. Based on limited ID labeled data and sufficient unlabeled data, we define a new setting called Weakly-Supervised Out-of-Distribution Detection (WSOOD). To solve the new problem, we propose an effective method called Topological Structure Learning (TSL). Firstly, TSL uses a contrastive learning method to build the initial topological structure space for ID and OOD data. Secondly, TSL mines effective topological connections in the initial topological space. Finally, based on limited ID labeled data and mined topological connections, TSL reconstructs the topological structure in a new topological space to increase the separability of ID and OOD instances. Extensive studies on several representative datasets show that TSL remarkably outperforms the state-of-the-art, verifying the validity and robustness of our method in the new setting of WSOOD.
Rundong He, Rongxue Li, Zhongyi Han, Xihong Yang, Yilong Yin
ACM Multimedia3
2023 LHAct: Rectifying Extremely Low and High Activations for Out-of-Distribution Detection
abstract
In recent years, out-of-distribution (OOD) detection has emerged as a crucial research area, especially when deploying AI products in real-world scenarios. OOD detection researchers have made significant efforts to mitigate the adverse effects of abnormal activation values (abbr. activations) that refer to the outputs of the activation function acted on feature maps. Since abnormal activations would cause difficulty in separating ID and OOD data, the previous unified solution is to rectify the extremely high abnormal activations by clipping them with a pre-defined threshold or filtering them with a low-pass filter. However, it ignores the extremely low abnormal activations, and the proposed rectification strategy is always suboptimal because the used rectification function is non-convergence or high-intensity convergence, leading to under-rectification or over-rectification. In this paper, we propose an approach called Rectifying Extremely Low and High Activations (LHAct). LHAct includes a newly-designed function to rectify the extremely low and high activations at the same time. Specifically, LHAct increases the difference of means between ID and OOD activation distributions while decreasing their variances after processing the original activations. Our theoretical analyses demonstrate that LHAct significantly enhances the separability of ID and OOD data. By conducting extensive experiments, we demonstrate that LHAct surpasses previous activation-based methods significantly and generalizes well to other architectures and OOD scores. Code is available at: https://github.com/ystyuan/LHAct.git.
Rundong He, Zhongyi Han, Yilong Yin
ACM Multimedia3
2023 Subclass-Dominant Label Noise: A Counterexample for the Success of Early Stopping
abstract
In this paper, we empirically investigate a previously overlooked and widespread type of label noise, subclass-dominant label noise (SDN). Our findings reveal that, during the early stages of training, deep neural networks can rapidly memorize mislabeled examples in SDN. This phenomenon poses challenges in effectively selecting confident examples using conventional early stopping techniques. To address this issue, we delve into the properties of SDN and observe that long-trained representations are superior at capturing the high-level semantics of mislabeled examples, leading to a clustering effect where similar examples are grouped together. Based on this observation, we propose a novel method called NoiseCluster that leverages the geometric structures of long-trained representations to identify and correct SDN. Our experiments demonstrate that NoiseCluster outperforms state-of-the-art baselines on both synthetic and real-world datasets, highlighting the importance of addressing SDN in learning with noisy labels. The code is available at https://github.com/tmllab/2023_NeurIPS_SDN.
Yingbin Bai, Zhongyi Han, Erkun Yang, Jun Yu 0001, Bo Han 0003, Dadong Wang, Tongliang Liu
NeurIPS2
2023 Abductive subconcept learning
Zhongyi Han, Le-Wen Cai, Wang-Zhou Dai, Yu-Xuan Huang, Benzheng Wei, Yilong Yin
Sci. China Inf. Sci.1
2023 Towards Accurate and Robust Domain Adaptation Under Multiple Noisy Environments
abstract
In many non-stationary environments, machine learning algorithms usually confront the distribution shift scenarios. Previous domain adaptation methods have achieved great success. However, they would lose algorithm robustness in multiple noisy environments where the examples of source domain become corrupted by label noise, feature noise, or open-set noise. In this paper, we report our attempt toward achieving noise-robust domain adaptation. We first give a theoretical analysis and find that different noises have disparate impacts on the expected target risk. To eliminate the effect of source noises, we propose offline curriculum learning minimizing a newly-defined empirical source risk. We suggest a proxy distribution-based margin discrepancy to gradually decrease the noisy distribution distance to reduce the impact of source noises. We propose an energy estimator for assessing the outlier degree of open-set-noise examples to defeat the harmful influence. We also suggest robust parameter learning to mitigate the negative effect further and learn domain-invariant feature representations. Finally, we seamlessly transform these components into an adversarial network that performs efficient joint optimization for them. A series of empirical studies on the benchmark datasets and the COVID-19 screening task show that our algorithm remarkably outperforms the state-of-the-art, with over 10% accuracy improvements in some transfer tasks.
Zhongyi Han, Xian-Jin Gui, Haoliang Sun, Yilong Yin, Shuo Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Neighborhood-based credibility anchor learning for universal domain adaptation
Wan Su, Zhongyi Han, Rundong He, Benzheng Wei, Xueying He, Yilong Yin
Pattern Recognit.2
2023 Lunar Phase Function Oversampling Correction Method for FY3D/MERSI On-Orbit Calibration
abstract
In order to further improve the accuracy of the lunar radiation correction, an oversampling correction method based on the lunar phase function was firstly proposed in this paper. This method avoided the dependence of the MODIS algorithm on the relative space position and velocity accuracy, and overcame the limitation of the classical SeaWIFS algorithm only applicable to low lunar phase angles, in order to realize the full lunar phase observation of the sky, earth and space. Based on the least square fitting of the actual image data together with the analytic solution of the lunar phase function, the exact expression of the oversampling correction coefficient was given, and the spatial resolution of the remote sensor was improved to the sub-pixel level by the numerical difference calculation. The effectiveness of the novel method was validated through the Terra MODIS freemoon data with the lunar phase angle range of [55°,80°], and the lunar phase fitting function was highly consistent with the lunar phase shape observed by the remote sensor, together with the normalized calibration coefficient calculated by the novel method coinciding basically with the results of the MODIS maneuver calibration data with the lunar phase angle range of about 55°, which demonstrated the competitive efficiency and accuracy of the novel method. To evaluate the on-orbit stability of the FY3D/MERSI satellite, the long time series of the normalized calibration coefficient was obtained by the oversampling correction method, and the calibration uncertainty of the effective data was analyzed, showing no significant change for the gain of the FY3D/MERSI in most bands. This research had significance for the lunar observation together with the on-orbit calibration and deep space detection.
Zhongyi Han, Yichao Zheng, Jingjing Ai, Hanlin Xiao, Xiuqing Hu, Chengli Qi, Gongju Liu, Zhaoming Bai
IEEE Trans. Geosci. Remote. Sens.2
2022 Not All Parameters Should Be Treated Equally: Deep Safe Semi-supervised Learning under Class Distribution Mismatch
abstract
Deep semi-supervised learning (SSL) aims to utilize a sizeable unlabeled set to train deep networks, thereby reducing the dependence on labeled instances. However, the unlabeled set often carries unseen classes that cause the deep SSL algorithm to lose generalization. Previous works focus on the data level that they attempt to remove unseen class data or assign lower weight to them but could not eliminate their adverse effects on the SSL algorithm. Rather than focusing on the data level, this paper turns attention to the model parameter level. We find that only partial parameters are essential for seen-class classification, termed safe parameters. In contrast, the other parameters tend to fit irrelevant data, termed harmful parameters. Driven by this insight, we propose Safe Parameter Learning (SPL) to discover safe parameters and make the harmful parameters inactive, such that we can mitigate the adverse effects caused by unseen-class data. Specifically, we firstly design an effective strategy to divide all parameters in the pre-trained SSL model into safe and harmful ones. Then, we introduce a bi-level optimization strategy to update the safe parameters and kill the harmful parameters. Extensive experiments show that SPL outperforms the state-of-the-art SSL methods on all the benchmarks by a large margin. Moreover, experiments demonstrate that SPL can be integrated into the most popular deep SSL networks and be easily extended to handle other cases of class distribution mismatch.
Rundong He, Zhongyi Han, Yang Yang 0074, Yilong Yin
AAAI2
2022 SNAIL: Semi-Separated Uncertainty Adversarial Learning for Universal Domain Adaptation
Zhongyi Han, Wan Su, Rundong He, Yilong Yin
ACML1
2022 Transferable Discriminative Learning for Medical Open-Set Domain Adaptation: Application to Pneumonia Classification
abstract
Previous pneumonia classification algorithms have succeeded in the clinic under closed and static environments. However, in the real world, the emergence of new categories (e.g., COVID-19) and changes in data distribution will cause the existing methods to lose their robustness. In this paper, we formalize this problem as medical open-set domain adaptation under open and dynamic environments. The critical challenge of this problem is to accurately detect the open class samples with subtle differences from the common class. To achieve that, we propose transferable discriminative learning that remarkably achieves robust pneumonia classification with distribution shift and open class emerging. First, we propose the transferable high-density clustering module to detect open class samples and obtain reliable common class samples by considering the density degree. Secondly, we present the transferable triplet loss to enlarge the semantic feature difference between common class and open class samples. Finally, we design the transferable scoring function to detect open class samples effectively. A series of empirical studies show that our algorithm remarkably outperforms state-of-the-art methods. This result demonstrates its potential as a clinical tool for medical open-set domain adaptation.
Wan Su, Zhongyi Han, Yilong Yin
BIBM3
2022 Safe-Student for Safe Deep Semi-Supervised Learning with Unseen-Class Unlabeled Data
abstract
Deep semi-supervised learning (SSL) methods aim to take advantage of abundant unlabeled data to improve the algorithm performance. In this paper, we consider the problem of safe SSL scenario where unseen-class instances appear in the unlabeled data. This setting is essential and commonly appears in a variety of real applications. One intuitive solution is removing these unseen-class instances after detecting them during the SSL process. Nevertheless, the performance of unseen-class identification is limited by the small number of labeled data and ignoring the availability of unlabeled data. To take advantage of these unseen-class data and ensure performance, we propose a safe SSL method called SAFE-STUDENT from the teacher-student view. Firstly, a new scoring function called energy-discrepancy (ED) is proposed to help the teacher model improve the security of instances selection. Then, a novel unseen-class label distribution learning mechanism mitigates the unseen-class perturbation by calibrating the unseen-class label distribution. Finally, we propose an iterative optimization strategy to facilitate teacher-student network learning. Extensive studies on several representative datasets show that SAFE-STUDENT remarkably outperforms the state-of-the-art, verifying the feasibility and robustness of our method in the under-explored problem.
Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin
CVPR2
2022 Exploring Domain-Invariant Parameters for Source Free Domain Adaptation
abstract
Source-free domain adaptation (SFDA) newly emerges to transfer the relevant knowledge of a well-trained source model to an unlabeled target domain, which is critical in various privacy-preserving scenarios. Most existing methods focus on learning the domain-invariant representations depending solely on the target data, leading to the obtained representations are target-specific. In this way, they cannot fully address the distribution shift problem across domains. In contrast, we provide a fascinating insight: rather than attempting to learn domain-invariant representations, it is better to explore the domain-invariant parameters of the source model. The motivation behind this insight is clear: the domain-invariant representations are dominated by only partial parameters of an available deep source model. We devise the Domain-Invariant Parameter Exploring (DIPE) approach to capture such domain-invariant parameters in the source model to generate domain-invariant representations. A distinguishing method is developed correspondingly for two types of parameters, i.e., domain-invariant and domain-specific parameters, as well as an effective update strategy based on the clustering correction technique and a target hypothesis is proposed. Extensive experiments verify that DIPE successfully exceeds the current state-of-the-art models on many domain adaptation datasets.
Zhongyi Han, Yongshun Gong, Yilong Yin
CVPR2
2022 RONF: Reliable Outlier Synthesis under Noisy Feature Space for Out-of-Distribution Detection
abstract
Out-of-distribution~(OOD) detection is fundamental to guaranteeing the reliability of multimedia applications during deployment in the open world. However, due to the lack of supervision signals from OOD data, the current model easily outputs overconfident predictions to OOD data during the inference phase. Several previous methods rely on large-scale auxiliary OOD datasets for model regularization. However, obtaining suitable and clean large-scale auxiliary OOD datasets is usually challenging. In this paper, we present Reliable Outlier synthesis under Noisy Feature space (RONF), which synthesizes reliable virtual outliers in noisy feature space to provide supervision signals for model regularization. Specifically, RONF first introduces a novel virtual outlier synthesis strategy Boundary Feature Mixup (BFM), which mixes up samples from the low-likelihood region of the class-conditional distribution in the feature space. However, the feature space is noisy due to the spurious features, which cause unreliable outlier synthesizing. To mitigate this problem, RONF then introduces Optimal Parameter Learning (OPL) to obtain desirable features and remove spurious features. Alongside, RONF proposes a provable and effective scoring function called Energy with Energy Discrepancy (EED) for the uncertainty measurement of OOD data. Extensive studies on several representative datasets of multimedia applications show that RONF outperforms the state-of-the-arts remarkably
Rundong He, Zhongyi Han, Xiankai Lu, Yilong Yin
ACM Multimedia2
2022 Towards safe and robust weakly-supervised anomaly detection under subpopulation shift
Rundong He, Zhongyi Han, Yilong Yin
Knowl. Based Syst.2
2022 Learning to rectify for robust learning with noisy labels
Haoliang Sun, Chenhui Guo, Qi Wei 0004, Zhongyi Han, Yilong Yin
Pattern Recognit.4
2022 Learning Transferable Parameters for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) enables a learning machine to adapt from a labeled source domain to an unlabeled target domain under the distribution shift. Thanks to the strong representation ability of deep neural networks, recent remarkable achievements in UDA resort to learning domain-invariant features. Intuitively, the goal is that a good feature representation and the hypothesis learned from the source domain can generalize well to the target domain. However, the learning processes of domain-invariant features and source hypotheses inevitably involve domain-specific information that would degrade the generalizability of UDA models on the target domain. The lottery ticket hypothesis proves that only partial parameters are essential for generalization. Motivated by it, we find in this paper that only partial parameters are essential for learning domain-invariant information. Such parameters are termed transferable parameters that can generalize well in UDA. In contrast, the rest parameters tend to fit domain-specific details and often cause the failure of generalization, which are termed untransferable parameters. Driven by this insight, we propose Transferable Parameter Learning (TransPar) to reduce the side effect of domain-specific information in the learning process and thus enhance the memorization of domain-invariant information. Specifically, according to the distribution discrepancy degree, we divide all parameters into transferable and untransferable ones in each training iteration. We then perform separate update rules for the two types of parameters. Extensive experiments on image classification and regression tasks (keypoint detection) show that TransPar outperforms prior arts by non-trivial margins. Moreover, experiments demonstrate that TransPar can be integrated into the most popular deep UDA networks and be easily extended to handle any data distribution shift scenarios.
Zhongyi Han, Haoliang Sun, Yilong Yin
IEEE Trans. Image Process.1
2021 Unifying neural learning and symbolic reasoning for spinal medical report generation
Zhongyi Han, Benzheng Wei, Xiaoming Xi, Bo Chen 0013, Yilong Yin, Shuo Li 0001
Medical Image Anal.1
2020 Towards Accurate and Robust Domain Adaptation under Noisy Environments
abstract
In non-stationary environments, learning machines usually confront the domain adaptation scenario where the data distribution does change over time. Previous domain adaptation works have achieved great success in theory and practice. However, they always lose robustness in noisy environments where the labels and features of examples from the source domain become corrupted. In this paper, we report our attempt towards achieving accurate noise-robust domain adaptation. We first give a theoretical analysis that reveals how harmful noises influence unsupervised domain adaptation. To eliminate the effect of label noise, we propose an offline curriculum learning for minimizing a newly-defined empirical source risk. To reduce the impact of feature noise, we propose a proxy distribution based margin discrepancy. We seamlessly transform our methods into an adversarial network that performs efficient joint optimization for them, successfully mitigating the negative influence from both data corruption and distribution shift. A series of empirical studies show that our algorithm remarkably outperforms state of the art, over 10% accuracy improvements in some domain adaptation tasks under noisy environments.
Zhongyi Han, Xian-Jin Gui, Chaoran Cui, Yilong Yin
IJCAI1
2020 Recursive narrative alignment for movie narrating
Zhongyi Han, Hongbo Wu, Benzheng Wei, Yilong Yin, Shuo Li 0001
Sci. China Inf. Sci.1
2020 DRAN: Deep recurrent adversarial network for automated pancreas segmentation
abstract
Automated pancreas segmentation in abdominal computed tomography (CT) scans is of high clinical relevance (i.e. pancreas cancer diagnosis and prognosis), but extremely difficult because the pancreas is a soft, small, and flexible abdominal organ with high anatomical variability, which causes the previous segmentation methods to result in low precision. In this study, the authors present a new deep recurrent adversarial network (DRAN) to tackle this challenge. DRAN contains three steps: (i) preserving global resolution of CT scans and modifying the receptive field of kernel adaptively through a dilated convolution autoencoder module; (ii) modelling contextual spatial correlation between neighbouring CT scan patches benefits from a specially designed local long short‐term memory module; and (iii) improving the performance and generalisation by leveraging an adversarial module, which can constrain the spatial smoothness consistency between continuous CT scans based on the long‐range spatial interaction. The system is evaluated on a dataset of 80 manually segmented CT volumes, using four‐fold cross‐validation. Its performance surpasses other state‐of‐the‐art methods, with the Dice similarity coefficient of and pixel‐wise accuracy of . Also, they perform a qualitative evaluation by an expert further revealing the effectiveness and potential of their DRAN as a clinical segmentation tool.
Yang Ning, Zhongyi Han, Caiming Zhang 0001
IET Image Process.2
2020 MMCL-Net: Spinal disease diagnosis in global mode using progressive multi-task joint learning
Yanfei Hong, Benzheng Wei, Zhongyi Han, Xiang Li 0114, Yuanjie Zheng, Shuo Li 0001
Neurocomputing3
2020 Accurate Screening of COVID-19 Using Attention-Based Deep 3D Multiple Instance Learning
abstract
Automated Screening of COVID-19 from chest CT is of emergency and importance during the outbreak of SARS-CoV-2 worldwide in 2020. However, accurate screening of COVID-19 is still a massive challenge due to the spatial complexity of 3D volumes, the labeling difficulty of infection areas, and the slight discrepancy between COVID-19 and other viral pneumonia in chest CT. While a few pioneering works have made significant progress, they are either demanding manual annotations of infection areas or lack of interpretability. In this paper, we report our attempt towards achieving highly accurate and interpretable screening of COVID-19 from chest CT with weak labels. We propose an attention-based deep 3D multiple instance learning (AD3D-MIL) where a patient-level label is assigned to a 3D chest CT that is viewed as a bag of instances. AD3D-MIL can semantically generate deep 3D instances following the possible infection area. AD3D-MIL further applies an attention-based pooling approach to 3D instances to provide insight into each instance's contribution to the bag label. AD3D-MIL finally learns Bernoulli distributions of the bag-level labels for more accessible learning. We collected 460 chest CT examples: 230 CT examples from 79 patients with COVID-19, 100 CT examples from 100 patients with common pneumonia, and 130 CT examples from 130 people without pneumonia. A series of empirical studies show that our algorithm achieves an overall accuracy of 97.9%, AUC of 99.0%, and Cohen kappa score of 95.7%. These advantages endow our algorithm as an efficient assisted tool in the screening of COVID-19.
Zhongyi Han, Benzheng Wei, Yanfei Hong, Jinyu Cong, Haifeng Wei
IEEE Trans. Medical Imaging1
2018 Automated Pancreas Segmentation Using Recurrent Adversarial Learning
Yang Ning, Zhongyi Han, Caiming Zhang 0001
BIBM2
2018 Towards Automatic Report Generation in Spine Radiology Using Weakly Supervised Framework
Zhongyi Han, Benzheng Wei, Stephanie Leung, Jonathan Chung 0002, Shuo Li 0001
MICCAI (4)1
2018 Spine-GAN: Semantic segmentation of multiple spinal structures
Zhongyi Han, Benzheng Wei, Ashley Mercado, Stephanie Leung, Shuo Li 0001
Medical Image Anal.1