VLDB 2026 Research / reviewers in the wild / expert
Yawen Huang
dblp:122/0805
· DBLP profile ↗
89ranked-venue papers
12as first author
83since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 8 first-author · 50 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 40 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 13 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unsupervised domain adaptation without source data for visual classification via adaptive confidence-driven mechanism
Ziyun Cai, Jie Song 0014, Yawen Huang, Changhui Hu 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Efficient Multiuser Searchable and Revocable Data Sharing Scheme for IoVabstractThe Internet of Vehicles (IoV), as a critical component of future intelligent transportation systems, enables vehicles to generate and exchange massive volumes of sensing data in real time. To facilitate efficient data sharing and alleviate the burden of local storage, vehicle data is typically encrypted and outsourced to the cloud, with data availability ensured through searchable encryption. However, existing IoV data sharing frameworks offer limited support for secure multi-keyword search and dynamic user revocation. Moreover, most current solutions rely on public-key searchable encryption, which compromises the efficiency required for data processing in IoV environments. To address these challenges, we propose an efficient multi-user searchable and revocable data sharing scheme (EMUSR-SSE) tailored for IoV. EMUSR-SSE extends symmetric searchable encryption to support efficient multi-user access control and trustworthy multi-keyword search by designing a proxy matrix transformation mechanism integrated with Intel SGX. To enable flexible and scalable user revocation, EMUSR-SSE introduces a revocable dynamic sparse Merkle tree structure to manage multi-user permissions effectively. Furthermore, by leveraging the trusted execution environment within SGX and encrypting each data item individually, EMUSR-SSE significantly mitigates the risk of key leakage. Security analysis and experimental evaluations demonstrate that EMUSR-SSE achieves high efficiency and strong practical performance in dynamic IoV environments. Ping Wang 0086, Fei Tang 0001, Haining Luo, Fengjie Peng, Huihui Zhu 0001, Ankui Jing, Yawen Huang |
IEEE Internet Things J. | 7 |
| 2026 | Semi-supervised crowd counting from unlabeled data
Haoran Duan 0001, Yawen Huang, Yang Long 0001, Xian Wu 0001, Feiyue Huang, Shaoxin Li 0001 |
Pattern Recognit. | 2 |
| 2026 | ConRF: Zero-shot stylization of 3D scenes with conditioned radiation fieldsabstract• We propose a novel method that leverages CLIP for zero-shot 3D scene artistic style transfer by a single condition (i.e. image or text). • We introduce a mapping network to alleviate the ambiguity in CLIP features related to style. • We present a 3D selection volume that allows for localized style manipulation within 3D scenes, expanding the possibilities in scene stylization and manipulation. Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D scenes utilizing text or visual input as conditioning factors. We introduce ConRF, a novel method of zero-shot stylization. Specifically, due to the ambiguity of CLIP features, we employ a conversion process that maps the CLIP feature space to the style space of a pre-trained VGG network and then refine the CLIP multi-modal knowledge into a style transfer neural radiation field. Additionally, we use a 3D volumetric representation to perform local style transfer. By combining these operations, ConRF offers the capability to utilize either text or images as references, resulting in the generation of sequences with novel views enhanced by global or local stylization. Our experiment demonstrates that ConRF outperforms other existing methods for 3D scene and single-text stylization in terms of visual quality. Code is available: https://xingy038.github.io/ConRF/ . Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001 |
Pattern Recognit. | 5 |
| 2026 | Open-world Weakly-Supervised Object Localization
Jinheng Xie, Zhaochuan Luo, Rouyi Li, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Yang Zhang 0012, LinLin Shen, Zheng Shou 0001 |
Pattern Recognit. | 4 |
| 2026 | Harmonized medical federated learning via redundancy-aware client consistency
Jingjun Yi, Yuexiang Li, Qi Bi, Wei Ji 0011, Huimin Huang 0002, Yawen Huang, Yefeng Zheng 0001, Feiyue Huang |
Pattern Recognit. | 7 |
| 2026 | MUSCLE: A New Perspective to Multi-Scale Fusion for Medical Image Classification Based on the Theory of EvidenceabstractIn the field of medical image analysis, medical image classification is one of the most fundamental and critical tasks. Current researches often rely on the off-the-shelf backbone networks derived from the field of computer vision, hoping to achieve satisfactory classification performance for medical images. However, given the characteristics of medical images, such as scattered distribution and varying sizes of lesions, features extracted with a single scale from the existing backbones often fail to perform accurate medical image classification. To this end, we propose a novel multi-scale learning paradigm, namely MUlti-SCale Learning with trusted Evidences (MUSCLE), which extracts and integrates features from different scales based on shape the theory of evidence, to generate the more comprehensive feature representation for the medical image classification task. Particularly, the proposed MUSCLE first estimates the uncertainties of features extracted from different scales/stages of the classification backbone as the evidences, and accordingly form the opinions regarding to the feature trustworthiness via a set of evidential deep neural networks. Then, these opinions on different scales of features are ensembled to yield an aggregated opinion, which can be used to adaptively tune the weights of multi-scale features for scatteredly distributed and size-varying lesions, and consequently improve the network capacity for accurate medical image classification. Our MUSCLE paradigm has been evaluated on five publicly available medical image datasets. The experimental results show that the proposed MUSCLE not only improves the accuracy of the original backbone network, but also enhances the reliability and interpretability of model decisions with the trusted evidences (https://github.com/Q4CS/MUSCLE). Junlai Qiu, Junyue Cao, Yawen Huang, Ziwei Zhu 0005, Fubo Wang, Cheng Lu 0001, Yuexiang Li, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2026 | SORT-LFR: Revisiting SORT for Multi-Object Tracking in Low-Frame-Rate VideosabstractFor certain applications like highway surveillance systems, only low-frame-rate videos are recorded, which presents a huge challenge to existing trackers, as objects tend to undergo far more abrupt changes in location, motion, and appearance between successive frames compared to normal frame rates. To handle the above challenges, we propose a novel approach, namely$\mathbb {SORT}$-$\mathbb {LFR}$, for$\mathbb {S}$imple$\mathbb {O}$nline and$\mathbb {R}$ealtime$\mathbb {T}$racking in$\mathbb {L}$ow-$\mathbb {F}$rame-$\mathbb {R}$ate videos, which consists of following techniques: 1) A feature-prior association strategy to improve the capability to track new objects with significant displacements; 2) A Kalman filter using acceleration in state space (accel-fused Kalman filter) to improve the motion estimation capability for non-constant velocity moving objects; 3) A detection-guided adaptive exponential moving average (DG-AEMA) feature update mechanism to enhance feature temporal modeling capability for tracked objects; 4) A trajectory-covariance threshold tuning (TCTT) method to filter out incorrect association results. Through these techniques, the proposed SORT achieves 91.8 HOTA, 92.6 MOTA and 93.9 IDF1, which surpass all state-of-the-art trackers on the public CityFlow and our private HighwayTrack datasets under the low-frame-rate setting. Yawen Huang, Yubei Lin, Ziwei Zhu 0005, Xingming Zhang 0001, Yang Liu 0182, Yuexiang Li, Yefeng Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | DGFamba: Learning Flow Factorized State Space for Visual Domain GeneralizationabstractDomain generalization aims to learn a representation from the source domain, which can be generalized to arbitrary unseen target domains. A fundamental challenge for visual domain generalization is the domain gap caused by the dramatic style variation whereas the image content is stable. The realm of selective state space, exemplified by VMamba, demonstrates its global receptive field in representing the content. However, the way exploiting the domain-invariant property for selective state space is rarely explored. In this paper, we propose a novel Flow Factorized State Space model, dubbed as DGFamba, for visual domain generalization. To maintain domain consistency, we innovatively map the style-augmented and the original state embeddings by flow factorization. In this latent flow space, each state embedding from a certain style is specified by a latent probability path. By aligning these probability paths in the latent space, the state embeddings are able to represent the same content distribution regardless of the style differences. Extensive experiments conducted on various visual domain generalization settings show its state-of-the-art performance. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li |
AAAI | 6 |
| 2025 | NightAdapter: Learning a Frequency Adapter for Generalizable Night-time Scene SegmentationabstractNight-time scene segmentation is a critical yet challenging task in the real-world applications, primarily due to the complicated lighting conditions. However, existing methods lack sufficient generalization ability to unseen nighttime scenes with varying illumination. In light of this issue, we focus on investigating generalizable paradigms for night-time scene segmentation and propose an efficient fine-tuning scheme, dubbed NightAdapter, alleviating the domain gap across various scenes. Interestingly, different properties embedded in the day-time and night-time features can be characterized by the bands after discrete sine transform, which can be categorized into illumination-sensitive/-insensitive bands. Hence, our NightAdapter is powered by two appealing designs: (1) Illumination-Insensitive Band Adaptation that provides a foundation for understanding the prior, enhancing the robustness to illumination shifts; (2) Illumination-Sensitive Band Adaptation that fine-tunes the randomized frequency bands, mitigating the domain gap between the day-time and various night-time scenes. As a consequence, illumination-insensitive enhancement improves the domain invariance, while illumination-sensitive diminution strengthens the domain shift between different scenes. NightAdapter yields significant improvements over the state-of-the-art methods under various day-to-night, night-to-night, and in-domain night segmentation experiments. Source code is available at https://github.com/BiQiWHU/NightAdapter. Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
CVPR | 6 |
| 2025 | Enhancing Federated Domain Adaptation via Multi-Granular Fine-Grained AlignmentabstractTraditional unsupervised multi-source domain adaptation usually assumes that all source domain data can be utilized during training. Unfortunately, due to practical concerns such as privacy, data storage, and computational costs, data from different source domains are often isolated from each other. To address this issue, we propose a federated domain adaptation framework based on fine-grained alignment. This method achieves domain adaptation at the model level through iterative training of source and target domains, thereby avoiding the direct use of source domain data. Specifically, our approach employs specialized techniques at various stages—model construction, pseudo-label generation, and model training—to handle fine-grained features that are often overlooked. This enables the model to effectively remove irrelevant information and learn more discriminative features, thus narrowing the distribution gap between domains. Extensive experimental results demonstrate the effectiveness of our proposed method across multiple datasets. Ziyun Cai, Shangshang Song, Jie Song 0014, Yawen Huang, Changhui Hu 0001, Xiaoyuan Jing |
ICASSP | 4 |
| 2025 | A Simple Yet Mighty Hartley Diffusion Versatilist for Generalizable Dense Vision Tasks
Qi Bi, Jingjun Yi, Huimin Huang 0002, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICCV | 7 |
| 2025 | GaussianReg: Rapid 2D/3D Registration for Emergency Surgery Via Explicit 3D Modeling with Gaussian Primitives
Weihao Yu 0004, Xiaoqing Guo, Xinyu Liu 0001, Yifan Liu 0010, Hao Zheng 0008, Yawen Huang, Yixuan Yuan |
ICCV | 6 |
| 2025 | Hierarchical Recovery of Convolutional Neural Networks via Self-embedding Watermarking
Yawen Huang, Huaicong Zhang |
ICICS (2) | 1 |
| 2025 | Source-Free Domain Adaptation via Transformer-based Object-centric PerceptionabstractIn this paper, we investigate the Source-Free Domain Adaptation (SFDA), where a well-trained model adapts to an unlabeled target domain without access to source data. Previous SFDA methods mainly relied on convolutional neural networks, which struggle with domain shifts due to their local focus. To address this, we propose the Object-centric Perception Source-Free Transformer (OP-SFT), which leverages the self-attention mechanism of Transformers to focus on relevant target regions, improving adaptability to domain shifts. We also introduce self-supervised knowledge distillation to enhance semantic perception and a confidence-based k-means clustering method for more accurate pseudo-label generation. Extensive experiments demonstrate that our OP-SFT achieves significant adaptation performance across four widely-used domain adaptation benchmark datasets compared to other state-of-the-art baselines. The code is available at https://github.com/Weilong-Gao/OP-SFT. Ziyun Cai, Weilong Gao, Yawen Huang, Jie Song 0014, Changhui Hu 0001, Tengfei Zhang 0001 |
ICME | 3 |
| 2025 | Make Multi-source Task Greater Again: Adaptive Causal Diffusion StrategyabstractMulti-source Domain Adaptation (MSDA) aims to adapt models trained on multiple labeled source domains to an unlabeled target domain. Recent MSDA methods based on Generative Adversarial Networks (GANs) implicitly capture the image distribution, which can lead to limited sample fidelity and result in misalignment of pixel-level information between the sources and the target domain. Moreover, when samples from different sources interact during training, significant misalignment across various source domains can occur. In this study, we introduce a novel MSDA framework called Adaptive Causal Diffusion Networks (ACDN) to address these challenges. ACDN integrates a diffusive domain adaptation model for effective, high-fidelity adaptation between the source and target domains, incorporating Granger-causal inference to ensure that the assigned weights for each source domain are closely related to their respective contributions to the decision-making process. Experimental results show that ACDN outperforms existing methods significantly across real-world domain adaptation benchmarks. Ziyun Cai, Yawen Huang, Jie Song 0014, Changhui Hu 0001, Tengfei Zhang 0001 |
ICME | 2 |
| 2025 | Multi-Scale Tubularity-Aware U-NetabstractU-Net architectures have made great progress in dealing with semantic segmentation tasks. However, existing frameworks have not yet possessed the ability of capturing sufficient local and contextual dependencies of tubular structures. The reasons are two-fold. First, traditional square convolutions are inherently limited to model irregular pixel changes due to their fixed geometric structures. Second, there exist semantic gaps among the multi-scale tubularity features and between stages of their encoding and the decoding. To mitigate these issues, we propose multi-scale tubularity-aware U-Net, by coupling a novel tubularity deformable convolution (TdConv) embedding and a dual attention Transformer (DaTrans) alternative to skip connection. On the one hand, TdConv embedding iteratively learns the deformation offsets of convolution itself in both directions along the tubular structure. On the other hand, DaTrans connection endows skip connections with attention mechanism from both multi-scale local pixel and cross-scale global semantic perspectives. Hinging on the local irregularity perception and the global semantic association, our method enables to analyze tubular structures appeared in complex contexts and at different scales. Extensive experiments show that our approach outperforms state-of-the-art techniques, including different U-Net variants, for various datasets on several tasks including road extraction and vessel segmentation. Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang |
ICME | 6 |
| 2025 | TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
Minye Shao, Xingyu Miao, Haoran Duan 0001, Zeyu Wang 0009, Jingkun Chen, Yawen Huang, Xian Wu 0001, Jingjing Deng 0001, Yang Long 0001, Yefeng Zheng 0001 |
MICCAI (4) | 6 |
| 2025 | BrainSegDMIF: A Dynamic Fusion-enhanced SAM for Brain Lesion Segmentation
Hongming Wang, Huimin Huang 0002, Jiaxuan Jiang 0001, Hao Zheng 0008, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001, Jinping Xu |
ACM Multimedia | 8 |
| 2025 | AtlantisGS: Underwater Sparse-View Scene Reconstruction via Gaussian Splatting
Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Haolan Zhan, Yixian Shen, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
ACM Multimedia | 8 |
| 2025 | D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide ImagesabstractDiffusion-based virtual staining methods of histopathology images have demonstrated outstanding potential for stain normalization and cross-dye staining (e.g., hematoxylin-eosin to immunohistochemistry). However, achieving pathology-correct cross-dye virtual staining with versatile tone controls poses significant challenges due to the difficulty of decoupling the given pathology and tone conditions. This issue would cause non-pathologic regions to be mistakenly stained like pathologic ones, and vice versa, which we term “pathology leakage.” To address this issue, we propose diffusion virtual staining Transformer (D-VST), a new framework with versatile tone control for cross-dye virtual staining. Specifically, we introduce a pathology encoder in conjunction with a tone encoder, combined with a two-stage curriculum learning scheme that decouples pathology and tone conditions, to enable tone control while eliminating pathology leakage. Further, to extend our method for billion-pixel whole slide image (WSI) staining, we introduce a novel frequency-aware adaptive patch sampling strategy for high-quality yet efficient inference of ultra-high resolution images in a zero-shot manner. Integrating these two innovative components facilitates a pathology-correct, tone-controllable, cross-dye WSI virtual staining process. Extensive experiments on three virtual staining tasks that involve translating between four different dyes demonstrate the superiority of our approach in generating high-quality and pathologically accurate images compared to existing methods based on generative adversarial networks and diffusion models. Our code and trained models will be released. Shurong Yang, Dong Wei 0004, Yihuang Hu, Qiong Peng, Yawen Huang, Xian Wu 0001, Yefeng Zheng 0001, Liansheng Wang 0002 |
NeurIPS | 6 |
| 2025 | Degradation-Aware Dynamic Schrödinger Bridge for Unpaired Image RestorationabstractImage restoration is a fundamental task in computer vision and machine learning, which learns a mapping between the clear images and the degraded images under various conditions (e.g., blur, low-light, haze).
Yet, most existing image restoration methods are highly restricted by the requirement of degraded and clear image pairs, which limits the generalization and feasibility to enormous real-world scenarios without paired images.
To address this bottleneck, we propose a Degradation-aware Dynamic Schr\"{o}dinger Bridge (DDSB) for unpaired image restoration.
Its general idea is to learn a Schr\"{o}dinger Bridge between clear and degraded image distribution,
while at the same time emphasizing the physical degradation priors to reduce the accumulation of errors during the restoration process.
A Degradation-aware Optimal Transport (DOT) learning scheme is accordingly devised.
Training a degradation model to learn the inverse restoration process is particularly challenging, as it must be applicable across different stages of the iterative restoration process.
A Dynamic Transport with Consistency (DTC) learning objective is further proposed to reduce the loss of image details in the early iterations and therefore refine the degradation model.
Extensive experiments on multiple image degradation tasks show its state-of-the-art performance over the prior arts. Jingjun Yi, Qi Bi, Hao Zheng 0008, Huimin Huang 0002, Yixian Shen, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 8 |
| 2025 | Learning a Cross-Modal Schrödinger Bridge for Visual Domain GeneralizationabstractDomain generalization aims to train models that perform robustly on unseen target domains without access to target data.
The realm of vision-language foundation model has opened a new venue owing to its inherent out-of-distribution generalization capability.
However, the static alignment to class-level textual anchors remains insufficient to handle the dramatic distribution discrepancy from diverse domain-specific visual features.
In this work, we propose a novel cross-domain Schrödinger Bridge (SB) method, namely SBGen, to handle this challenge, which explicitly formulates the stochastic semantic evolution, to gain better generalization to unseen domains.
Technically, the proposed \texttt{SBGen} consists of three key components: (1) \emph{text-guided domain-aware feature selection} to isolate semantically aligned image tokens; (2) \emph{stochastic cross-domain evolution} to simulate the SB dynamics via a learnable time-conditioned drift; and (3) \emph{stochastic domain-agnostic interpolation} to construct semantically grounded feature trajectories.
Empirically, \texttt{SBGen} achieves state-of-the-art performance on domain generalization in both classification and segmentation. This work highlights the importance of modeling domain shifts as structured stochastic processes grounded in semantic alignment. Hao Zheng 0008, Jingjun Yi, Qi Bi, Huimin Huang 0002, Haolan Zhan, Yawen Huang, Yuexiang Li, Xian Wu 0001, Yefeng Zheng 0001 |
NeurIPS | 6 |
| 2025 | Adaptive margin for unsupervised domain adaptation without source data
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Changhui Hu 0001, Xiaoyuan Jing |
Comput. Vis. Image Underst. | 2 |
| 2025 | Multi-Source Domain Adaptation by Causal-Guided Adaptive Multimodal Diffusion Networks
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Yefeng Zheng 0001, Dong Yue 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions
Yawen Huang, Huimin Huang 0002, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 1 |
| 2025 | D3T: Dual-Domain Diffusion Transformer in Triplanar Latent Space for 3D Incomplete-View CT Reconstruction
Xuhui Liu, Hong Li 0016, Yawen Huang, Xiantong Zhen, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 4 |
| 2025 | CLIMS++: Cross Language Image Matching with Automatic Context Discovery for Weakly Supervised Semantic Segmentation
Jinheng Xie, Songhe Deng, Xianxu Hou, Zhaochuan Luo, LinLin Shen, Yawen Huang, Yefeng Zheng 0001, Zheng Shou 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | A Recovery-Mechanism-Driven Wireless Group Key Generation Protocol for Multiuser ScenariosabstractPhysical-layer key generation (PKG) leveraging the reciprocity of wireless channel provides an effective approach for key agreement among resource-constrained Internet of Things devices. However, current researches on PKG predominantly focus on pairwise communication scenarios, and there remain challenges in achieving group key generation for multiuser scenarios. In this article, we propose a novel recovery mechanism-driven wireless group key generation protocol to facilitate key sharing in the star network typology. Specifically, the root node will assign each member node its unique group key component before initiating group key distribution. Subsequently, all group key components are distributed to member nodes using a forward error correction mechanism, which helps reduce system overhead. Finally, all member nodes utilize a recovery mechanism and their respective group key component to obtain the same complete group key, thereby achieving group key distribution. Compared to existing schemes, our protocol can avoid the significant information leakage caused by repeated distribution of the same group key, thereby enhancing security. We further design and implement a practical wireless group key generation system using ESP32. Additionally, a group channel state information (CSI) extraction tool for multiuser channel measurements is developed. Experimental results demonstrate that our protocol can generate the group key with high randomness while benefiting from good channel reciprocity, making it suitable for cryptographic applications in multiuser communication scenarios. Huaicong Zhang, Yawen Huang, Jiabao Yu, Boqian Liu, Aiqun Hu |
IEEE Internet Things J. | 2 |
| 2025 | Learning Generalized Medical Image Representation by Decoupled Feature QueriesabstractMedical images are usually collected from multiple clinical centers with various types of scanners. When confronted with such significant cross-domain distribution discrepancy, a deep network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. Such channel redundancy limits the expressive capability of a representation, resulting in less preferable generalization ability. To address this fundamental yet challenging issue, we propose a novel decoupled feature as query (DFQ) framework for domain generalized medical image representation learning. Its general idea is to leverage the channel-wise decoupled deep features as queries. Particularly, a deep instance whitening transform with restricted isometry is proposed, which enforces each channel orthogonal to the rest channels after decoupling. Besides, the long-range dependency between decoupled deep and shallow features is implicitly constrained to minimize channel redundancy throughout training. Extensive experiments show its state-of-the-art performance on three medical domain generalization tasks with four modalities. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | GAD: Domain generalized diabetic retinopathy grading by grade-aware de-stylization
Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
Pattern Recognit. | 5 |
| 2025 | Attention-driven acoustic properties learning for underwater target ranging
Xiaohui Chu, Hantao Zhou, Yan Zhang 0109, Yachao Zhang 0001, Runze Hu, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001, Rongrong Ji |
Pattern Recognit. | 7 |
| 2025 | SP-SLAM: Neural Real-Time Dense SLAM With Scene PriorsabstractNeural implicit representations have recently shown promising progress in dense Simultaneous Localization And Mapping (SLAM). However, existing works have shortcomings in terms of reconstruction quality and real-time performance, mainly due to inflexible scene representation strategy without leveraging any prior information. In this paper, we introduce SP-SLAM, a novel neural RGB-D SLAM system that performs tracking and mapping in real-time. SP-SLAM computes depth images and establishes sparse voxel-encoded scene priors near the surface reconstruction. Simultaneously, we employ triplanes to store scene appearance information, striking a balance between achieving high-quality geometric texture mapping and minimizing memory consumption. Furthermore, in SP-SLAM, we introduce an effective optimization strategy for mapping, allowing the system to continuously optimize the poses of all historical input frames during runtime without increasing computational overhead. We conduct extensive evaluations on five benchmark datasets (Replica, ScanNet, TUM RGB-D, Synthetic RGB-D, 7-Scenes). The results demonstrate that, compared to existing methods, we achieve superior tracking accuracy and reconstruction quality, while running at a significantly faster speed. Zhen Hong, Haoran Duan 0001, Yawen Huang, Zhenyu Wen, Xiang Wu 0012, Wei Xiang 0001, Yefeng Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DUSA-UNet: Dual Sparse Attentive U-Net for Multiscale Road Network ExtractionabstractThe challenges of road network segmentation demand an algorithm capable of adapting to the sparse and irregular shapes, as well as the diverse context, which often leads traditional encoding-decoding methods and simple Transformer embeddings to failure. We introduce a computationally efficient and powerful framework for elegant road-aware segmentation. Our method, called DUSA-UNet, effectively encodes fine-grained local road connectivity and holistic global topological semantics while decoding multiscale road network information. DUSA-UNet offers a novel alternative to the U-Net architecture by integrating connectivity attention, which can exploit intra-road interactions across multi-level sampling features with reduced computational complexity. This local interaction serves as valuable prior information for learning global interactions between road networks and the background through another integrality attention mechanism. The two forms of sparse attention are arranged alternatively and complementarily, and trained jointly, resulting in performance improvements without significant increases in computational complexity. Extensive experiments on various datasets with different resolutions, including Massachusetts, DeepGlobe, SpaceNet, and Large-Scale remote sensing images, demonstrate that DUSA-UNet outperforms state-of-the-art techniques. Our approach represents a significant advancement in the field of road network extraction, providing a computationally feasible solution that achieves high-quality segmentation results. Jie Song 0014, Ziyun Cai, Liang Xiao 0001, Yawen Huang, Yefeng Zheng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Rethinking Brain Tumor Segmentation From the Frequency Domain PerspectiveabstractPrecise segmentation of brain tumors, particularly contrast-enhancing regions visible in post-contrast MRI (areas highlighted by contrast agent injection), is crucial for accurate clinical diagnosis and treatment planning but remains challenging. However, current methods exhibit notable performance degradation in segmenting these enhancing brain tumor areas, largely due to insufficient consideration of MRI-specific tumor features such as complex textures and directional variations. To address this, we propose the Harmonized Frequency Fusion Network (HFF-Net), which rethinks brain tumor segmentation from a frequency-domain perspective. To comprehensively characterize tumor regions, we develop a Frequency Domain Decomposition (FDD) module that separates MRI images into low-frequency components, capturing smooth tumor contours and high-frequency components, highlighting detailed textures and directional edges. To further enhance sensitivity to tumor boundaries, we introduce an Adaptive Laplacian Convolution (ALC) module that adaptively emphasizes critical high-frequency details using dynamically updated convolution kernels. To effectively fuse tumor features across multiple scales, we design a Frequency Domain Cross-Attention (FDCA) integrating semantic, positional, and slice-specific information. We further validate and interpret frequency-domain improvements through visualization, theoretical reasoning, and experimental analyses. Extensive experiments on four public datasets demonstrate that HFF-Net achieves an average relative improvement of 4.48% (ranging from 2.39% to 7.72%) in the mean Dice scores across the three major subregions, and an average relative improvement of 7.33% (ranging from 5.96% to 8.64%) in the segmentation of contrast-enhancing tumor regions, while maintaining favorable computational efficiency and clinical applicability. Our code is available at: https://github.com/VinyehShaw/HFF. Minye Shao, Zeyu Wang 0009, Haoran Duan 0001, Yawen Huang, Bing Zhai, Shizheng Wang, Yang Long 0001, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | A Semantic-Consistent Few-Shot Modulation Recognition Framework for IoT ApplicationsabstractThe rapid growth of the Internet of Things (IoT) has led to the widespread adoption of the IoT networks in numerous digital applications. To counter physical threats in these systems, automatic modulation classification (AMC) has emerged as an effective approach for identifying the modulation format of signals in noisy environments. However, identifying those threats can be particularly challenging due to the scarcity of labeled data, which is a common issue in various IoT applications, such as anomaly detection for unmanned aerial vehicles (UAVs) and intrusion detection in the IoT networks. Few-shot learning (FSL) offers a promising solution by enabling models to grasp the concepts of new classes using only a limited number of labeled samples. However, prevalent FSL techniques are primarily tailored for tasks in the computer vision domain and are not suitable for the wireless signal domain. Instead of designing a new FSL model, this work suggests a novel approach that enhances wireless signals to be more efficiently processed by the existing state-of-the-art (SOTA) FSL models. We present the semantic-consistent signal pretransformation (ScSP), a parameterized transformation architecture that ensures signals with identical semantics exhibit similar representations. ScSP is designed to integrate seamlessly with various SOTA FSL models for signal modulation recognition and supports commonly used deep learning backbones. Our evaluation indicates that ScSP boosts the performance of numerous SOTA FSL models, while preserving flexibility. Jie Su 0001, Zhenyu Wen, Fangda Guo, Yiming Wu 0009, Zhen Hong, Haoran Duan 0001, Yawen Huang, Rajiv Ranjan 0001, Yefeng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | UniHead: Unifying Multi-Perception for Detection HeadsabstractThe detection head constitutes a pivotal component within object detectors, tasked with executing both classification and localization functions. Regrettably, the commonly used parallel head often lacks omni perceptual capabilities, such as deformation perception (DP), global perception (GP), and cross-task perception (CTP). Despite numerous methods attempting to enhance these abilities from a single aspect, achieving a comprehensive and unified solution remains a significant challenge. In response to this challenge, we develop an innovative detection head, termed UniHead, to unify three perceptual abilities simultaneously. More precisely, our approach: 1) introduces DP, enabling the model to adaptively sample object features; 2) proposes a dual-axial aggregation transformer (DAT) to adeptly model long-range dependencies, thereby achieving GP; and 3) devises a cross-task interaction transformer (CIT) that facilitates interaction between the classification and localization branches, thus aligning the two tasks. As a plug-and-play method, the proposed UniHead can be conveniently integrated with existing detectors. Extensive experiments on the COCO dataset demonstrate that our UniHead can bring significant improvements to many detectors. For instance, the UniHead can obtain +2.7 AP gains in RetinaNet, +2.9 AP gains in FreeAnchor, and +2.1 AP gains in GFL. The code is available at https://github.com/zht8506/UniHead. Hantao Zhou, Rui Yang 0040, Yachao Zhang 0001, Haoran Duan 0001, Yawen Huang, Runze Hu, Xiu Li 0001, Yefeng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Learning Generalized Medical Image Segmentation from Decoupled Feature QueriesabstractDomain generalized medical image segmentation requires models to learn from multiple source domains and generalize well to arbitrary unseen target domain. Such a task is both technically challenging and clinically practical, due to the domain shift problem (i.e., images are collected from different hospitals and scanners). Existing methods focused on either learning shape-invariant representation or reaching consensus among the source domains. An ideal generalized representation is supposed to show similar pattern responses within the same channel for cross-domain images. However, to deal with the significant distribution discrepancy, the network tends to capture similar patterns by multiple channels, while different cross-domain patterns are also allowed to rest in the same channel. To address this issue, we propose to leverage channel-wise decoupled deep features as queries. With the aid of cross-attention mechanism, the long-range dependency between deep and shallow features can be fully mined via self-attention and then guides the learning of generalized representation. Besides, a relaxed deep whitening transformation is proposed to learn channel-wise decoupled features in a feasible way. The proposed decoupled fea- ture query (DFQ) scheme can be seamlessly integrate into the Transformer segmentation model in an end-to-end manner. Extensive experiments show its state-of-the-art performance, notably outperforming the runner-up by 1.31% and 1.98% with DSC metric on generalized fundus and prostate benchmarks, respectively. Source code is available at https://github.com/BiQiWHU/DFQ. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 5 |
| 2024 | Federated Learning via Input-Output Collaborative DistillationabstractFederated learning (FL) is a machine learning paradigm in which distributed local nodes collaboratively train a central model without sharing individually held private data. Existing FL methods either iteratively share local model parameters or deploy co-distillation. However, the former is highly susceptible to private data leakage, and the latter design relies on the prerequisites of task-relevant real data. Instead, we propose a data-free FL framework based on local-to-central collaborative distillation with direct input and output space exploitation. Our design eliminates any requirement of recursive local parameter exchange or auxiliary task-relevant data to transfer knowledge, thereby giving direct privacy control to local users. In particular, to cope with the inherent data heterogeneity across locals, our technique learns to distill input on which each local model produces consensual yet unique results to represent each expertise. Our proposed FL framework achieves notable privacy-utility trade-offs with extensive experiments on image classification and segmentation tasks under various real-world heterogeneous federated learning settings on both natural and medical images. Code is available at https://github.com/lsl001006/FedIOD. Shanglin Li, Yuxiang Bao, Barry Yao, Yawen Huang, Ziyan Wu 0001, Baochang Zhang 0001, Yefeng Zheng 0001, David S. Doermann |
AAAI | 5 |
| 2024 | Combinatorial CNN-Transformer Learning with Manifold Constraints for Semi-supervised Medical Image SegmentationabstractSemi-supervised learning (SSL), as one of the dominant methods, aims at leveraging the unlabeled data to deal with the annotation dilemma of supervised learning, which has attracted much attentions in the medical image segmentation. Most of the existing approaches leverage a unitary network by convolutional neural networks (CNNs) with compulsory consistency of the predictions through small perturbations applied to inputs or models. The penalties of such a learning paradigm are that (1) CNN-based models place severe limitations on global learning; (2) rich and diverse class-level distributions are inhibited. In this paper, we present a novel CNN-Transformer learning framework in the manifold space for semi-supervised medical image segmentation. First, at intra-student level, we propose a novel class-wise consistency loss to facilitate the learning of both discriminative and compact target feature representations. Then, at inter-student level, we align the CNN and Transformer features using a prototype-based optimal transport method. Extensive experiments show that our method outperforms previous state-of-the-art methods on three public medical image segmentation benchmarks. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
AAAI | 2 |
| 2024 | Going Beyond Multi-Task Dense Prediction with Synergy Embedding ModelsabstractMulti-task visual scene understanding aims to leverage the relationships among a set of correlated tasks, which are solved simultaneously by embedding them within a unified network. However, most existing methods give rise to two primary concerns from a task-level perspective: (1) the lack of task-independent correspondences for distinct tasks, and (2) the neglect of explicit task-consensual dependencies among various tasks. To address these issues, we propose a novel synergy embedding models (SEM), which goes beyond multi-task dense prediction by leveraging two innovative designs: the intra-task hierarchy-adaptive module and the inter-task EM-interactive module. Specifically, the constructed intra-task module incorporates hierarchy-adaptive keys from multiple stages, enabling the efficient learning of specialized visual patterns with an optimal trade-off. In addition, the developed inter-task module learns interactions from a compact set of mutual bases among various tasks, benefiting from the expectation maximization (EM) algorithm. Extensive empirical evidence from two public benchmarks, NYUD-v2 and PASCAL-Context, demonstrates that SEM consistently outperforms state-of-the-art approaches across a range of metrics. Huimin Huang 0002, Yawen Huang, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hao Zheng 0008, Yuexiang Li, Yefeng Zheng 0001 |
CVPR | 2 |
| 2024 | Tune-an-Ellipse: CLIP Has Potential to Find what you WantabstractVisual prompting of large vision language models such as CLIP exhibits intriguing zero-shot capabilities. A manually drawn red circle, commonly used for highlighting, can guide CLIP's attention to the surrounding region, to identify specific objects within an image. Without precise object proposals, however, it is insufficient for localization. Our novel, simple yet effective approach, i.e., Differentiable Visual Prompting, enables CLIP to zero-shot localize: given an image and a text prompt describing an object, we first pick a rendered ellipse from uniformly distributed anchor ellipses on the image grid via visual prompting, then use three loss functions to tune the ellipse coefficients to encap-sulate the target region gradually. This yields promising ex-perimental results for referring expression comprehension without precisely specified object proposals. In addition, we systematically present the limitations of visual prompting inherent in CLIP and discuss potential solutions. Jinheng Xie, Songhe Deng, Bing Li 0024, Yawen Huang, Yefeng Zheng 0001, Jürgen Schmidhuber, Bernard Ghanem, LinLin Shen, Zheng Shou 0001 |
CVPR | 5 |
| 2024 | Self-Supervised Cross-Level Consistency Learning For Fundus Image ClassificationabstractThe rapid development of intelligent systems for eye disease diagnosis decreases the risk of people suffering from vision impairment. However, the superior discrimination ability of existing retinal disease diagnosis methods heavily relies on the large-scale high-quality annotations. In this work, we adapt the self-supervised technique for fundus image classification with the merits of bypassing the over-dependence of labeled data. Unlike most current self-supervised approaches, which only learn global pre-text representations from view-level, our method further incorporates the region-level representations into the learning process, since the pathological changes in fundus images are usually subtle and scattered. Specifically, we propose a novel self-supervised cross-level consistency learning scheme (S2C2L), which leverages both view-level and region-level representations of a vision Transformer to improve the robustness of extracted self-supervised representation. A diagnosis perception module (DPM) is constructed to enhance the activation of local pathological regions from both region and view levels, and a cross-level consistency loss is dedicated to align the representations from both levels. Extensive experiments on iChallenge-AMD, LAG and APTOS2019 datasets validate the state-of-the-art performance of our method for three common eye diseases. Qi Bi, Hao Zheng 0008, Xu Sun 0006, Jingjun Yi, Wentian Zhang, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ICASSP | 6 |
| 2024 | Dual Variational Knowledge Attention for Class Incremental Vision TransformerabstractClass incremental learning (CIL) strives to emulate the human cognitive process of continuously learning and adapting to new tasks while retaining knowledge from past experiences. Despite significant advancements in this field, Transformer-based models have not fully leveraged the potential of attention mechanisms to balance the transferable knowledge between tokens and the associated information. This paper addresses this gap by using a dual variational knowledge attention (DVKA) mechanism within a Transformer-based encoder-decoder framework, tailored for CIL. DVKA mechanism aims to manage the information flow through the attention maps, ensuring a balanced representation of all classes, and mitigating the risk of information dilution as new classes are incrementally introduced. This method, leverage the information bottleneck and mutual information principle, selectively filters less relevant information, directing the model’s focus towards the most significant details for each class. The DVKA is designed with two distinct attentions: one focused on the feature level and the other on the token dimension. The feature-focused attention aims to purify the complex nature of various classification tasks, ensuring a comprehensive representation of both old and new tasks. The token-focused attention mechanism highlights specific tokens, facilitating local discrimination among disparate patches and fostering global coordination for a spectrum of task tokens. Our work is a major stride towards improving transformer models for class incremental learning, presenting a theoretical rationale and effective experimental results on three widely-used datasets. Haoran Duan 0001, Rui Sun 0010, Varun Ojha 0001, Tejal Shah, Zhuoxu Huang, Zizhou Ouyang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001 |
IJCNN | 7 |
| 2024 | Hallucinated Style Distillation for Single Domain Generalization in Medical Image Segmentation
Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Shaoxin Li 0001, Yuexiang Li, Yefeng Zheng 0001, Feiyue Huang |
MICCAI (10) | 6 |
| 2024 | Learning Spectral-Decomposited Tokens for Domain Generalized Semantic SegmentationabstractThe rapid development of Vision Foundation Model (VFM) brings inherent out-domain generalization for a variety of down-stream tasks. Among them, domain generalized semantic segmentation (DGSS) holds unique challenges as the cross-domain images share common pixel-wise content information but vary greatly in terms of the style. In this paper, we present a novel Spectral-dEcomposed Token (SET) learning framework to advance the frontier. Delving into further than existing fine-tuning token & frozen backbone paradigm, the proposed SET especially focuses on the way learning style-invariant features from these learnable tokens. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction. Particularly, the frozen VFM features are first decomposed into the phase and amplitude components in the frequency space, which mainly contain the information of content and style, respectively, and then separately processed by learnable tokens for task-specific information extraction.After the decomposition, style variation primarily impacts the token-based feature enhancement within the amplitude branch. To address this issue, we further develop an attention optimization method to bridge the gap between style-affected representation and static tokens during inference. Extensive cross-domain experiments show its state-of-the-art performance. Jingjun Yi, Qi Bi, Hao Zheng 0008, Haolan Zhan, Wei Ji 0011, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 6 |
| 2024 | Samba: Severity-aware Recurrent Modeling for Cross-domain Medical Image GradingabstractDisease grading is a crucial task in medical image analysis. Due to the continuous progression of diseases, i.e., the variability within the same level and the similarity between adjacent stages, accurate grading is highly challenging.
Furthermore, in real-world scenarios, models trained on limited source domain datasets should also be capable of handling data from unseen target domains.
Due to the cross-domain variants, the feature distribution between source and unseen target domains can be dramatically different, leading to a substantial decrease in model performance.
To address these challenges in cross-domain disease grading, we propose a Severity-aware Recurrent Modeling (Samba) method in this paper.
As the core objective of most staging tasks is to identify the most severe lesions, which may only occupy a small portion of the image, we propose to encode image patches in a sequential and recurrent manner.
Specifically, a state space model is tailored to store and transport the severity information by hidden states.
Moreover, to mitigate the impact of cross-domain variants, an Expectation-Maximization (EM) based state recalibration mechanism is designed to map the patch embeddings into a more compact space.
We model the feature distributions of different lesions through the Gaussian Mixture Model (GMM) and reconstruct the intermediate features based on learnable severity bases.
Extensive experiments show the proposed Samba outperforms the VMamba baseline by an average accuracy of 23.5\%, 5.6\% and 4.1\% on the cross-domain grading of fatigue fracture, breast cancer and diabetic retinopathy, respectively.
Source code is available at \url{https://github.com/BiQiWHU/Samba}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Wei Ji 0011, Haolan Zhan, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 6 |
| 2024 | Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationabstractThe emerging vision foundation model (VFM) has inherited the ability to generalize to unseen images.
Nevertheless, the key challenge of domain-generalized semantic segmentation (DGSS) lies in the domain gap attributed to the cross-domain styles, i.e., the variance of urban landscape and environment dependencies.
Hence, maintaining the style-invariant property with varying domain styles becomes the key bottleneck in harnessing VFM for DGSS.
The frequency space after Haar wavelet transformation provides a feasible way to decouple the style information from the domain-invariant content, since the content and style information are retained in the low- and high- frequency components of the space, respectively.
To this end, we propose a novel Frequency-Adapted (FADA) learning scheme to advance the frontier.
Its overall idea is to separately tackle the content and style information by frequency tokens throughout the learning process.
Particularly, the proposed FADA consists of two branches, i.e., low- and high- frequency branches. The former one is able to stabilize the scene content, while the latter one learns the scene styles and eliminates its impact to DGSS.
Experiments conducted on various DGSS settings show the state-of-the-art performance of our FADA and its versatility to a variety of VFMs.
Source code is available at \url{https://github.com/BiQiWHU/FADA}. Qi Bi, Jingjun Yi, Hao Zheng 0008, Haolan Zhan, Yawen Huang, Wei Ji 0011, Yuexiang Li, Yefeng Zheng 0001 |
NeurIPS | 5 |
| 2024 | Triplet-branch network with contrastive prior-knowledge embedding for disease grading
Yuexiang Li, Yawen Huang, Jingxin Liu 0005, Yi Lin 0009, Dong Wei 0004, Qirui Zhang 0004, Kai Ma 0002, Guangming Lu 0001, Yefeng Zheng 0001 |
Artif. Intell. Medicine | 4 |
| 2024 | Multi-Constraint Transferable Generative Adversarial Networks for Cross-Modal Brain Image Synthesis
Yawen Huang, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Guo-Jun Qi, Ling Shao 0001, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Wearable-based behaviour interpolation for semi-supervised human activity recognitionabstractWhile traditional feature engineering for Human Activity Recognition (HAR) involves a trial-and-error process, deep learning has emerged as a preferred method for high-level representations of sensor-based human activities. However, most deep learning-based HAR requires a large amount of labelled data and extracting HAR features from unlabelled data for effective deep learning training remains challenging. We, therefore, introduce a deep semi-supervised HAR approach, MixHAR, which concurrently uses labelled and unlabelled activities. Our MixHAR employs a linear interpolation mechanism to blend labelled and unlabelled activities while addressing both inter- and intra-activity variability. A unique challenge identified is the activity-intrusion problem during mixing, for which we propose a mixing calibration mechanism to mitigate it in the feature embedding space. Additionally, we rigorously explored and evaluated the five conventional/popular deep semi-supervised technologies on HAR, acting as the benchmark of deep semi-supervised HAR. Our results demonstrate that MixHAR significantly improves performance, underscoring the potential of deep semi-supervised techniques in HAR. Haoran Duan 0001, Varun Ojha 0001, Shizheng Wang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001, Yefeng Zheng 0001 |
Inf. Sci. | 5 |
| 2024 | Improving vision transformer for medical image classification via token-wise perturbation
Yuexiang Li, Yawen Huang, Nanjun He, Kai Ma 0002, Yefeng Zheng 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Attention Cycle-consistent universal network for More Universal Domain Adaptation
Ziyun Cai, Yawen Huang, Tengfei Zhang 0001, Xiaoyuan Jing, Yefeng Zheng 0001, Ling Shao 0001 |
Pattern Recognit. | 2 |
| 2024 | CTNeRF: Cross-time Transformer for dynamic neural radiance field from monocular video
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001 |
Pattern Recognit. | 5 |
| 2024 | Anomaly detection via gating highway connection for retinal fundus images
Wentian Zhang, Jinheng Xie, Yawen Huang, Yu Zhang 0185, Yuexiang Li, Ramachandra Raghavendra, Yefeng Zheng 0001 |
Pattern Recognit. | 4 |
| 2024 | Affine Collaborative Normalization: A shortcut for adaptation in medical image analysis
Chuyan Zhang, Yuncheng Yang, Hao Zheng 0008, Yawen Huang, Yefeng Zheng 0001, Yun Gu |
Pattern Recognit. | 4 |
| 2024 | DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost VolumeabstractSelf-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. However, this assumption does not hold in dynamic objects in scenarios, leading to errors during the view synthesis stage, such as feature mismatch and occlusion, which can significantly reduce the accuracy of the generated depth maps. To address this problem, we propose a novel dynamic cost volume that exploits residual optical flow to describe moving objects, improving incorrectly occluded regions in static cost volumes used in previous work. Nevertheless, the dynamic cost volume inevitably generates extra occlusions and noise, thus we alleviate this by designing a fusion module that makes static and dynamic cost volumes compensate for each other. In other words, occlusion from the static volume is refined by the dynamic volume, and incorrect information from the dynamic volume is eliminated by the static volume. Furthermore, we propose a pyramid distillation loss to reduce photometric error inaccuracy at low resolutions and an adaptive photometric error loss to alleviate the flow direction of the large gradient in the occlusion regions. We conducted extensive experiments on the KITTI and Cityscapes datasets, and the results demonstrate that our model outperforms previously published baselines for self-supervised monocular depth estimation. Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long 0001, Yefeng Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Cross-Modal Vertical Federated Learning for MRI ReconstructionabstractFederated learning enables multiple hospitals to cooperatively learn a shared model without privacy disclosure. Existing methods often take a common assumption that the data from different hospitals have the same modalities. However, such a setting is difficult to fully satisfy in practical applications, since the imaging guidelines may be different between hospitals, which makes the number of individuals with the same set of modalities limited. To this end, we formulate this practical-yet-challenging cross-modal vertical federated learning task, in which data from multiple hospitals have different modalities with a small amount of multi-modality data collected from the same individuals. To tackle such a situation, we develop a novel framework, namely Federated Consistent Regularization constrained Feature Disentanglement (Fed-CRFD), for boosting MRI reconstruction by effectively exploring the overlapping samples (i.e., same patients with different modalities at different hospitals) and solving the domain shift problem caused by different modalities. Particularly, our Fed-CRFD involves an intra-client feature disentangle scheme to decouple data into modality-invariant and modality-specific features, where the modality-invariant features are leveraged to mitigate the domain shift problem. In addition, a cross-client latent representation consistency constraint is proposed specifically for the overlapping samples to further align the modality-invariant features extracted from different modalities. Hence, our method can fully exploit the multi-source data from hospitals while alleviating the domain shift problem. Extensive experiments on two typical MRI datasets demonstrate that our network clearly outperforms state-of-the-art MRI reconstruction methods. Yunlu Yan, Hong Wang 0021, Yawen Huang, Nanjun He, Lei Zhu 0003, Yong Xu 0001, Yuexiang Li, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement LearningabstractMedical image segmentation is a critical task for clinical diagnosis and research. However, dealing with highly imbalanced data remains a significant challenge in this domain, where the region of interest (ROI) may exhibit substantial variations across different slices. This presents a significant hurdle to medical image segmentation, as conventional segmentation methods may either overlook the minority class or overly emphasize the majority class, ultimately leading to a decrease in the overall generalization ability of the segmentation results. To overcome this, we propose a novel approach based on multi-step reinforcement learning, which integrates prior knowledge of medical images and pixel-wise segmentation difficulty into the reward function. Our method treats each pixel as an individual agent, utilizing diverse actions to evaluate its relevance for segmentation. To validate the effectiveness of our approach, we conduct experiments on four imbalanced medical datasets, and the results show that our approach surpasses other state-of-the-art methods in highly imbalanced scenarios. These findings hold substantial implications for clinical diagnosis and research. Feiyang Yang, Haoran Duan 0001, Feilong Xu, Yawen Huang, Xiaoli Zhang 0001, Yang Long 0001, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Prototype Correlation Matching and Class- Relation Reasoning for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation has achieved great progress in improving accuracy and efficiency of medical analysis in the biomedical imaging field. However, most existing methods cannot explore inter-class relations among base and novel medical classes to reason unseen novel classes. Moreover, the same kind of medical class has large intra-class variations brought by diverse appearances, shapes and scales, thus causing ambiguous visual characterization to degrade generalization performance of these existing methods on unseen novel classes. To address the above challenges, in this paper, we propose a Prototype correlation Matching and Class-relation Reasoning (i.e., PMCR) model. The proposed model can effectively mitigate false pixel correlation matches caused by large intra-class variations while reasoning inter-class relations among different medical classes. Specifically, in order to address false pixel correlation match brought by large intra-class variations, we propose a prototype correlation matching module to mine representative prototypes that can characterize diverse visual information of different appearances well. We aim to explore prototypelevel rather than pixel-level correlation matching between support and query features via optimal transport algorithm to tackle false matches caused by intra-class variations. Meanwhile, in order to explore inter-class relations, we design a class-relation reasoning module to segment unseen novel medical objects via reasoning inter-class relations between base and novel classes. Such inter-class relations can be well propagated to semantic encoding of local query features to improve few-shot segmentation performance. Quantitative comparisons illustrates the large performance improvement of our model over other baseline methods. Hongliu Li, Yajun Gao, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | ClassFormer: Exploring Class-Aware Dependency with Transformer for Medical Image SegmentationabstractVision Transformers have recently shown impressive performances on medical image segmentation. Despite their strong capability of modeling long-range dependencies, the current methods still give rise to two main concerns in a class-level perspective: (1) intra-class problem: the existing methods lacked in extracting class-specific correspondences of different pixels, which may lead to poor object coverage and/or boundary prediction; (2) inter-class problem: the existing methods failed to model explicit category-dependencies among various objects, which may result in inaccurate localization. In light of these two issues, we propose a novel transformer, called ClassFormer, powered by two appealing transformers, i.e., intra-class dynamic transformer and inter-class interactive transformer, to address the challenge of fully exploration on compactness and discrepancy. Technically, the intra-class dynamic transformer is first designed to decouple representations of different categories with an adaptive selection mechanism for compact learning, which optimally highlights the informative features to reflect the salient keys/values from multiple scales. We further introduce the inter-class interactive transformer to capture the category dependency among different objects, and model class tokens as the representative class centers to guide a global semantic reasoning. As a consequence, the feature consistency is ensured with the expense of intra-class penalization, while inter-class constraint strengthens the feature discriminability between different categories. Extensive empirical evidence shows that ClassFormer can be easily plugged into any architecture, and yields improvements over the state-of-the-art methods in three public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hong Wang 0021, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001 |
AAAI | 8 |
| 2023 | Combating Mode Collapse via Offline Manifold Entropy EstimationabstractGenerative Adversarial Networks (GANs) have shown compelling results in various tasks and applications in recent years. However, mode collapse remains a critical problem in GANs. In this paper, we propose a novel training pipeline to address the mode collapse issue of GANs. Different from existing methods, we propose to generalize the discriminator as feature embedding and maximize the entropy of distributions in the embedding space learned by the discriminator. Specifically, two regularization terms, i.e., Deep Local Linear Embedding (DLLE) and Deep Isometric feature Mapping (DIsoMap), are introduced to encourage the discriminator to learn the structural information embedded in the data, such that the embedding space learned by the discriminator can be well-formed. Based on the well-learned embedding space supported by the discriminator, a non-parametric entropy estimator is designed to efficiently maximize the entropy of embedding vectors, playing as an approximation of maximizing the entropy of the generated distribution. By improving the discriminator and maximizing the distance of the most similar samples in the embedding space, our pipeline effectively reduces the mode collapse without sacrificing the quality of generated samples. Extensive experimental results show the effectiveness of our method which outperforms the GAN baseline, MaF-GAN on CelebA (9.13 vs. 12.43 in FID) and surpasses the recent state-of-the-art energy-based model on the ANIMEFACE dataset (2.80 vs. 2.26 in Inception score). Bing Li 0024, Haoqian Wu, Hanbang Liang, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001 |
AAAI | 5 |
| 2023 | SemiCVT: Semi-Supervised Convolutional Vision Transformer for Semantic SegmentationabstractSemi-supervised learning improves data efficiency of deep models by leveraging unlabeled samples to alleviate the reliance on a large set of labeled samples. These successes concentrate on the pixel-wise consistency by using convolutional neural networks (CNNs) but fail to address both global learning capability and class-level features for unlabeled data. Recent works raise a new trend that Transformer achieves superior performance on the entire feature map in various tasks. In this paper, we unify the current dominant Mean-Teacher approaches by reconciling intra-model and inter-model properties for semi-supervised segmentation to produce a novel algorithm, SemiCVT, that absorbs the quintessence of CNNs and Transformer in a comprehensive way. Specifically, we first design a parallel CNN-Transformer architecture (CVT) with introducing an intra-model local-global interaction schema (LGI) in Fourier domain for full integration. The inter-model class-wise consistency is further presented to complement the class-level statistics of CNNs and Transformer in a cross-teaching manner. Extensive empirical evidence shows that SemiCVT yields consistent improvements over the state-of-the-art methods in two public benchmarks. Huimin Huang 0002, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Hong Wang 0021, Yawen Huang, Yefeng Zheng 0001 |
CVPR | 8 |
| 2023 | AdaptiveMix: Improving GAN Training via Feature Space ShrinkageabstractDue to the outstanding capability for data generation, Generative Adversarial Networks (GANs) have attracted considerable attention in unsupervised learning. However, training GANs is difficult, since the training distribution is dynamic for the discriminator, leading to unstable image representation. In this paper, we address the problem of training GANs from a novel perspective, i.e., robust image classification. Motivated by studies on robust image representation, we propose a simple yet effective module, namely AdaptiveMix, for GANs, which shrinks the regions of training data in the image representation space of the discriminator. Considering it is intractable to directly bound feature space, we propose to construct hard samples and narrow down the feature distance between hard and easy samples. The hard samples are constructed by mixing a pair of training images. We evaluate the effectiveness of our AdaptiveMix with widely-used and state-of-the-art GAN architectures. The evaluation results demonstrate that our AdaptiveMix can facilitate the training of GANs and effectively improve the image quality of generated samples. We also show that our AdaptiveMix can be further applied to image classification and Out-Of-Distribution (OOD) detection tasks, by equipping it with state-of-the-art methods. Extensive experiments on seven publicly available datasets show that our method effectively boosts the performance of baselines. The code is publicly available at https://github.com/WentianZhang-ML/AdaptiveMix. Wentian Zhang, Bing Li 0024, Haoqian Wu, Nanjun He, Yawen Huang, Yuexiang Li, Bernard Ghanem, Yefeng Zheng 0001 |
CVPR | 6 |
| 2023 | Interactive Segmentation as Gaussian Process ClassificationabstractClick-based interactive segmentation (IS) aims to extract the target objects under user interaction. For this task, most of the current deep learning (DL)-based methods mainly follow the general pipelines of semantic segmentation. Albeit achieving promising performance, they do not fully and explicitly utilize and propagate the click information, inevitably leading to unsatisfactory segmentation results, even at clicked points. Against this issue, in this paper, we propose to formulate the IS task as a Gaussian process (GP)-based pixel-wise binary classification model on each image. To solve this model, we utilize amortized variational inference to approximate the intractable GP posterior in a data-driven manner and then decouple the approximated GP posterior into double space forms for efficient sampling with linear complexity. Then, we correspondingly construct a GP classification framework, named GPCIS, which is integrated with the deep kernel learning mechanism for more flexibility. The main specificities of the proposed GPCIS lie in: 1) Under the explicit guidance of the derived GP posterior, the information contained in clicks can be finely propagated to the entire image and then boost the segmentation; 2) The accuracy of predictions at clicks has good theoretical support. These merits of GPCIS as well as its good generality and high efficiency are substantiated by comprehensive experiments on several benchmarks, as compared with representative methods both quantitatively and qualitatively. Codes will be released at https://github.com/zmhhlnz/GPCIS_CVPR2023. Hong Wang 0021, Qian Zhao 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001 |
CVPR | 5 |
| 2023 | FemtoDet: An Object Detection Baseline for Energy Versus Performance TradeoffsabstractEfficient detectors for edge devices are often optimized for parameters or speed count metrics, which remain in weak correlation with the energy of detectors. However, some vision applications of convolutional neural networks, such as always-on surveillance cameras, are critical for energy constraints. This paper aims to serve as a baseline by designing detectors to reach tradeoffs between energy and performance from two perspectives: 1) We extensively analyze various CNNs to identify low-energy architectures, including selecting activation functions, convolutions operators, and feature fusion structures on necks. These underappreciated details in past work seriously affect the energy consumption of detectors; 2) To break through the dilemmatic energy-performance problem, we propose a balanced detector driven by energy using discovered low-energy components named FemtoDet. In addition to the novel construction, we improve FemtoDet by considering convolutions and training strategy optimizations. Specifically, we develop a new instance boundary enhancement (IBE) module for convolution optimization to overcome the contradiction between the limited capacity of CNNs and detection tasks in diverse spatial representations, and propose a recursive warm-restart (RecWR) for optimizing training strategy to escape the sub-optimization of light-weight detectors by considering the data shift produced in popular augmentations. As a result, FemtoDet with only 68.77k parameters achieves a competitive score of 46.3 AP50 on PASCAL VOC and 1.11 W & 64.47 FPS on Qualcomm Snapdragon 865 CPU platforms. Extensive experiments on COCO and TJUDHD datasets indicate that the proposed method achieves competitive results in diverse scenes. Peng Tu, Guo Ai, Yuexiang Li, Yawen Huang, Yefeng Zheng 0001 |
ICCV | 5 |
| 2023 | BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained DiffusionabstractRecent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using other modalities as conditions, considerable paired data, e.g., box/mask-image pairs, and fine-tuning time are required for nurturing models. As such paired data is time-consuming and labor-intensive to acquire and restricted to a closed set, this potentially becomes the bottleneck for applications in an open world. This paper focuses on the simplest form of user-provided conditions, e.g., box or scribble. To mitigate the aforementioned problem, we propose a training-free method to control objects and contexts in the synthesized images adhering to the given spatial conditions. Specifically, three spatial constraints, i.e., Inner-Box, Outer-Box, and Corner Constraints, are designed and seamlessly integrated into the denoising step of diffusion models, requiring no additional training and massive annotated layout data. Extensive experimental results demonstrate that the proposed constraints can control what and where to present in the images while retaining the ability of Diffusion models to synthesize with high fidelity and diverse concept coverage. Jinheng Xie, Yuexiang Li, Yawen Huang, Wentian Zhang, Yefeng Zheng 0001, Zheng Shou 0001 |
ICCV | 3 |
| 2023 | Semi-Supervised Convolutional Vision Transformer with Bi-Level Uncertainty Estimation for Medical Image SegmentationabstractSemi-supervised learning (SSL) has attracted much attention in the field of medical image segmentation, which enables to alleviate the heavy burden of labelling pixel-wise annotation by extracting knowledge from unlabeled data. The existing methods basically benefit from the success of convolutional neural networks (CNNs) by keeping consistency of the predictions under small perturbations imposed on the networks or inputs. Two main concerns arise when learning such a paradigm: (1) CNNs tend to retain discriminative local features, neglecting global dependency and thus leading to inaccurate localization; (2) CNNs omit reliable feature-level and pixel-level information, resulting in sketchy pseudo-labels, especially around the confusing boundary. In this paper, we revisit the model of semi-supervised learning and develop a novel CNN-Transformer learning framework that allows for effective segmentation of medical images by producing complementary and reliable features and pseudo-label with bi-level uncertainty. Motivated by the uncertainty estimation to gain insight on feature discrimination, we explore the statistical and geometrical properties of features on network optimization and thus launching an alignment method in a more accurate and stable way. We attach equal significance to pixel-level uncertainty estimation for alleviating the influence of unreliable pseudo-labels in the training progress and advocating the reliability of predictions. Experimental results show that our method significantly surpasses existing semi-supervised approaches on two public medical image segmentation datasets. Huimin Huang 0002, Yawen Huang, Shiao Xie, Lanfen Lin, Ruofeng Tong 0001, Yen-Wei Chen 0001, Yuexiang Li, Yefeng Zheng 0001 |
ACM Multimedia | 2 |
| 2023 | Dynamically Masked Discriminator for GANsabstractTraining Generative Adversarial Networks (GANs) remains a challenging problem. The discriminator trains the generator by learning the distribution of real/generated data. However, the distribution of generated data changes throughout the training process, which is difficult for the discriminator to learn. In this paper, we propose a novel method for GANs from the viewpoint of online continual learning. We observe that the discriminator model, trained on historically generated data, often slows down its adaptation to the changes in the new arrival generated data, which accordingly decreases the quality of generated results. By treating the generated data in training as a stream, we propose to detect whether the discriminator slows down the learning of new knowledge in generated data. Therefore, we can explicitly enforce the discriminator to learn new knowledge fast. Particularly, we propose a new discriminator, which automatically detects its retardation and then dynamically masks its features, such that the discriminator can adaptively learn the temporally-vary distribution of generated data. Experimental results show our method outperforms the state-of-the-art approaches. Wentian Zhang, Bing Li 0024, Jinheng Xie, Yawen Huang, Yuexiang Li, Yefeng Zheng 0001, Bernard Ghanem |
NeurIPS | 5 |
| 2023 | Convolutional neural networks tamper detection and location based on fragile watermarking
Yawen Huang, Hongying Zheng, Di Xiao 0001 |
Appl. Intell. | 1 |
| 2023 | FedMed-GAN: Federated domain translation on unsupervised cross-modality brain image synthesis
Jinbao Wang 0001, Guoyang Xie, Yawen Huang, Jiayi Lyu, Feng Zheng 0001, Yefeng Zheng 0001, Yaochu Jin |
Neurocomputing | 3 |
| 2023 | MIL-ViT: A multiple instance vision transformer for fundus image classification
Qi Bi, Xu Sun 0006, Kai Ma 0002, Cheng Bian, Munan Ning, Nanjun He, Yawen Huang, Yuexiang Li, Hanruo Liu, Yefeng Zheng 0001 |
J. Vis. Commun. Image Represent. | 8 |
| 2023 | A deep weakly semi-supervised framework for endoscopic lesion segmentation
Hong Wang 0021, Haoqin Ji, Yuexiang Li, Nanjun He, Dong Wei 0004, Yawen Huang, Xinrong Chen, Yefeng Zheng 0001, Hongmeng Yu |
Medical Image Anal. | 8 |
| 2023 | Learning from multiple annotators for medical image segmentationabstractSupervised machine learning methods have been widely developed for segmentation tasks in recent years. However, the quality of labels has high impact on the predictive performance of these algorithms. This issue is particularly acute in the medical image domain, where both the cost of annotation and the inter-observer variability are high. Different human experts contribute estimates of the "actual" segmentation labels in a typical label acquisition process, influenced by their personal biases and competency levels. The performance of automatic segmentation algorithms is limited when these noisy labels are used as the expert consensus label. In this work, we use two coupled CNNs to jointly learn, from purely noisy observations alone, the reliability of individual annotators and the expert consensus label distributions. The separation of the two is achieved by maximally describing the annotator's "unreliable behavior" (we call it "maximally unreliable") while achieving high fidelity with the noisy training data. We first create a toy segmentation dataset using MNIST and investigate the properties of the proposed algorithm. We then use three public medical imaging segmentation datasets to demonstrate our method's efficacy, including both simulated (where necessary) and real-world annotations: 1) ISBI2015 (multiple-sclerosis lesions); 2) BraTS (brain tumors); 3) LIDC-IDRI (lung abnormalities). Finally, we create a real-world multiple sclerosis lesion dataset (QSMSC at UCL: Queen Square Multiple Sclerosis Center at UCL, UK) with manual segmentations from 4 different annotators (3 radiologists with different level skills and 1 expert to generate the expert consensus label). In all datasets, our method consistently outperforms competing methods and relevant baselines, especially when the number of annotations is small and the amount of disagreement is large. The studies also reveal that the system is capable of capturing the complicated spatial characteristics of annotators' mistakes. Le Zhang 0005, Ryutaro Tanno, Moucheng Xu, Yawen Huang, Kevin Bronik, Joseph Jacob, Yefeng Zheng 0001, Ling Shao 0001, Olga Ciccarelli, Frederik Barkhof, Daniel C. Alexander |
Pattern Recognit. | 4 |
| 2023 | Blind Super-Resolution of 3D MRI via Unsupervised Domain TransformationabstractHigh-resolution medical images can be effectively used for clinical diagnosis. However, the acquisition of high-resolution images is difficult and often limited by medical instruments. Super-resolution (SR) methods provide a solution, where high-resolution (HR) images can be reconstructed from low-resolution (LR) ones. Most of existing deep neural networks for 3D SR medical images trained in a non-blind process, where LR images are directly degraded from HR data via a pre-determined downscale method. Such approaches rely heavily on the assumed degradation model, resulting in inevitable deviations in real clinical practice. Blind super-resolution, as a more attractive research line for this field, aims to generate HR images from LR inputs containing unknown degradation. Towards generalizing SR models for diverse types of degradation, we propose a robust blind SR of 3D medical images in an unsupervised manner with domain correction and upscaling treatment. First, a CycleGAN-based architecture is implemented to generate the LR data from the source domain to the target one for domain correction. Then, an upscaling network is learned via pre-determined HR-LR couples for reconstruction. The proposed framework is able to automatically learn noisy and blurry correction kernels for unpaired 3D SR magnetic resonance images (MRI). Our method achieves better and more robust performances in reconstruction of HR images from LR MRI with multiple unknown degradation processes, and show its superiority to other state-of-the-art supervised models and cycle-consistency based methods, especially in severe distortion cases. Hexiang Zhou, Yawen Huang, Yuexiang Li, Yi Zhou 0007, Yefeng Zheng 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | GuidedMix-Net: Semi-supervised Semantic Segmentation by Using Labeled Images as ReferenceabstractSemi-supervised learning is a challenging problem which aims to construct a model by learning from limited labeled examples. Numerous methods for this task focus on utilizing the predictions of unlabeled instances consistency alone to regularize networks. However, treating labeled and unlabeled data separately often leads to the discarding of mass prior knowledge learned from the labeled examples. In this paper, we propose a novel method for semi-supervised semantic segmentation named GuidedMix-Net, by leveraging labeled information to guide the learning of unlabeled instances. Specifically, GuidedMix-Net employs three operations: 1) interpolation of similar labeled-unlabeled image pairs; 2) transfer of mutual information; 3) generalization of pseudo masks. It enables segmentation models can learning the higher-quality pseudo masks of unlabeled data by transfer the knowledge from labeled samples to unlabeled data. Along with supervised learning for labeled data, the prediction of unlabeled data is jointly learned with the generated pseudo masks from the mixed data. Extensive experiments on PASCAL VOC 2012, and Cityscapes demonstrate the effectiveness of our GuidedMix-Net, which achieves competitive segmentation accuracy and significantly improves the mIoU over 7$\%$ compared to previous approaches. Peng Tu, Yawen Huang, Feng Zheng 0001, Zhenyu He 0001, Liujuan Cao, Ling Shao 0001 |
AAAI | 2 |
| 2022 | Generalized Brain Image Synthesis with Transferable Convolutional Sparse Coding Networks
Yawen Huang, Feng Zheng 0001, Xu Sun 0006, Yuexiang Li, Ling Shao 0001, Yefeng Zheng 0001 |
ECCV (34) | 1 |
| 2022 | Point Beyond Class: A Benchmark for Weakly Semi-supervised Abnormality Localization in Chest X-Rays
Haoqin Ji, Yuexiang Li, Jinheng Xie, Nanjun He, Yawen Huang, Dong Wei 0004, Xinrong Chen, LinLin Shen, Yefeng Zheng 0001 |
MICCAI (3) | 6 |
| 2022 | Orientation-Shared Convolution Representation for CT Metal Artifact Learning
Hong Wang 0021, Qi Xie 0002, Yuexiang Li, Yawen Huang, Deyu Meng, Yefeng Zheng 0001 |
MICCAI (6) | 4 |
| 2022 | mmFormer: Multimodal Medical Transformer for Incomplete Multimodal Learning of Brain Tumor Segmentation
Yao Zhang 0010, Nanjun He, Jiawei Yang 0002, Yuexiang Li, Dong Wei 0004, Yawen Huang, Yang Zhang 0002, Zhiqiang He 0002, Yefeng Zheng 0001 |
MICCAI (5) | 6 |
| 2022 | FedMed-ATL: Misaligned Unpaired Cross-Modality Neuroimage Synthesis via Affine Transform LossabstractThe existence of completely aligned and paired multi-modal neuroimaging data has proved its effectiveness in the diagnosis of brain diseases. However, collecting the full set of well-aligned and paired data is impractical, since the practical difficulties may include high cost, long time acquisition, image corruption, and privacy issues. Previously, the misaligned unpaired neuroimaging data (termed as MUD) are generally treated as noisy labels. However, such a noisy label-based method fails to accomplish well when misaligned data occurs distortions severely. For example, the angle of rotation is different. In this paper, we propose a novel federated self-supervised learning (FedMed) for brain image synthesis. An affine transform loss (ATL) was formulated to make use of severely distorted images without violating privacy legislation for the hospital. We then introduce a new data augmentation procedure for self-supervised training and fed it into three auxiliary heads, namely auxiliary rotation, auxiliary translation, and auxiliary scaling heads. The proposed method demonstrates the advanced performance in both the quality of our synthesized results under a severely misaligned and unpaired data setting, and better stability than other GAN-based algorithms. The proposed method also reduces the demand for deformable registration while encouraging to leverage the misaligned and unpaired data. Experimental results verify the outstanding performance of our learning paradigm compared to other state-of-the-art approaches. Jinbao Wang 0001, Guoyang Xie, Yawen Huang, Yefeng Zheng 0001, Yaochu Jin, Feng Zheng 0001 |
ACM Multimedia | 3 |
| 2022 | Compressive Sensing-Based Image Encryption and Authentication in Edge-Clouds
Hongying Zheng, Yawen Huang, Di Xiao 0001 |
MMM (2) | 2 |
| 2021 | Brain Image Synthesis With Unsupervised Multivariate Canonical CSCl4NetabstractRecent advances in neuroscience have highlighted the effectiveness of multi-modal medical data for investigating certain pathologies and understanding human cognition. However, obtaining full sets of different modalities is limited by various factors, such as long acquisition times, high examination costs and artifact suppression. In addition, the complexity, high dimensionality and heterogeneity of neuroimaging data remains another key challenge in leveraging existing randomized scans effectively, as data of the same modality is often measured differently by different machines. There is a clear need to go beyond the traditional imaging-dependent process and synthesize anatomically specific target-modality data from a source in-put. In this paper, we propose to learn dedicated features that cross both intre- and intra-modal variations using a novel CSCℓ4Net. Through an initial unification of intra-modal data in the feature maps and multivariate canonical adaptation, CSC ℓ4Net facilitates feature-level mutual transformation. The positive definite Riemannian manifold-penalized data fidelity term further enables CSCℓ4Net to re-construct missing measurements according to transformed features. Finally, the maximization ℓ4-norm boils down to a computationally efficient optimization problem. Extensive experiments validate the ability and robustness of our CSC ℓ4Net compared to the state-of-the-art methods on multiple datasets. Yawen Huang, Feng Zheng 0001, Matthew R. Scott, Ling Shao 0001 |
CVPR | 1 |
| 2020 | Super-Resolution and Inpainting with Degraded and Upgraded Generative Adversarial NetworksabstractImage super-resolution (SR) and image inpainting are two topical problems in medical image processing. Existing methods for solving the problems are either tailored to recovering a high-resolution version of the low-resolution image or focus on filling missing values, thus inevitably giving rise to poor performance when the acquisitions suffer from multiple degradations. In this paper, we explore the possibility of super-resolving and inpainting images to handle multiple degradations and therefore improve their usability. We construct a unified and scalable framework to overcome the drawbacks of propagated errors caused by independent learning. We additionally provide improvements over previously proposed super-resolution approaches by modeling image degradation directly from data observations rather than bicubic downsampling. To this end, we propose HLH-GAN, which includes a high-to-low (H-L) GAN together with a low-to-high (L-H) GAN in a cyclic pipeline for solving the medical image degradation problem. Our comparative evaluation demonstrates that the effectiveness of the proposed method on different brain MRI datasets. In addition, our method outperforms many existing super-resolution and inpainting approaches. Yawen Huang, Feng Zheng 0001, Junyu Jiang, Xiaoqian Wang 0001, Ling Shao 0001 |
IJCAI | 1 |
| 2020 | MCMT-GAN: Multi-Task Coherent Modality Transferable GAN for 3D Brain Image SynthesisabstractThe ability to synthesize multi-modality data is highly desirable for many computer-aided medical applications, e.g. clinical diagnosis and neuroscience research, since rich imaging cohorts offer diverse and complementary information unraveling human tissues. However, collecting acquisitions can be limited by adversary factors such as patient discomfort, expensive cost and scanner unavailability. In this paper, we propose a multi-task coherent modality transferable GAN (MCMT-GAN) to address this issue for brain MRI synthesis in an unsupervised manner. Through combining the bidirectional adversarial loss, cycle-consistency loss, domain adapted loss and manifold regularization in a volumetric space, MCMT-GAN is robust for multi-modality brain image synthesis with visually high fidelity. In addition, we complement discriminators collaboratively working with segmentors which ensure the usefulness of our results to segmentation task. Experiments evaluated on various cross-modality synthesis show that our method produces visually impressive results with substitutability for clinical post-processing and also exceeds the state-of-the-art methods. Yawen Huang, Feng Zheng 0001, Runmin Cong, Matthew R. Scott, Ling Shao 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Cross-Modality Image Synthesis via Weakly Coupled and Geometry Co-Regularized Joint Dictionary LearningabstractMulti-modality medical imaging is increasingly used for comprehensive assessment of complex diseases in either diagnostic examinations or as part of medical research trials. Different imaging modalities provide complementary information about living tissues. However, multi-modal examinations are not always possible due to adversary factors, such as patient discomfort, increased cost, prolonged scanning time, and scanner unavailability. In additionally, in large imaging studies, incomplete records are not uncommon owing to image artifacts, data corruption or data loss, which compromise the potential of multi-modal acquisitions. In this paper, we propose a weakly coupled and geometry co-regularized joint dictionary learning method to address the problem of cross-modality synthesis while considering the fact that collecting the large amounts of training data is often impractical. Our learning stage requires only a few registered multi-modality image pairs as training data. To employ both paired images and a large set of unpaired data, a cross-modality image matching criterion is proposed. Then, we propose a unified model by integrating such a criterion into the joint dictionary learning and the observed common feature space for associating cross-modality data for the purpose of synthesis. Furthermore, two regularization terms are added to construct robust sparse representations. Our experimental results demonstrate superior performance of the proposed model over state-of-the-art methods. Yawen Huang, Ling Shao 0001, Alejandro F. Frangi |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Simultaneous Super-Resolution and Cross-Modality Synthesis of 3D Medical Images Using Weakly-Supervised Joint Convolutional Sparse CodingabstractMagnetic Resonance Imaging (MRI) offers high-resolution in vivo imaging and rich functional and anatomical multimodality tissue contrast. In practice, however, there are challenges associated with considerations of scanning costs, patient comfort, and scanning time that constrain how much data can be acquired in clinical or research studies. In this paper, we explore the possibility of generating high-resolution and multimodal images from low-resolution single-modality imagery. We propose the weakly-supervised joint convolutional sparse coding to simultaneously solve the problems of super-resolution (SR) and cross-modality image synthesis. The learning process requires only a few registered multimodal image pairs as the training set. Additionally, the quality of the joint dictionary learning can be improved using a larger set of unpaired images. To combine unpaired data from different image resolutions/modalities, a hetero-domain image alignment term is proposed. Local image neighborhoods are naturally preserved by operating on the whole image domain (as opposed to image patches) and using joint convolutional sparse coding. The paired images are enhanced in the joint learning process with unpaired data and an additional maximum mean discrepancy term, which minimizes the dissimilarity between their feature distributions. Experiments show that the proposed method outperforms state-of-the-art techniques on both SR reconstruction and simultaneous SR and cross-modality synthesis. Yawen Huang, Ling Shao 0001, Alejandro F. Frangi |
CVPR | 1 |
| 2017 | DOTE: Dual cOnvolutional filTer lEarning for Super-Resolution and Cross-Modality Synthesis in MRI
Yawen Huang, Ling Shao 0001, Alejandro F. Frangi |
MICCAI (3) | 1 |
| 2016 | Color object recognition via cross-domain learning on RGB-D imagesabstractThis paper addresses the object recognition problem using multiple-domain inputs. We present a novel approach that utilizes labeled RGB-D data in the training stage, where depth features are extracted for enhancing the discriminative capability of the original learning system that only relies on RGB images. The highly dissimilar source and target domain data are mapped into a unified feature space through transfer at both feature and classifier levels. In order to alleviate cross-domain discrepancy, we employ a state-of-the-art domain-adaptive dictionary learning algorithm that updates image representations in both domains and the classifier parameters simultaneously. The proposed method is trained on a RGB-D Object dataset and evaluated on the Caltech-256 dataset. Experimental results suggest that our approach can lead to significant performance gain over the state-of-the-art methods. Yawen Huang, Fan Zhu 0001, Ling Shao 0001, Alejandro F. Frangi |
ICRA | 1 |