EDBT 2026 Demo / reviewers in the wild / expert
Zhiyong Huang 0004
dblp:181/2754-4
· DBLP profile ↗
23ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0001-6368-1008ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BACFormer: A robust boundary-aware transformer for medical image segmentation
Zhiyong Huang 0004, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Knowl. Based Syst. | 1 |
| 2026 | Bridging teacher-student representation domains for manifold-aware knowledge distillation
Shuai Miao, Zhiyong Huang 0004, Daidi Zhong, Mingyang Hou, Penghao Jia, Yan Yan 0022, Yushi Liu 0001 |
Knowl. Based Syst. | 2 |
| 2026 | Eliminating domain-related confounding factors in cross-domain one-shot medical image segmentation via causal inference
Mingyang Hou, Zhiyong Huang 0004, Daidi Zhong, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Medical Image Anal. | 2 |
| 2026 | Integrating visual and language cues via state space models for medical image segmentation
Mingyang Hou, Zhiyong Huang 0004, Daidi Zhong, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Neural Networks | 2 |
| 2026 | Enhancing Radiography-Report foundation model via Multi-View masked contrastive learning
Daidi Zhong, Zhiyong Huang 0004, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
Pattern Recognit. | 3 |
| 2026 | Overcoming Limitations in One-Shot Semantic Segmentation via Adaptive Visual-Text Guided Prototype Relationship OptimizationabstractThe scarcity of high-quality annotated data in medical imaging significantly constrains the performance of deep learning-based segmentation models. While few-shot medical image segmentation (FSMIS) has emerged as a promising solution, existing methods exhibit critical limitations when handling scenarios with inter-class similarity between foreground-background regions and intra-class heterogeneity within foreground objects. Current prototype-based approaches focus primarily on the holistic extraction of the prototype from support images, failing to distinguish subtle anatomical variations and complex feature representations effectively. The AVT-ProNet features three innovative components: 1) An Adaptive Visual-Text Prototype Generation (AVPG) module leveraging CLIP’s cross-modal guide capabilities through adaptive prompting strategies; 2) a graph-based multiregion prototyping relationship optimization (GMPRO) module establishing structural relationships between decomposed subregion prototypes via graph neural networks; 3) a foreground-background prototyping contrast learning (FBPCL) strategy implementing dual-space optimization through inter-class separation and intra-class compactness. The synergistic integration of multi-modal guidance, structural relationship modeling, and contrastive prototype refinement enables our framework to overcome existing limitations in FSMIS. Comprehensive evaluations across multiple clinical scenarios (CHAOS, SABS, and CMR datasets under diverse training configurations) demonstrate superior performance over state-of-the-art approaches, including PANet, CAT-Net, DMAP, and recent PAMI baselines, Source code is available at https://github.com/394481125/AVT-ProNet. Mingyang Hou, Zhiyong Huang 0004, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Multi-Domain Patch-Differentiated Transformer for vehicle re-identification
Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001, Daming Sun, Hans Gregersen |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | WTSF-ReID: Depth-driven Window-oriented Token Selection and Fusion for multi-modality vehicle re-identification with knowledge consistency constraint
Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Hessian-based mixed-precision quantization with transition aware training for neural networks
Zhiyong Huang 0004, Yunlan Zhao, Mingyang Hou, Shengdong Hu |
Neural Networks | 1 |
| 2025 | SCFMUNet: A fusion architecture based on multi-scale state space model and channel attention for medical image segmentation
Zhiyong Huang 0004, Mingyang Hou, Shiyao Zhou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
Neural Networks | 1 |
| 2025 | Feature-Tuning Hierarchical Transformer via token communication and sample aggregation constraint for object re-identification
Zhiyong Huang 0004, Mingyang Hou, Jiaming Pei, Yan Yan 0022, Yushi Liu 0001, Daming Sun |
Neural Networks | 2 |
| 2025 | Representation Selective Coupling via Token Sparsification for Multi-Spectral Object Re-IdentificationabstractTo tackle the challenge of single-spectral object re-identification in complex and dynamic lighting scenarios, multi-spectral object re-identification, which integrates visible light and infrared information, is gradually taking the lead. Nevertheless, the significant heterogeneity across spectra causes formidable obstacles for this task. Most existing approaches alleviate inter-spectral disparities by amalgamating representations from different spectra, ignoring the selection of spectrum-specific crucial information. To address this issue, we propose a novel Representation Selective Coupling Network (RSCNet) for multi-spectral object re-identification. Specifically, we design an Attention-Fourier Token Sparsification (AFTS) module to adaptively sparse and join tokens from multi-spectral images in the attention domain and Fourier domain. This not only preserves spectrum-specific crucial information but also reduces inter-spectral gaps by selective coupling of multi-spectral representation. Meanwhile, to further align multi-spectral information and guide the model to learn more discriminative representation, we propose an Information Unification Constraint (IUC) learning strategy. Both feature-level information constraint and distribution-level information constraint are simultaneously deployed in IUC. Finally, we conduct extensive experiments on three multi-spectral object re-identification benchmarks, and the experimental results verify the effectiveness of our proposed method. Zhiyong Huang 0004, Mingyang Hou, Jiaming Pei, Yan Yan 0022, Yushi Liu 0001, Daming Sun |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | MCECF: A Multiscale Complementary Enhanced Context Fusion Network for Remote Sensing Change DetectionabstractRemote sensing change detection (RSCD) holds significant research value in remote sensing (RS) image processing. In recent years, many researchers have achieved remarkable results in RSCD tasks using methods based on convolutional neural networks (CNNs) or Transformers. Considering the limited receptive field of CNN models and the high computational cost of Transformers, many researchers have combined the two approaches, yielding promising results. However, most current RSCD-based models focus solely on change and temporal information, overlooking their complementary relationship. Additionally, some multiscale feature fusion methods emphasize enhancing individual scales while neglecting the correlations between different scales. To address the above issues, we propose a multiscale complementary enhanced context fusion (MCECF) network. The network first introduces a global-local context aggregation module (GLCAM) to capture global-local context information while extracting multilevel feature maps. Subsequently, a complementary enhancement difference module (CEDM) is employed to complementarily aggregate the captured change and temporal information of bi-temporal RS image features. To fully leverage the correlations between multiscale features, a progressive decoder comprising a supervised spatial attention (SSA) mechanism and a multiscale complementary enhanced fusion module (MCEFM) was developed. Moreover, to tackle the disparity between changed and unchanged regions, a dual-branch dynamic attention fusion module (DAFM) was designed to enhance the model’s adaptability to diverse scenarios. We conducted comparative experiments on five RSCD datasets against nine state-of-the-art (SOTA) methods, and the results confirmed the effectiveness of the proposed MCECF in RSCD tasks. Our code will be made available athttps://github.com/kakuqikaduo/MCECF Zhiyong Huang 0004, Hongjiang Qiu, Mingyang Hou, Jiahong Wang, Yan Yan 0022, Yushi Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | MSCD-VM-UNet: A Vision Mamba Combining Multi-Scale Global and Local Feature Extraction With Cross-Domain Feature Fusion for Medical Image SegmentationabstractAccurate segmentation of tissues and lesions is essential for diagnosis and treatment. State Space Models (SSMs) have gained attention for their linear complexity and ability to model long-range dependencies. However, the existing Mamba architecture relies on direct skip connections, which limits its ability to integrate multi-scale and multi-level features and handle boundary details effectively. To address these limitations, we propose the MSCD-VM-UNet architecture, which incorporates three novel modules: the Spatial Group Multi-Scale Attention Module (SGMAM), the Cross-Domain Feature Fusion Module (CDFFM), and the Attention-Based Feature Injection Module (ABFIM). The SGMAM captures multi-scale global and local information and adaptively adjusts feature importance to highlight key regions while suppressing noise. The CDFFM enhances boundary and detail handling by aligning semantic features from both the frequency and spatial domains. The ABFIM utilizes attention mechanisms to adaptively fuse and weigh features from different scales and semantics, promoting feature collaboration and improving the model's robustness in complex tasks. Experiments on multiple datasets show that these modules significantly enhance the accuracy of MSCD-VM-UNet, setting a new benchmark for medical image segmentation. Zhiyong Huang 0004, Mingyang Hou, Yan Yan 0022, Yushi Liu 0001, Hans Gregersen |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Attention-guided fusion of transformers and CNNs for enhanced medical image segmentation
Shiyao Zhou, Zhiyong Huang 0004, Yuqin He, Yunlan Zhao |
Vis. Comput. | 3 |
| 2024 | CSwT-SR: Conv-Swin Transformer for Blind Remote Sensing Image Super-Resolution With Amplitude-Phase Learning and Structural Detail Alternating LearningabstractImage super-resolution (SR) stands as a pivotal process in the domains of image processing and computer vision, finding diverse applications in film, television, photography, surveillance, medical imaging, and remote sensing. In the context of remote sensing images (RSIs), the inherent challenge arises from low spatial resolution caused by factors such as sensor noise, orbit height, and weather conditions, necessitating SR reconstruction. An evident limitation of prevailing methods lies in their dependence on idealized fixed degradation models, which fail to capture the intricate degradation processes unique to remote sensing scenes. In response to these constraints, this article introduces an innovative blind image super-resolution reconstruction method tailored for remote sensing images. The proposed approach integrates convolution with a transformer and incorporates an amplitude-phase learning module (ALM) to comprehensively capture local and long-range dependencies while enhancing frequency information. The iterative optimization strategy refines texture information by carefully balancing structural and detail elements. Key contributions include a holistic approach to remote sensing image SR, ALM integration for precise feature representation, and the introduction of a patch-based frequency loss mechanism for evaluating frequency-domain features. Rigorous experiments demonstrate that compared with other state-of-the-art (SOTA) methods, the proposed algorithm delivers SR results with exceptional visual perception quality across three distinct remote sensing datasets. Mingyang Hou, Zhiyong Huang 0004, Yan Yan 0022, Yunlan Zhao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Semantic-Oriented Feature Coupling Transformer for Vehicle Re-Identification in Intelligent Transportation SystemabstractMore robust intelligent transportation systems including autonomous driving systems are in full flourish with the revolution of deep learning and the 6G wireless communication network. Vehicle Re-Identification, an indispensable branch of the intelligent transportation system, aims to retrieve specific vehicles captured from non-overlapping cameras. However, this is fundamentally challenging with the substantial inter-class similarity and substantial intra-class divergence. Embedding semantic information into vehicle re-identification task has gained ample interest, but the performance needs to be further improved. This work proposes a semantic-oriented feature coupling transformer (SOFCT) for vehicle re-identification as a solution. Specifically, the knowledge-based transformer is first embedded to model images with discriminative attributes. Second, original patches are divided into five semantic groups via semantics-patches coupling, and the feature extractions for different semantics are performed in the semantic feature extraction (SFE) transformer. Third, patch features are weighted via semantics-patches coupling in the patch feature weighting (PFW) transformer, the weighted feature is fed into subsequent encoders to excavate information. Finally, two groups of learnable semantics are embedded to automatically learn semantic features in the learnable semantic extraction (LSE) transformer. Experiments demonstrate that the proposed SOFCT method surpasses other state-of-the-arts with the mAP/Rank-1 of 80.7%/96.6%, 89.8%/84.5%, 86.4%/80.9%, and 84.3%/78.7% on VeRi776 and VehicleID. Zhiyong Huang 0004, Jiaming Pei, Lamia Tahsin, Daming Sun |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Transformer-based feature interactor for person re-identification with margin self-punishment loss
Zhiyong Huang 0004, Pinzhong Qin, Lamia Tahsin |
Image Vis. Comput. | 1 |
| 2022 | Learning diverse and deep clues for person reidentification
Wen-cheng Qin, Baojin Huang, Pinzhong Qin, Zhiyong Huang 0004, Daidi Zhong |
Image Vis. Comput. | 4 |
| 2022 | Gaussian-based probability fusion for person re-identification with Taylor angular margin loss
Zhiyong Huang 0004, Tianhui Guan, Wen-cheng Qin, Lamia Tahsin, Daming Sun |
Neural Comput. Appl. | 1 |
| 2022 | Joining features by global guidance with bi-relevance trihard loss for person re-identification
Wen-cheng Qin, Zhiyong Huang 0004, Lamia Tahsin, Daming Sun, Yuanhong Zhong |
Neural Comput. Appl. | 3 |
| 2022 | TriEP: Expansion-Pool TriHard Loss for Person Re-Identification
Wen-cheng Qin, Lamia Tahsin, Zhiyong Huang 0004 |
Neural Process. Lett. | 4 |
| 2022 | Deep Constraints Space via Channel Alignment for Visible-Infrared Person Re-identificationabstractReducing the inter-modality gap has been the core of visible-infrared person re-identification (VI-ReID). Most of the existing methods directly constrain the features to reduce the gap between the different modalities. However, the consistency of channel semantic information across different modalities is ignored. In this letter, channel alignment along with severe semantic constraints is proposed to alleviate the significant distribution differences between modalities. To obtain the optimum channel matching mode, a novel channel alignment mechanism termed Channel Instance Level Alignment (CILA) is proposed at the shallow layer, and matched channel features are constrained by the proposed Channel Instance Alignment (CIA) Loss. In the middle layer, we divide the features into multiple hierarchically-aware channel clusters and align the channel clusters by Channel Cluster Alignment (CCA) Loss. The proposed method is validated on SYSU-MM01 and RegDB, and extensive experiments show that the proposed method achieves competitive performance compared with the state of the arts. Wen-cheng Qin, Baojin Huang, Zhiyong Huang 0004, Lamia Tahsin, Daming Sun |
IEEE Signal Process. Lett. | 3 |