Shiqi Huang 0001

dblp:99/10462-1 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-1348-3817ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 PromptReg: Interactive Registration by "Corresponding Prompts" for Segment Anything Model (SAM)
abstract
Effectively establishing correspondence between two images is at the centre of image registration methods. Spatially omnipresent representations, including dense displacement fields (DDFs) and spatial (non-)rigid transformations, have been used to parameterise such correspondence. Alternatively, region-based representation uses paired regions of interest (ROIs) to represent region-level correspondence, while retaining its local and dense representation capability at pixel/voxel level if required. Thus, registration can be re-envisioned as a problem of segmenting corresponding paired ROIs in the to-be-registered images. In this work, we utilize models such as SAM, which are pre-trained on substantive datasets, to segment ROIs of the same class from two images, for a new training-free, non-iterative registration algorithm. First, a "corresponding prompt problem" is posed to find a corresponding Prompt Y on Image Y, given any vision Prompt X on Image X, such that the two respectively prompt-conditioned segmentations are a pair of corresponding ROIs from the two images. Second, we propose an "inverse prompt" solution to the corresponding prompt problem, by inverting Prompt X to the Image Y prompt space, where the Jacobian of prototypical features is used. Third, we propose a new registration algorithm that identifies multiple paired corresponding ROIs, by marginalizing the inverted Prompt X over both prompt and spatial spaces, random sampling Prompt X and spatial warping Image X. Comprehensive experiments were conducted on five applications of registering 3D prostate MR, 3D abdomen CT, 3D lung CT, 2D histopathology and, as a non-medical example, 2D aerial images. Based on metrics including Dice and target registration errors on anatomical structures, the proposed registration outperforms both intensity-based iterative algorithms and learning-based networks, even yielding competitive performance with weakly-supervised registration which requires fully-segmented training data.
Shiqi Huang 0001, Tingfa Xu, Jianan Li 0001, Shaheer U. Saeed, Ziyi Shen, Dean C. Barratt, Yipeng Hu
IEEE Trans. Image Process.1
2025 Register Anything: Estimating "Corresponding Prompts" for Segment Anything Model
Shiqi Huang 0001, Tingfa Xu, Wen Yan 0005, Dean C. Barratt, Yipeng Hu
MICCAI (4)1
2025 Competing for Pixels: A Self-Play Algorithm for Weakly-Supervised Semantic Segmentation
abstract
Weakly-supervised semantic segmentation (WSSS) methods, reliant on image-level labels indicating object presence, lack explicit correspondence between labels and regions of interest (ROIs), posing a significant challenge. Despite this, WSSS methods have attracted attention due to their much lower annotation costs compared to fully-supervised segmentation. Leveraging reinforcement learning (RL) self-play, we propose a novel WSSS method that gamifies image segmentation of a ROI. We formulate segmentation as a competition between two agents that compete to select ROI-containing patches until exhaustion of all such patches. The score at each time-step, used to compute the reward for agent training, represents likelihood of object presence within the selection, determined by an object presence detector pre-trained using only image-level binary classification labels of object presence. Additionally, we propose a game termination condition that can be called by either side upon exhaustion of all ROI-containing patches, followed by the selection of a final patch from each. Upon termination, the agent is incentivised if ROI-containing patches are exhausted or disincentivised if a ROI-containing patch is found by the competitor. This competitive setup ensures minimisation of over- or under-segmentation, a common problem with WSSS methods. Extensive experimentation across four datasets demonstrates significant performance improvements over recent state-of-the-art methods.
Shaheer U. Saeed, Shiqi Huang 0001, João Ramalhinho, Iani J. M. B. Gayo, Nina Montaña Brown, Ester Bonmati, Stephen P. Pereira, Brian R. Davidson, Dean C. Barratt, Matthew J. Clarkson, Yipeng Hu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Dynamic Client Distillation for Semi-Supervised Federated Learning in a Realistic Scenario
abstract
Recent advancements in semi-supervised federated learning (SSFL) have significantly enhanced public health services by enabling medical institutions to share model updates via a central server. However, most SSFL approaches are based on conservative assumptions, such as labels-at-server and labels-at-client, which fail to fully capture the complex and diverse data distributions inherent in medical institutions. To address this limitation, we introduce a novel application of SSFL tailored to a realistic client data scenario, encompassing clients with fully-labeled, partially-labeled, and fully-unlabeled data. This approach effectively navigates varying levels of data annotation by maximizing the utility of unlabeled samples within the client federation. To tackle the challenges posed by such a complex scenario, we propose a new SSFL framework, FedCD. FedCD incorporates three client-distilled models, each corresponding to a distinct client data distribution, alongside server-client federation. First, each client-distilled model condenses the diverse parameters of the client federation into robust knowledge through distillation. The contribution of each client model is then dynamically adjusted based on its proximity to the client-distilled model, ensuring that the framework adapts to the heterogeneous characteristics of individual clients. By aggregating client-distilled models, FedCD implements model drift correction, effectively mitigating parameter drift across heterogeneous models. This dynamic federated approach not only harnesses unlabeled data efficiently but also accommodates diverse annotation levels while adapting to varying data distributions. Extensive experiments on two medical image segmentation tasks and one classification task demonstrate the superiority of our method, highlighting its ability to address realistic challenges in medical data scenarios.
Tingfa Xu, Shiqi Huang 0001, Jianan Li 0001
IEEE Trans. Medical Imaging3
2024 One Registration is Worth Two Segmentations
Shiqi Huang 0001, Tingfa Xu, Ziyi Shen, Shaheer U. Saeed, Wen Yan 0005, Dean C. Barratt, Yipeng Hu
MICCAI (12)1
2023 Rethinking Few-Shot Medical Segmentation: A Vector Quantization View
abstract
The existing few-shot medical segmentation networks share the same practice that the more prototypes, the better performance. This phenomenon can be theoretically interpreted in Vector Quantization (VQ) view: the more prototypes, the more clusters are separated from pixel-wise feature points distributed over the full space. However, as we further think about few-shot segmentation with this perspective, it is found that the clusterization of feature points and the adaptation to unseen tasks have not received enough attention. Motivated by the observation, we propose a learning VQ mechanism consisting of grid-format VQ (GFVQ), self-organized VQ (SOVQ) and residual oriented VQ (ROVQ). To be specific, GFVQ generates the prototype matrix by averaging square grids over the spatial extent, which uniformly quantizes the local details; SOVQ adaptively assigns the feature points to different local classes and creates a new representation space where the learnable local prototypes are updated with a global view; ROVQ introduces residual information to fine-tune the aforementioned learned local prototypes without retraining, which benefits the generalization performance for the irrelevance to the training task. We empirically show that our VQ framework yields the state-of-the-art performance over abdomen, cardiac and prostate MRI datasets and expect this work will provoke a rethink of the current few-shot medical segmentation model design. Our code will soon be publicly available.
Shiqi Huang 0001, Tingfa Xu, Feng Mu, Jianan Li 0001
CVPR1
2023 Expert-Guided Knowledge Distillation for Semi-Supervised Vessel Segmentation
abstract
In medical image analysis, blood vessel segmentation is of considerable clinical value for diagnosis and surgery. The predicaments of complex vascular structures obstruct the development of the field. Despite many algorithms have emerged to get off the tight corners, they rely excessively on careful annotations for tubular vessel extraction. A practical solution is to excavate the feature information distribution from unlabeled data. This work proposes a novel semi-supervised vessel segmentation framework, named EXP-Net, to navigate through finite annotations. Based on the training mechanism of the Mean Teacher model, we innovatively engage an expert network in EXP-Net to enhance knowledge distillation. The expert network comprises knowledge and connectivity enhancement modules, which are respectively in charge of modeling feature relationships from global and detailed perspectives. In particular, the knowledge enhancement module leverages the vision transformer to highlight the long-range dependencies among multi-level token components; the connectivity enhancement module maximizes the properties of topology and geometry by skeletonizing the vessel in a non-parametric manner. The key components are dedicated to the conditions of weak vessel connectivity and poor pixel contrast. Extensive evaluations show that our EXP-Net achieves state-of-the-art performance on subcutaneous vessel, retinal vessel, and coronary artery segmentations.
Tingfa Xu, Shiqi Huang 0001, Feng Mu, Jianan Li 0001
IEEE J. Biomed. Health Informatics3
2023 SCANet: A Unified Semi-Supervised Learning Framework for Vessel Segmentation
abstract
Automatic subcutaneous vessel imaging with near-infrared (NIR) optical apparatus can promote the accuracy of locating blood vessels, thus significantly contributing to clinical venipuncture research. Though deep learning models have achieved remarkable success in medical image segmentation, they still struggle in the subfield of subcutaneous vessel segmentation due to the scarcity and low-quality of annotated data. To relieve it, this work presents a novel semi-supervised learning framework, SCANet, that achieves accurate vessel segmentation through an alternate training strategy. The SCANet is composed of a multi-scale recurrent neural network that embeds coarse-to-fine features and two auxiliary branches, a consistency decoder and an adversarial learning branch, responsible for strengthening fine-grained details and eliminating differences between ground-truths and predictions, respectively. Equipped with a novel semi-supervised alternate training strategy, the three components work collaboratively, enabling SCANet to accurately segment vessel regions with only a handful of labeled data and abounding unlabeled data. Moreover, to mitigate the shortage of annotated data in this field, we provide a new subcutaneous vessel dataset, VESSEL-NIR. Extensive experiments on a wide variety of tasks, including the segmentation of subcutaneous vessels, retinal vessels, and skin lesions, well demonstrate the superiority and generality of our approach.
Tingfa Xu, Ziyang Bian, Shiqi Huang 0001, Feng Mu, Bo Huang 0012, Yuze Xiao, Jianan Li 0001
IEEE Trans. Medical Imaging4
2022 Pixel-Adaptive Field-of-View for Remote Sensing Image Segmentation
abstract
Mineral segmentation of satellite imagery is crucial to mining surveying and monitoring. Conventional deep segmentation networks extract features at every position with a fixed field-of-view. Nevertheless, the rich content in large mineral scenes, which causes dramatically different local characteristics across regions, may require features with spatially varying field-of-view to achieve accurate segmentation. In light of this, we propose a novel Pixel-Adaptive Field-of-View (PA-FoV) module to adjust the field-of-view of a given feature map in a pixel-wise manner. Specifically, it refines the features at each position by a weighted aggregation of the outputs from atrous convolutions with different dilation rates adaptively depending on the position-specific content. The module works in a plug-and-play manner and can be flexibly inserted into any arbitrary backbone network or segmentation head, to boost feature representation and in turn improve the result of segmentation. Moreover, in order to mitigate the scarcity of labeled data, we further establish a benchmark remote sensing mineral dataset, dubbed RSMI, to facilitate research in this field. Extensive experiments show a simple addition of our PA-FoV module provides solid improvements on top of strong baselines, achieving state-of-the-art performance.
Feng Mu, Jianan Li 0001, Shiqi Huang 0001, Yongzhuo Pan, Tingfa Xu
IEEE Geosci. Remote. Sens. Lett.4
2022 RTNet: Relation Transformer Network for Diabetic Retinopathy Multi-Lesion Segmentation
abstract
Automatic diabetic retinopathy (DR) lesions segmentation makes great sense of assisting ophthalmologists in diagnosis. Although many researches have been conducted on this task, most prior works paid too much attention to the designs of networks instead of considering the pathological association for lesions. Through investigating the pathogenic causes of DR lesions in advance, we found that certain lesions are closed to specific vessels and present relative patterns to each other. Motivated by the observation, we propose a relation transformer block (RTB) to incorporate attention mechanisms at two main levels: a self-attention transformer exploits global dependencies among lesion features, while a cross-attention transformer allows interactions between lesion and vessel features by integrating valuable vascular information to alleviate ambiguity in lesion detection caused by complex fundus structures. In addition, to capture the small lesion patterns first, we propose a global transformer block (GTB) which preserves detailed information in deep network. By integrating the above blocks of dual-branches, our network segments the four kinds of lesions simultaneously. Comprehensive experiments on IDRiD and DDR datasets well demonstrate the superiority of our approach, which achieves competitive performance compared to state-of-the-arts.
Shiqi Huang 0001, Jianan Li 0001, Yuze Xiao, Tingfa Xu
IEEE Trans. Medical Imaging1