VLDB 2026 Research / reviewers in the wild / expert
Jiayu Huo
dblp:144/2201
· DBLP profile ↗
12ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-9811-0368ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Medical SegmentationabstractRapid advancements in medical image segmentation performance have been significantly driven by the development of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). These models follow discriminative pixel-wise classification learning paradigm and often have limited ability to generalize across diverse medical imaging datasets. In this manuscript, we introduce Generative Medical Segmentation (GMS), a novel generative approach to perform image segmentation. GMS employs a robust pre-trained vision foundation model to extract latent representations for images and corresponding ground truth masks, followed by a lightweight model that learns a mapping function from the image to the mask in the latent space. Once trained, the model can generate estimated segmentation masks using the pre-trained vision foundation model to decode the predicted latent mask representation back into image space. The design of GMS leads to fewer trainable parameters in the model, reducing the risk of overfitting and enhancing its generalization capability. Our experimental analysis across five open-source datasets in different medical imaging domains demonstrates GMS outperforms existing discriminative and generative segmentation models. Furthermore, GMS is able to generalize well across datasets of the same imaging modality from different centers. Our experiments suggest GMS offers a scalable and effective solution for medical image segmentation. Jiayu Huo, Xi Ouyang, Sébastien Ourselin, Rachel Sparks |
AAAI | 1 |
| 2025 | Motion-Boundary-Driven Unsupervised Surgical Instrument Segmentation in Low-Quality Optical Flow
Yang Liu 0271, Peiran Wu, Jiayu Huo, Gongyu Zhang, Christos Bergeles, Rachel Sparks, Prokar Dasgupta, Alejandro Granados, Sébastien Ourselin |
MICCAI (9) | 3 |
| 2025 | Self-supervised brain lesion generation for effective data augmentation of medical imagesabstractAccurate brain lesion delineation is important for planning neurosurgical treatment. Automatic brain lesion segmentation methods based on convolutional neural networks have demonstrated remarkable performance. However, neural network performance is constrained by the lack of large-scale well-annotated training datasets. In this manuscript, we propose a comprehensive framework to efficiently generate new samples for training a brain lesion segmentation model. We first train a self-supervised lesion generator based on the adversarial autoencoder to model lesion appearance and shape. Next, we utilize a novel image composition algorithm, Soft Poisson Blending, to seamlessly combine synthetic lesions and brain images to obtain training samples. Finally, to effectively train the brain lesion segmentation model with augmented images we introduce a new prototype consistence regularization to align real and synthetic features. Our framework is validated by extensive experiments on two public brain lesion segmentation datasets: ATLAS v2.0 and Shift MS. Our method outperforms existing brain image data augmentation schemes. For instance, our method improves the Dice from 50.36% to 60.23% compared to the UNet with conventional data augmentation techniques for the ATLAS v2.0 dataset. Jiayu Huo, Sébastien Ourselin, Rachel Sparks |
Neural Networks | 1 |
| 2024 | MatchSeg: Towards Better Segmentation via Reference Image MatchingabstractRecently, automated medical image segmentation methods based on deep learning have achieved great success. However, they heavily rely on large annotated datasets, which are costly and time-consuming to acquire. Few-shot learning aims to overcome the need for annotated data by using a small labeled dataset, known as a support set, to guide predicting labels for new, unlabeled images, known as the query set. Inspired by this paradigm, we introduce MatchSeg, a novel framework that enhances medical image segmentation through strategic reference image matching. We leverage contrastive language-image pre-training (CLIP) to select highly relevant samples when defining the support set. Additionally, we design a Joint Attention module to strengthen the interaction between support and query features, facilitating a more effective knowledge transfer between support and query sets. We validated our method across four public datasets. Experimental results demonstrate superior segmentation performance and powerful domain generalization ability of MatchSeg against existing methods for domain-specific and cross-domain segmentation tasks. Our code is made available at https://github.com/keeplearning-again/MatchSeg Jiayu Huo, Ruiqiang Xiao, Yang Liu 0271, Sébastien Ourselin, Rachel Sparks |
BIBM | 1 |
| 2023 | SKiT: a Fast Key Information Video Transformer for Online Surgical Phase RecognitionabstractThis paper introduces SKiT, a fast Key information Transformer for phase recognition of videos. Unlike previous methods that rely on complex models to capture long-term temporal information, SKiT accurately recognizes high-level stages of videos using an efficient key pooling operation. This operation records important key information by retaining the maximum value recorded from the beginning up to the current video frame, with a time complexity of ${\mathcal{O}}\left( 1 \right)$. Experimental results on Cholec80 and AutoLaparo surgical datasets demonstrate the ability of our model to recognize phases in an online manner. SKiT achieves higher performance than state-of-the-art methods with an accuracy of 92.5% and 82.9% on Cholec80 and AutoLaparo, respectively, while running the temporal model eight times faster (7ms v.s. 55ms) than LoViT, which uses ProbSparse to capture global information. We highlight that the inference time of SKiT is constant, and independent from the input length, making it a stable choice for keeping a record of important global information, that appears on long surgical videos, essential for phase recognition. To sum up, we propose an effective and efficient model for surgical phase recognition that leverages key global information. This has an intrinsic value when performing this task in an online manner on long surgical videos for stable real-time surgical recognition systems. Yang Liu 0271, Jiayu Huo, Jingjing Peng, Rachel Sparks, Prokar Dasgupta, Alejandro Granados, Sébastien Ourselin |
ICCV | 2 |
| 2023 | HENet: Hierarchical Enhancement Network for Pulmonary Vessel Segmentation in Non-contrast CT Images
Xiao Zhang 0028, Dongdong Gu, Sheng Wang 0014, Jiayu Huo, Zhihao Jiang 0001, Feng Shi 0001, Zhong Xue, Yiqiang Zhan, Xi Ouyang, Dinggang Shen |
MICCAI (3) | 5 |
| 2022 | Automatic Grading Assessments for Knee MRI Cartilage Defects via Self-ensembling Semi-supervised Learning with Dual-Consistency
Jiayu Huo, Xi Ouyang, Liping Si, Kai Xuan, Sheng Wang 0014, Weiwu Yao, Dahong Qian, Zhong Xue, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Medical Image Anal. | 1 |
| 2021 | Learning Hierarchical Attention for Weakly-Supervised Chest X-Ray Abnormality Localization and DiagnosisabstractWe consider the problem of abnormality localization for clinical applications. While deep learning has driven much recent progress in medical imaging, many clinical challenges are not fully addressed, limiting its broader usage. While recent methods report high diagnostic accuracies, physicians have concerns trusting these algorithm results for diagnostic decision-making purposes because of a general lack of algorithm decision reasoning and interpretability. One potential way to address this problem is to further train these models to localize abnormalities in addition to just classifying them. However, doing this accurately will require a large amount of disease localization annotations by clinical experts, a task that is prohibitively expensive to accomplish for most applications. In this work, we take a step towards addressing these issues by means of a new attention-driven weakly supervised algorithm comprising a hierarchical attention mining framework that unifies activation- and gradient-based visual attention in a holistic manner. Our key algorithmic innovations include the design of explicit ordinal attention constraints, enabling principled model training in a weakly-supervised fashion, while also facilitating the generation of visual-attention-driven model explanations by means of localization cues. On two large-scale chest X-ray datasets (NIH ChestX-ray14 and CheXpert), we demonstrate significant localization performance improvements over the current state of the art while also achieving competitive classification performance. Xi Ouyang, Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, Jiayu Huo, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng |
IEEE Trans. Medical Imaging | 5 |
| 2020 | SLIR: Synthesis, localization, inpainting, and registration for image-guided thermal ablation of liver tumors
Dongming Wei, Sahar Ahmad, Jiayu Huo, Pu Huang 0001, Pew-Thian Yap, Zhong Xue, Jianqi Sun, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 3 |
| 2020 | Dual-Sampling Attention Network for Diagnosis of COVID-19 From Community Acquired PneumoniaabstractThe coronavirus disease (COVID-19) is rapidly spreading all over the world, and has infected more than 1,436,000 people in more than 200 countries and territories as of April 9, 2020. Detecting COVID-19 at early stage is essential to deliver proper healthcare to the patients and also to protect the uninfected population. To this end, we develop a dual-sampling attention network to automatically diagnose COVID-19 from the community acquired pneumonia (CAP) in chest computed tomography (CT). In particular, we propose a novel online attention module with a 3D convolutional network (CNN) to focus on the infection regions in lungs when making decisions of diagnoses. Note that there exists imbalanced distribution of the sizes of the infection regions between COVID-19 and CAP, partially due to fast progress of COVID-19 after symptom onset. Therefore, we develop a dual-sampling strategy to mitigate the imbalanced learning. Our method is evaluated (to our best knowledge) upon the largest multi-center CT data for COVID-19 from 8 hospitals. In the training-validation stage, we collect 2186 CT scans from 1588 patients for a 5-fold cross-validation. In the testing stage, we employ another independent large-scale testing dataset including 2796 CT scans from 2057 patients. Results show that our algorithm can identify the COVID-19 images with the area under the receiver operating characteristic curve (AUC) value of 0.944, accuracy of 87.5%, sensitivity of 86.9%, specificity of 90.1%, and F1-score of 82.0%. With this performance, the proposed algorithm could potentially aid radiologists with COVID-19 diagnosis from CAP, especially in the early stage of the COVID-19 outbreak. Xi Ouyang, Jiayu Huo, Liming Xia, Jun Liu 0075, Zhanhao Mo, Fuhua Yan, Zhongxiang Ding, Bin Song 0002, Feng Shi 0001, Huan Yuan, Ying Wei 0009, Xiaohuan Cao, Yaozong Gao, Dijia Wu, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Synthesis and Inpainting-Based MR-CT Registration for Image-Guided Thermal Ablation of Liver Tumors
Dongming Wei, Sahar Ahmad, Jiayu Huo, Wen Peng, Yunhao Ge, Zhong Xue, Pew-Thian Yap, Dinggang Shen, Qian Wang 0001 |
MICCAI (5) | 3 |
| 2019 | Thermodynamic edge entropy in Alzheimer's disease
Jianjia Wang, Jiayu Huo, Lichi Zhang |
Pattern Recognit. Lett. | 2 |