EDBT 2026 Demo / reviewers in the wild / expert
Xi Ouyang
dblp:174/1636
· DBLP profile ↗
21ranked-venue papers
7as first author
15since 2021 · last 2025
0000-0002-7841-5359ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 5 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Medical SegmentationabstractRapid advancements in medical image segmentation performance have been significantly driven by the development of Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs). These models follow discriminative pixel-wise classification learning paradigm and often have limited ability to generalize across diverse medical imaging datasets. In this manuscript, we introduce Generative Medical Segmentation (GMS), a novel generative approach to perform image segmentation. GMS employs a robust pre-trained vision foundation model to extract latent representations for images and corresponding ground truth masks, followed by a lightweight model that learns a mapping function from the image to the mask in the latent space. Once trained, the model can generate estimated segmentation masks using the pre-trained vision foundation model to decode the predicted latent mask representation back into image space. The design of GMS leads to fewer trainable parameters in the model, reducing the risk of overfitting and enhancing its generalization capability. Our experimental analysis across five open-source datasets in different medical imaging domains demonstrates GMS outperforms existing discriminative and generative segmentation models. Furthermore, GMS is able to generalize well across datasets of the same imaging modality from different centers. Our experiments suggest GMS offers a scalable and effective solution for medical image segmentation. Jiayu Huo, Xi Ouyang, Sébastien Ourselin, Rachel Sparks |
AAAI | 2 |
| 2025 | Learning better contrastive view from radiologist's gaze
Sheng Wang 0014, Zihao Zhao 0002, Zixu Zhuang, Xi Ouyang, Lichi Zhang, Zheren Li, Chong Ma 0004, Tianming Liu 0001, Dinggang Shen, Qian Wang 0001 |
Pattern Recognit. | 4 |
| 2024 | Prompt-Based Segmentation Model of Anatomical Structures and Lesions in CT Images
Xi Ouyang, Dongdong Gu, Qianqian Chen 0002, Yiqiang Zhan, Xiang Sean Zhou, Feng Shi 0001, Zhong Xue, Dinggang Shen |
MICCAI (8) | 1 |
| 2024 | Exploiting Latent Classes for Medical Image Segmentation from Partially Labeled Datasets
Xiangyu Zhao 0003, Xi Ouyang, Lichi Zhang, Zhong Xue, Dinggang Shen |
MICCAI (8) | 2 |
| 2024 | Carotid Vessel Wall Segmentation Through Domain Aligner, Topological Learning, and Segment Anything Model for Sparse Annotation in MR ImagesabstractMedical image analysis poses significant challenges due to limited availability of clinical data, which is crucial for training accurate models. This limitation is further compounded by the specialized and labor-intensive nature of the data annotation process. For example, despite the popularity of computed tomography angiography (CTA) in diagnosing atherosclerosis with an abundance of annotated datasets, magnetic resonance (MR) images stand out with better visualization for soft plaque and vessel wall characterization. However, the higher cost and limited accessibility of MR, as well as time-consuming nature of manual labeling, contribute to fewer annotated datasets. To address these issues, we formulate a multi-modal transfer learning network, named MT-Net, designed to learn from unpaired CTA and sparsely-annotated MR data. Additionally, we harness the Segment Anything Model (SAM) to synthesize additional MR annotations, enriching the training process. Specifically, our method first segments vessel lumen regions followed by precise characterization of carotid artery vessel walls, thereby ensuring both segmentation accuracy and clinical relevance. Validation of our method involved rigorous experimentation on publicly available datasets from COSMOS and CARE-II challenge, demonstrating its superior performance compared to existing state-of-the-art techniques. Xibao Li, Xi Ouyang, Zhongxiang Ding, Yuyao Zhang 0005, Zhong Xue, Feng Shi 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2023 | HENet: Hierarchical Enhancement Network for Pulmonary Vessel Segmentation in Non-contrast CT Images
Xiao Zhang 0028, Dongdong Gu, Sheng Wang 0014, Jiayu Huo, Zhihao Jiang 0001, Feng Shi 0001, Zhong Xue, Yiqiang Zhan, Xi Ouyang, Dinggang Shen |
MICCAI (3) | 11 |
| 2023 | Image synthesis with disentangled attributes for chest X-ray nodule augmentation and detection
Zhenrong Shen 0001, Xi Ouyang, Bin Xiao 0010, Jie-Zhi Cheng, Dinggang Shen, Qian Wang 0001 |
Medical Image Anal. | 2 |
| 2023 | Knee Cartilage Defect Assessment by Graph Representation and Surface ConvolutionabstractKnee osteoarthritis (OA) is the most common osteoarthritis and a leading cause of disability. Cartilage defects are regarded as major manifestations of knee OA, which are visible by magnetic resonance imaging (MRI). Thus early detection and assessment for knee cartilage defects are important for protecting patients from knee OA. In this way, many attempts have been made on knee cartilage defect assessment by applying convolutional neural networks (CNNs) to knee MRI. However, the physiologic characteristics of the cartilage may hinder such efforts: the cartilage is a thin curved layer, implying that only a small portion of voxels in knee MRI can contribute to the cartilage defect assessment; heterogeneous scanning protocols further challenge the feasibility of the CNNs in clinical practice; the CNN-based knee cartilage evaluation results lack interpretability. To address these challenges, we model the cartilages structure and appearance from knee MRI into a graph representation, which is capable of handling highly diverse clinical data. Then, guided by the cartilage graph representation, we design a non-Euclidean deep learning network with the self-attention mechanism, to extract cartilage features in the local and global, and to derive the final assessment with a visualized result. Our comprehensive experiments show that the proposed method yields superior performance in knee cartilage defect assessment, plus its convenient 3D visualization for interpretability. Zixu Zhuang, Liping Si, Sheng Wang 0014, Kai Xuan, Xi Ouyang, Yiqiang Zhan, Zhong Xue, Lichi Zhang, Dinggang Shen, Weiwu Yao, Qian Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Automatic Grading Assessments for Knee MRI Cartilage Defects via Self-ensembling Semi-supervised Learning with Dual-Consistency
Jiayu Huo, Xi Ouyang, Liping Si, Kai Xuan, Sheng Wang 0014, Weiwu Yao, Dahong Qian, Zhong Xue, Qian Wang 0001, Dinggang Shen, Lichi Zhang |
Medical Image Anal. | 2 |
| 2022 | GAN-Guided Deformable Attention Network for Identifying Thyroid Nodules in Ultrasound ImagesabstractEarly detection and identification of malignant thyroid nodules, a vital precursory to the treatment, is a difficult task even for experienced clinicians. Many Computer-Aided Diagnose (CAD) systems have been developed to assist clinicians in performing this task on ultrasonic images. Learning-based CAD systems for thyroid nodules generally accommodate both nodule detection/ segmentation and fine-grained classification for its malignancy, and prior researches often treat aforementioned tasks in separate stages, leading to additional computational costs. In this paper, we utilize an online class activation mapping (CAM) mechanism to guide the network to learn discriminative features for identifying thyroid nodules in ultrasound images, called CAM attention network. It takes nodule masks as localization cues for direct spatial attention of the classification module, thereby avoiding isolated training for classification. Meanwhile, we propose a deformable convolution module to add offsets to the regular grid sampling locations in the standard convolution, guiding the network to capture more discriminative features of nodule areas. Furthermore, we use a generative adversarial network (GAN)to ensure reliable deformations of nodules from the deformable convolution module. Our proposed CAM attention network has already achieved the 2nd place in the classification task of TN-SCUI 2020, a MICCAI 2020 Challenge with the largest set of thyroid nodule ultrasound images according to our knowledge. The further inclusion of our proposed GAN-guided deformable module allows for capturing more fine-grained features between benign and malignant nodules, and further improves the classification accuracy to a new state-of-the-art level. Jintao Lu, Xi Ouyang, Xueda Shen, Zhiming Cui 0001, Qian Wang 0001, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 2 |
| 2022 | Follow My Eye: Using Gaze to Supervise Computer-Aided DiagnosisabstractWhen deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a properly functioning DNN. This demand for supervision data and labels is a major bottleneck in current medical image analysis, since collecting a large number of annotations from experienced experts can be time-consuming and expensive. In this paper, we demonstrate that the eye movement of radiologists reading medical images can be a new form of supervision to train the DNN-based computer-aided diagnosis (CAD) system. Particularly, we record the tracks of the radiologists' gaze when they are reading images. The gaze information is processed and then used to supervise the DNN's attention via an Attention Consistency module. To the best of our knowledge, the above pipeline is among the earliest efforts to leverage expert eye movement for deep-learning-based CAD. We have conducted extensive experiments on knee X-ray images for osteoarthritis assessment. The results show that our method can achieve considerable improvement in diagnosis performance, with the help of gaze supervision. Sheng Wang 0014, Xi Ouyang, Tianming Liu 0001, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Domain Generalization for Mammography Detection via Multi-style and Multi-view Contrastive Learning
Zheren Li, Zhiming Cui 0001, Sheng Wang 0014, Yuji Qi, Xi Ouyang, Qitian Chen, Yuezhi Yang, Zhong Xue, Dinggang Shen, Jie-Zhi Cheng |
MICCAI (7) | 5 |
| 2021 | Self-adversarial Learning for Detection of Clustered Microcalcifications in Mammograms
Xi Ouyang, Jifei Che, Qitian Chen, Zheren Li, Yiqiang Zhan, Zhong Xue, Qian Wang 0001, Jie-Zhi Cheng, Dinggang Shen |
MICCAI (7) | 1 |
| 2021 | Nodule Synthesis and Selection for Augmenting Chest X-ray Nodule Detection
Zhenrong Shen 0001, Xi Ouyang, Zhuochen Wang, Yiqiang Zhan, Zhong Xue, Qian Wang 0001, Jie-Zhi Cheng, Dinggang Shen |
PRCV (3) | 2 |
| 2021 | Learning Hierarchical Attention for Weakly-Supervised Chest X-Ray Abnormality Localization and DiagnosisabstractWe consider the problem of abnormality localization for clinical applications. While deep learning has driven much recent progress in medical imaging, many clinical challenges are not fully addressed, limiting its broader usage. While recent methods report high diagnostic accuracies, physicians have concerns trusting these algorithm results for diagnostic decision-making purposes because of a general lack of algorithm decision reasoning and interpretability. One potential way to address this problem is to further train these models to localize abnormalities in addition to just classifying them. However, doing this accurately will require a large amount of disease localization annotations by clinical experts, a task that is prohibitively expensive to accomplish for most applications. In this work, we take a step towards addressing these issues by means of a new attention-driven weakly supervised algorithm comprising a hierarchical attention mining framework that unifies activation- and gradient-based visual attention in a holistic manner. Our key algorithmic innovations include the design of explicit ordinal attention constraints, enabling principled model training in a weakly-supervised fashion, while also facilitating the generation of visual-attention-driven model explanations by means of localization cues. On two large-scale chest X-ray datasets (NIH ChestX-ray14 and CheXpert), we demonstrate significant localization performance improvements over the current state of the art while also achieving competitive classification performance. Xi Ouyang, Srikrishna Karanam, Ziyan Wu 0001, Terrence Chen, Jiayu Huo, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng |
IEEE Trans. Medical Imaging | 1 |
| 2020 | SAANet: Siamese action-units attention network for improving dynamic facial expression recognition
Daizong Liu, Xi Ouyang, Shuangjie Xu, Pan Zhou 0001, Kun He 0001, Shiping Wen 0001 |
Neurocomputing | 2 |
| 2020 | Dual-Sampling Attention Network for Diagnosis of COVID-19 From Community Acquired PneumoniaabstractThe coronavirus disease (COVID-19) is rapidly spreading all over the world, and has infected more than 1,436,000 people in more than 200 countries and territories as of April 9, 2020. Detecting COVID-19 at early stage is essential to deliver proper healthcare to the patients and also to protect the uninfected population. To this end, we develop a dual-sampling attention network to automatically diagnose COVID-19 from the community acquired pneumonia (CAP) in chest computed tomography (CT). In particular, we propose a novel online attention module with a 3D convolutional network (CNN) to focus on the infection regions in lungs when making decisions of diagnoses. Note that there exists imbalanced distribution of the sizes of the infection regions between COVID-19 and CAP, partially due to fast progress of COVID-19 after symptom onset. Therefore, we develop a dual-sampling strategy to mitigate the imbalanced learning. Our method is evaluated (to our best knowledge) upon the largest multi-center CT data for COVID-19 from 8 hospitals. In the training-validation stage, we collect 2186 CT scans from 1588 patients for a 5-fold cross-validation. In the testing stage, we employ another independent large-scale testing dataset including 2796 CT scans from 2057 patients. Results show that our algorithm can identify the COVID-19 images with the area under the receiver operating characteristic curve (AUC) value of 0.944, accuracy of 87.5%, sensitivity of 86.9%, specificity of 90.1%, and F1-score of 82.0%. With this performance, the proposed algorithm could potentially aid radiologists with COVID-19 diagnosis from CAP, especially in the early stage of the COVID-19 outbreak. Xi Ouyang, Jiayu Huo, Liming Xia, Jun Liu 0075, Zhanhao Mo, Fuhua Yan, Zhongxiang Ding, Bin Song 0002, Feng Shi 0001, Huan Yuan, Ying Wei 0009, Xiaohuan Cao, Yaozong Gao, Dijia Wu, Qian Wang 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 1 |
| 2019 | Weakly Supervised Segmentation Framework with Uncertainty: A Study on Pneumothorax Segmentation in Chest X-ray
Xi Ouyang, Zhong Xue, Yiqiang Zhan, Xiang Sean Zhou, Qian Wang 0001, Jie-Zhi Cheng |
MICCAI (6) | 1 |
| 2018 | Spatial Pyramid Pooling Mechanism in 3D Convolutional Network for Sentence-Level ClassificationabstractIn this paper, we investigate the usage of the convolutional neural network (CNN) to propose a novel end-to-end language processing structure to model textual data for this task. In particular, we propose a 3D CNN structure for the task, which is featured by spatial pyramid pooling (SPP). To our knowledge, it is the first time that 3D convolution and SPP structure are applied together in language processing issues. Compared with methods of 2D CNNs, the proposed method can effectively and efficiently capture the complicated internal relations in sentences. Furthermore, in previous work, the issue of sentence length variety is usually addressed by padding zero to make all sentences vectors to a fixed length, which causes too much redundant and useless noise. Inspired by the SPP structure for object detection in image processing, this issue can be well handled with the SPP, which divides the sentences into several length sections for respective pooling processing. Experiments are conducted for the task of sentence classification as well as relation classification. Experiments on Stanford Treebank, TREC, subj, and Yelp datasets demonstrate that our proposed method can outperform other state-of-the-art models, with respect to classification accuracy. Auxiliary attempts to leverage our method to SemEval-2010 Task 8 dataset further substantiate the model's capability of extracting features efficiently. Xi Ouyang, Kang Gu, Pan Zhou 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | ZipNet-GAN: Inferring Fine-grained Mobile Traffic Patterns via a Generative Adversarial Neural NetworkabstractLarge-scale mobile traffic analytics is becoming essential to digital infrastructure provisioning, public transportation, events planning, and other domains. Monitoring city-wide mobile traffic is however a complex and costly process that relies on dedicated probes. Some of these probes have limited precision or coverage, others gather tens of gigabytes of logs daily, which independently offer limited insights. Extracting fine-grained patterns involves expensive spatial aggregation of measurements, storage, and post-processing. In this paper, we propose a mobile traffic super-resolution technique that overcomes these problems by inferring narrowly localised traffic consumption from coarse measurements. We draw inspiration from image processing and design a deep-learning architecture tailored to mobile networking, which combines Zipper Network (ZipNet) and Generative Adversarial neural Network (GAN) models. This enables to uniquely capture spatio-temporal relations between traffic volume snapshots routinely monitored over broad coverage areas ('low-resolution') and the corresponding consumption at 0.05 km2 level ('high-resolution') usually obtained after intensive computation. Experiments we conduct with a real-world data set demonstrate that the proposed ZipNet(-GAN) infers traffic consumption with remarkable accuracy and up to 100X higher granularity as compared to standard probing, while outperforming existing data interpolation techniques. To our knowledge, this is the first time super-resolution concepts are applied to large-scale mobile traffic analysis and our solution is the first to infer fine-grained urban traffic patterns from coarse aggregates. Chaoyun Zhang, Xi Ouyang, Paul Patras |
CoNEXT | 2 |
| 2017 | Audio-visual emotion recognition using deep transfer learning and multiple temporal modelsabstractThis paper presents the techniques used in our contribution to Emotion Recognition in the Wild 2017 video based sub-challenge. The purpose of the sub-challenge is to classify the six basic emotions (angry, sad, happy, surprise, fear and disgust) and neutral. Our proposed solution utilizes three state-of-the-arts techniques to overcome the challenges for the wild emotion recognition. Deep network transfer learning is used for feature extraction. Spatial-temporal model fusion is to make full use of the complementary of different networks. Semi-auto reinforcement learning is for the optimization of fusion strategy based on dynamic outside feedbacks given by challenge organizers. The overall accuracy of the proposed approach on the challenge test dataset is 57.2%, which is better than the challenge baseline of 40.47% . Xi Ouyang, Shigenori Kawaai, Ester Gue Hua Goh, Shengmei Shen, Wan Ding, Huaiping Ming, Dong-Yan Huang |
ICMI | 1 |