VLDB 2026 Research / reviewers in the wild / expert
Wufeng Xue
dblp:79/10698
· DBLP profile ↗
36ranked-venue papers
13as first author
21since 2021 · last 2026
0000-0002-9776-0975ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 8 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ctrl-GenAug: Controllable Generative Augmentation for Medical Sequence Classification
Haoran Dou, Shijing Chen, Ao Chang, Weiran Long, Erjiao Xu, Alejandro F. Frangi, Ruobing Huang, Wufeng Xue, Dong Ni 0001 |
Int. J. Comput. Vis. | 15 |
| 2026 | MBAS2024: A large-scale benchmark for multi-class bi-atrial segmentation in multi-center contrast-enhanced MRIsabstractAtrial fibrillation (AF), the most common cardiac arrhythmia, affects one in three adults over 45 years of age. Improving its treatment requires a better understanding of bi-atrial anatomy. Existing benchmarks have focused on the left atrial (LA) cavity, overlooking the fundamental challenges posed by bi-atrial anatomy, most notably the thin atrial walls, which are critical for substrate-guided ablation planning in patients with atrial fibrillation. To address these limitations, the Multi-class Bi-Atrial Segmentation 2024 Challenge (MBAS2024) introduced the first large-scale, multi-class benchmark for simultaneous segmentation of the LA cavity, right atrial (RA) cavity, and bi-atrial walls from late gadolinium-enhanced (LGE) MRI. We systematically evaluated 13 state-of-the-art methods on the world's largest curated bi-atrial dataset, comprising 175 3D multi-center scans with expert-validated annotations, providing a comprehensive assessment of current methodological capabilities and limitations. Key findings include: segmentation of the LA and RA cavities is generally robust to image quality, whereas atrial wall delineation is highly sensitive to image degradation. Performance varies across centers, indicating limited generalization of atrial wall segmentation across different acquisition protocols. Model architecture, rather than hyperparameter tuning, is the primary driver of performance, with U-Net-based models and emerging state-space models (e.g., UMambaBot) achieving higher accuracy at modest computational cost. Segmentation accuracy also varies along the slice dimension, with central slices segmented more reliably. Finally, hybrid labeling strategies-separating LA and RA cavities while merging bi-atrial walls into a single class-consistently improve performance. The MBAS2024 challenge establishes a foundational benchmark for bi-atrial segmentation, providing validated baselines and actionable insights to guide the development of clinically relevant, efficient, and anatomically aware segmentation algorithms to improve targeted ablation in patients with AF. Fangqiang Xu, James Kennelly, Alexander M. Zolotarev, Caroline H. Roney, Michal Nohel, Constantin Ulrich, Bryan Anenberg, Peter Chang, Yu Hon On, Marta Varela, Claas Thesing, Abhirup Banerjee, Enrique Almar-Munoz, Markus Tiefenthaler, Susana Merino-Caviedes, Emmanuel C. Nnadozie, Abdul Qayyum 0002, Moona Mazher, Waqas Anwaar, Wufeng Xue, Jingsu Kang, Lucas Beveridge, Malitha Gunawardhana, Kunihiko Kiuchi, Martin K. Stiles, Jichao Zhao |
Medical Image Anal. | 20 |
| 2026 | UINO-FSS: Unifying Representation Learning and Few-Shot Segmentation via Hierarchical Distillation and Mamba-HyperCorrelationabstractFew-shot semantic segmentation has attracted growing interest for its ability to generalize to novel object categories using only a few annotated samples. To address data scarcity, recent methods incorporate multiple foundation models to improve feature transferability and segmentation performance. However, they often rely on dual-branch architectures that combine pre-trained encoders to leverage complementary strengths, a design that limits flexibility and efficiency. This raises a fundamental question: "can we build a unified model that integrates knowledge from different foundation architectures?" Achieving this is, however, challenging due to the misalignment between class-agnostic segmentation capabilities and fine-grained discriminative representations. To this end, we present UINO-FSS (pronounced //), a novel framework built on the key observation that early-stage DINOv2 features exhibit distribution consistency with SAM's output embeddings. This consistency enables the integration of both models' knowledge into a single-encoder architecture via coarse-to-fine multimodal distillation. In particular, our segmenter consists of three core components: a bottleneck adapter for embedding alignment, a meta-visual prompt generator that leverages dense similarity volumes and semantic embeddings, and a mask decoder. Using hierarchical cross-model distillation, we effectively transfer SAM's knowledge into the segmenter, further enhanced by Mamba-based 4D correlation mining on support-query pairs. Extensive experiments show that UINO-FSS achieves new state-of-the-art results on COCO- $20^{i}$ under the 1-shot setting, with an mIoU of 64.5% (+2.2%), while also delivering competitive performance on PASCAL- $5^{i}$ . Zhiyue Tang, Wufeng Xue, Junkai Ji, LinLin Shen |
IEEE Trans. Image Process. | 3 |
| 2025 | EchoONE: Segmenting Multiple Echocardiography Planes in One ModelabstractIn clinical practice of echocardiography examinations, multiple planes containing the heart structures of different view are usually required in screening, diagnosis and treatment of cardiac disease. AI models for echocardiogra-phy have to be tailored for each specific plane due to the dramatic structure differences, thus resulting in repetition development and extra complexity. Effective solution for such a multi-plane segmentation (MPS) problem is highly demanded for medical images, yet has not been well investigated. In this paper, we propose a novel solution, EchoONE, for this problem with an SAM-based segmentation architecture, a prior-composable mask learning (PC-Mask) module for semantic-aware dense prompt generation, and a learnable CNN-branch with a simple yet effective local feature fusion and adaption (LFFA) module for SAM adapting. We extensively evaluated our method on multiple internal and external echocardiography datasets and achieved consistently state-of-the-art performance for multi-source datasets with different heart planes. This is the first time the MPS problem has been solved in one model for echocardiography data. The code will be available at https://github.com/a2502503/EchoONE. Jiongtong Hu, Wufeng Xue, Jun Cheng 0006, Dong Ni 0001 |
CVPR | 2 |
| 2024 | HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models
Yuhao Huang 0001, Wufeng Xue, Haoran Dou, Jun Cheng 0006, Dong Ni 0001 |
MICCAI (7) | 3 |
| 2024 | Steerable Pyramid Transform Enables Robust Left Ventricle Quantification
Kede Ma, Wufeng Xue |
PRCV (14) | 3 |
| 2023 | MUVF-YOLOX: A Multi-modal Ultrasound Video Fusion Network for Renal Tumor Diagnosis
Dong Ni 0001, Wufeng Xue, Dongmei Zhu, Jun Cheng 0006 |
MICCAI (5) | 4 |
| 2023 | Mitral Regurgitation Quantification from Multi-channel Ultrasound Images via Deep Learning
Keming Tang, Zhenyi Ge, Rongbo Ling, Jun Cheng 0006, Wufeng Xue, Cuizhen Pan, Xianhong Shu, Dong Ni 0001 |
MICCAI (6) | 5 |
| 2023 | Wall Thickness Estimation from Short Axis Ultrasound Images via Temporal Compatible Deformation Learning
Guijuan Peng, Jialan Zheng, Jun Cheng 0006, Yuanyuan Sheng, Yingqi Zheng, Yumei Yang, Wufeng Xue, Dong Ni 0001 |
MICCAI (6) | 12 |
| 2023 | Inflated 3D Convolution-Transformer for Weakly-Supervised Carotid Stenosis Grading with Ultrasound Videos
Yuhao Huang 0001, Wufeng Xue, Xin Yang 0009, Yuxin Zou, Qilong Ying, Yuanji Zhang, Dong Ni 0001 |
MICCAI (2) | 3 |
| 2023 | RecON: Online learning for sensorless freehand 3D ultrasound reconstruction
Mingyuan Luo, Xin Yang 0009, Hongzhang Wang, Haoran Dou, Xindi Hu, Yuhao Huang 0001, Nishant Ravikumar, Songcheng Xu, Yuanji Zhang, Yi Xiong 0001, Wufeng Xue, Alejandro F. Frangi, Dong Ni 0001 |
Medical Image Anal. | 11 |
| 2023 | Co-learning of appearance and shape for precise ejection fraction estimation from echocardiographic sequences
Hongrong Wei, Junqiang Ma, Yongjin Zhou 0002, Wufeng Xue, Dong Ni 0001 |
Medical Image Anal. | 4 |
| 2023 | Semi-Supervised Representation Learning for Segmentation on Medical Volumes and SequencesabstractBenefiting from the massive labeled samples, deep learning-based segmentation methods have achieved great success for two dimensional natural images. However, it is still a challenging task to segment high dimensional medical volumes and sequences, due to the considerable efforts for clinical expertise to make large scale annotations. Self/semi-supervised learning methods have been shown to improve the performance by exploiting unlabeled data. However, they are still lack of mining local semantic discrimination and exploitation of volume/sequence structures. In this work, we propose a semi-supervised representation learning method with two novel modules to enhance the features in the encoder and decoder, respectively. For the encoder, based on the continuity between slices/frames and the common spatial layout of organs across subjects, we propose an asymmetric network with an attention-guided predictor to enable prediction between feature maps of different slices of unlabeled data. For the decoder, based on the semantic consistency between labeled data and unlabeled data, we introduce a novel semantic contrastive learning to regularize the feature maps in the decoder. The two parts are trained jointly with both labeled and unlabeled volumes/sequences in a semi-supervised manner. When evaluated on three benchmark datasets of medical volumes and sequences, our model outperforms existing methods with a large margin of 7.3% DSC on ACDC, 6.5% on Prostate, and 3.2% on CAMUS when only a few labeled data is available. Further, results on the M&M dataset show that the proposed method yields improvement without using any domain adaption techniques for data from unknown domain. Intensive evaluations reveal the effectiveness of representation mining, and superiority on performance of our method. The code is available at https://github.com/CcchenzJ/BootstrapRepresentation. Zejian Chen, Tianfu Wang 0001, Jun Cheng 0006, Wufeng Xue, Dong Ni 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2022 | Sketch guided and progressive growing GAN for realistic and editable ultrasound image synthesis
Jiamin Liang, Xin Yang 0009, Yuhao Huang 0001, Haoming Li 0008, Shuangchi He, Xindi Hu, Zejian Chen, Wufeng Xue, Jun Cheng 0006, Dong Ni 0001 |
Medical Image Anal. | 8 |
| 2022 | AWSnet: An auto-weighted supervision attention network for myocardial scar and edema segmentation in multi-sequence cardiac magnetic resonance images
Kai-Ni Wang, Xin Yang 0009, Juzheng Miao, Lei Li 0020, Wufeng Xue, Guangquan Zhou, Xiahai Zhuang, Dong Ni 0001 |
Medical Image Anal. | 7 |
| 2022 | Joint Landmark and Structure Learning for Automatic Evaluation of Developmental Dysplasia of the HipabstractThe ultrasound (US) screening of the infant hip is vital for the early diagnosis of developmental dysplasia of the hip (DDH). The US diagnosis of DDH refers to measuring alpha and beta angles that quantify hip joint development. These two angles are calculated from key anatomical landmarks and structures of the hip. However, this measurement process is not trivial for sonographers and usually requires a thorough understanding of complex anatomical structures. In this study, we propose a multi-task framework to learn the relationships among landmarks and structures jointly and automatically evaluate DDH. Our multi-task networks are equipped with three novel modules. Firstly, we adopt Mask R-CNN as the basic framework to detect and segment key anatomical structures and add one landmark detection branch to form a new multi-task framework. Secondly, we propose a novel shape similarity loss to refine the incomplete anatomical structure prediction robustly and accurately. Thirdly, we further incorporate the landmark-structure consistent prior to ensure the consistency of the bony rim estimated from the segmented structure and the detected landmark. In our experiments, 1231 US images of the infant hip from 632 patients are collected, of which 247 images from 126 patients are tested. The average errors in alpha and beta angles are 2.221${}^{\circ }$and 2.899${}^{\circ }$. About 93% and 85% estimates of alpha and beta angles have errors less than 5 degrees, respectively. Experimental results demonstrate that the proposed method can accurately and robustly realize the automatic evaluation of DDH, showing great potential for clinical application. Xindi Hu, Xin Yang 0009, Xu Zhou 0005, Wufeng Xue, Yan Cao 0002, Shengfeng Liu, Yuhao Huang 0001, Shuangping Guo, Dong Ni 0001, Ning Gu 0003 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Improved Segmentation of Echocardiography With Orientation-Congruency of Optical Flow and Motion-Enhanced SegmentationabstractQuantification of left ventricular (LV) ejection fraction (EF) from echocardiography depends upon the identification of endocardium boundaries as well as the calculation of end-diastolic (ED) and end-systolic (ES) LV volumes. It's critical to segment the LV cavity for precise calculation of EF from echocardiography. Most of the existing echocardiography segmentation approaches either only segment ES and ED frames without leveraging the motion information, or the motion information is only utilized as an auxiliary task. To address the above drawbacks, in this work, we propose a novel echocardiography segmentation method which can effectively utilize the underlying motion information by accurately predicting optical flow (OF) fields. First, we devised a feature extractor shared by the segmentation and the optical flow sub-tasks for efficient information exchange. Then, we proposed a new orientation congruency constraint for the OF estimation sub-task by promoting the congruency of optical flow orientation between successive frames. Finally, we design a motion-enhanced segmentation module for the final segmentation. Experimental results show that the proposed method achieved state-of-the-art performance for EF estimation, with a Pearson correlation coefficient of 0.893 and a Mean Absolute Error of 5.20% when validated with echo sequences of 450 patients. Wufeng Xue, Junqiang Ma, Ti Bai, Tianfu Wang 0001, Dong Ni 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Regional Cardiac Motion Scoring With Multi-Scale Motion-Based Spatial AttentionabstractRegional cardiac motion scoring aims to classify the motion status of each myocardium segment into one of the four categories (normal, hypokinetic, akinetic, and dyskinetic) from multiple short-axis MR sequences. It is essential for prognosis and early diagnosis for various cardiac diseases. However, the complex motion procedure of the myocardium and the invisible pattern differences pose great challenges, leading to low performance for automatic methods. Most existing works mitigate the task by differentiating the normal motion patterns from the abnormal ones, without fine-grained motion scoring. We propose an effective method for the task of cardiac motion scoring by connecting a bottom-up and another top-down branch with a novel motion-based spatial attention module in multi-scale space. Specifically, we use the convolution blocks for low-level feature extraction that acts as a bottom-up mechanism, and the task of optical flow for explicit motion extraction that acts as a top-down mechanism for high-level allocation of spatial attention. To this end, a newly designed Multi-scale Motion-based Spatial Attention (MMSA) module is used as the pivot connecting the bottom-up part and the top-down part, and adaptively weight the low-level features according to the motion information. Experimental results on a newly constructed dataset of 1440 myocardium segments from 90 subjects demonstrate that the proposed MMSA can accurately analyze the regional myocardium motion, with accuracies of 79.3% for 4-way motion scoring, 89.0% for abnormality detection, and correlation of 0.943 for estimation of motion score index. This work has great potential for practical assessmentof cardiac motion function. Wufeng Xue, Zejian Chen, Tianfu Wang 0001, Shuo Li 0001, Dong Ni 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Modality alignment contrastive learning for severity assessment of COVID-19 from lung ultrasound and clinical information
Wufeng Xue, Chunyan Cao, Yilian Duan, Haiyan Cao, Jian Wang 0099, Xumin Tao, Zejian Chen, Jinxiang Zhang, Xin Yang 0009, Ruobing Huang, Feixiang Xiang, Manjie You, Mingxing Xie |
Medical Image Anal. | 1 |
| 2021 | Left Ventricle Quantification Challenge: A Comprehensive Comparison and Evaluation of Segmentation and Regression for Mid-Ventricular Short-Axis Cardiac MR DataabstractAutomatic quantification of the left ventricle (LV) from cardiac magnetic resonance (CMR) images plays an important role in making the diagnosis procedure efficient, reliable, and alleviating the laborious reading work for physicians. Considerable efforts have been devoted to LV quantification using different strategies that include segmentation-based (SG) methods and the recent direct regression (DR) methods. Although both SG and DR methods have obtained great success for the task, a systematic platform to benchmark them remains absent because of differences in label information during model learning. In this paper, we conducted an unbiased evaluation and comparison of cardiac LV quantification methods that were submitted to the Left Ventricle Quantification (LVQuan) challenge, which was held in conjunction with the Statistical Atlases and Computational Modeling of the Heart (STACOM) workshop at the MICCAI 2018. The challenge was targeted at the quantification of 1) areas of LV cavity and myocardium, 2) dimensions of the LV cavity, 3) regional wall thicknesses (RWT), and 4) the cardiac phase, from mid-ventricle short-axis CMR images. First, we constructed a public quantification dataset Cardiac-DIG with ground truth labels for both the myocardium mask and these quantification targets across the entire cardiac cycle. Then, the key techniques employed by each submission were described. Next, quantitative validation of these submissions were conducted with the constructed dataset. The evaluation results revealed that both SG and DR methods can offer good LV quantification performance, even though DR methods do not require densely labeled masks for supervision. Among the 12 submissions, the DR method LDAMT offered the best performance, with a mean estimation error of 301 mm2for the two areas, 2.15 mm for the cavity dimensions, 2.03 mm for RWTs, and a 9.5% error rate for the cardiac phase classification. Three of the SG methods also delivered comparable performances. Finally, we discussed the advantages and disadvantages of SG and DR methods, as well as the unsolved problems in automatic cardiac quantification for clinical practice applications. Wufeng Xue, Jiahui Li 0005, Eric Kerfoot, James R. Clough, Ilkay Öksüz, Vicente Grau, Fumin Guo, Matthew Ng, Xiang Li 0001, Quanzheng Li, Lihong Liu, Ilias Grinias, Georgios Tziritas, Angélica Atehortúa, Mireille Garreau, Yeonggul Jang, Alejandro Debus, Enzo Ferrante, Guanyu Yang 0001, Tiancong Hua, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Agent With Warm Start and Adaptive Dynamic Termination for Plane Localization in 3D UltrasoundabstractAccurate standard plane (SP) localization is the fundamental step for prenatal ultrasound (US) diagnosis. Typically, dozens of US SPs are collected to determine the clinical diagnosis. 2D US has to perform scanning for each SP, which is time-consuming and operator-dependent. While 3D US containing multiple SPs in one shot has the inherent advantages of less user-dependency and more efficiency. Automatically locating SP in 3D US is very challenging due to the huge search space and large fetal posture variations. Our previous study proposed a deep reinforcement learning (RL) framework with an alignment module and active termination to localize SPs in 3D US automatically. However, termination of agent search in RL is important and affects the practical deployment. In this study, we enhance our previous RL framework with a newly designed adaptive dynamic termination to enable an early stop for the agent searching, saving at most 67% inference time, thus boosting the accuracy and efficiency of the RL framework at the same time. Besides, we validate the effectiveness and generalizability of our algorithm extensively on our in-house multi-organ datasets containing 433 fetal brain volumes, 519 fetal abdomen volumes, and 683 uterus volumes. Our approach achieves localization error of 2.52mm/10.26°, 2.48mm/10.39°, 2.02mm/10.48°, 2.00mm/14.57°, 2.61mm/9.71°, 3.09mm/9.58°, 1.49mm/7.54°for the transcerebellar, transventricular, transthalamic planes in fetal brain, abdominal plane in fetal abdomen, and mid-sagittal, transverse and coronal planes in uterus, respectively. Experimental results show that our method is general and has the potential to improve the efficiency and standardization of US scanning. Xin Yang 0009, Haoran Dou, Ruobing Huang, Wufeng Xue, Yuhao Huang 0001, Jikuan Qian, Yuanji Zhang, Huanjia Luo, Huizhi Guo, Tianfu Wang 0001, Yi Xiong 0001, Dong Ni 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Auto-weighting for Breast Cancer Classification in Multimodal Ultrasound
Jian Wang 0099, Juzheng Miao, Xin Yang 0009, Rui Li 0038, Guangquan Zhou, Yuhao Huang 0001, Wufeng Xue, Xiaohong Jia 0003, Jianqiao Zhou, Ruobing Huang, Dong Ni 0001 |
MICCAI (6) | 8 |
| 2020 | Temporal-Consistent Segmentation of Echocardiography with Co-learning from Appearance and Shape
Hongrong Wei, Yiqin Cao, Yongjin Zhou 0002, Wufeng Xue, Dong Ni 0001, Shuo Li 0001 |
MICCAI (2) | 5 |
| 2019 | FetusMap: Fetal Pose Estimation in 3D Ultrasound
Xin Yang 0009, Wenlong Shi, Haoran Dou, Jikuan Qian, Yi Wang 0031, Wufeng Xue, Shengli Li 0001, Dong Ni 0001, Pheng-Ann Heng |
MICCAI (5) | 6 |
| 2019 | Agent with Warm Start and Active Termination for Plane Localization in 3D Ultrasound
Haoran Dou, Xin Yang 0009, Jikuan Qian, Wufeng Xue, Xu Wang 0017, Lequan Yu, Yi Xiong 0001, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (5) | 4 |
| 2019 | Arterial Spin Labeling Images Synthesis From sMRI Using Unbalanced Deep Discriminant LearningabstractAdequate medical images are often indispensable in contemporary deep learning-based medical imaging studies, although the acquisition of certain image modalities may be limited due to several issues including high costs and patients issues. However, thanks to recent advances in deep learning techniques, the above tough problem can be substantially alleviated by medical images synthesis, by which various modalities including T1/T2/DTI MRI images, PET images, cardiac ultrasound images, retinal images, and so on, have already been synthesized. Unfortunately, the arterial spin labeling (ASL) image, which is an important fMRI indicator in dementia diseases diagnosis nowadays, has never been comprehensively investigated for the synthesis purpose yet. In this paper, ASL images have been successfully synthesized from structural magnetic resonance images for the first time. Technically, a novel unbalanced deep discriminant learning-based model equipped with new ResNet sub-structures is proposed to realize the synthesis of ASL images from structural magnetic resonance images. The extensive experiments have been conducted. Comprehensive statistical analyses reveal that: 1) this newly introduced model is capable to synthesize ASL images that are similar towards real ones acquired by actual scanning; 2) synthesized ASL images obtained by the new model have demonstrated outstanding performance when undergoing rigorous tests of region-based and voxel-based corrections of partial volume effects, which are essential in ASL images processing; and 3) it is also promising that the diagnosis performance of dementia diseases can be significantly improved with the help of synthesized ASL images obtained by the new model, based on a multi-modal MRI dataset containing 355 demented patients in this paper. Wei Huang 0013, Mingyuan Luo, Xi Liu 0008, Peng Zhang 0005, Huijun Ding, Wufeng Xue, Dong Ni 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Cardiac Motion Scoring with Segment- and Subject-Level Non-local Modeling
Wufeng Xue, Gary Brahm, Stephanie Leung, Olga Shmuilovich, Shuo Li 0001 |
MICCAI (2) | 1 |
| 2018 | Full left ventricle quantification via deep multitask relationships learning
Wufeng Xue, Gary Brahm, Sachin Pandey, Stephanie Leung, Shuo Li 0001 |
Medical Image Anal. | 1 |
| 2017 | Full Quantification of Left Ventricle via Deep Multitask Learning Network Respecting Intra- and Inter-Task Relatedness
Wufeng Xue, Andrea Lum, Ashley Mercado, Mark Landis, James Warrington, Shuo Li 0001 |
MICCAI (3) | 1 |
| 2017 | Direct Multitype Cardiac Indices Estimation via Joint Representation and Regression LearningabstractCardiac indices estimation is of great importance during identification and diagnosis of cardiac disease in clinical routine. However, estimation of multitype cardiac indices with consistently reliable and high accuracy is still a great challenge due to the high variability of cardiac structures and the complexity of temporal dynamics in cardiac MR sequences. While efforts have been devoted into cardiac volumes estimation through feature engineering followed by a independent regression model, these methods suffer from the vulnerable feature representation and incompatible regression model. In this paper, we propose a semi-automated method for multitype cardiac indices estimation. After the manual labeling of two landmarks for ROI cropping, an integrated deep neural network Indices-Net is designed to jointly learn the representation and regression models. It comprises two tightly-coupled networks, such as a deep convolution autoencoder for cardiac image representation, and a multiple output convolution neural network for indices regression. Joint learning of the two networks effectively enhances the expressiveness of image representation with respect to cardiac indices, and the compatibility between image representation and indices regression, thus leading to accurate and reliable estimations for all the cardiac indices. When applied with five-fold cross validation on MR images of 145 subjects, Indices-Net achieves consistently low estimation error for LV wall thicknesses (1.44 ± 0.71 mm) and areas of cavity and myocardium (204 ± 133 mm2). It outperforms, with significant error reductions, segmentation method (55.1% and 17.4%), and two-phase direct volume-only methods (12.7% and 14.6%) for wall thicknesses and areas, respectively. These advantages endow the proposed method a great potential in clinical cardiac function assessment. Wufeng Xue, Ali Islam, Mousumi Bhaduri, Shuo Li 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Blind Image Quality Assessment Using Joint Statistics of Gradient Magnitude and Laplacian FeaturesabstractBlind image quality assessment (BIQA) aims to evaluate the perceptual quality of a distorted image without information regarding its reference image. Existing BIQA models usually predict the image quality by analyzing the image statistics in some transformed domain, e.g., in the discrete cosine transform domain or wavelet domain. Though great progress has been made in recent years, BIQA is still a very challenging task due to the lack of a reference image. Considering that image local contrast features convey important structural information that is closely related to image perceptual quality, we propose a novel BIQA model that utilizes the joint statistics of two types of commonly used local contrast features: 1) the gradient magnitude (GM) map and 2) the Laplacian of Gaussian (LOG) response. We employ an adaptive procedure to jointly normalize the GM and LOG features, and show that the joint statistics of normalized GM and LOG features have desirable properties for the BIQA task. The proposed model is extensively evaluated on three large-scale benchmark databases, and shown to deliver highly competitive performance with state-of-the-art BIQA models, as well as with some well-known full reference image quality assessment models. Wufeng Xue, Xuanqin Mou, Lei Zhang 0006, Alan C. Bovik, Xiangchu Feng |
IEEE Trans. Image Process. | 1 |
| 2014 | Gradient Magnitude Similarity Deviation: A Highly Efficient Perceptual Image Quality IndexabstractIt is an important task to faithfully evaluate the perceptual quality of output images in many applications, such as image compression, image restoration, and multimedia streaming. A good image quality assessment (IQA) model should not only deliver high quality prediction accuracy, but also be computationally efficient. The efficiency of IQA metrics is becoming particularly important due to the increasing proliferation of high-volume visual data in high-speed networks. We present a new effective and efficient IQA model, called gradient magnitude similarity deviation (GMSD). The image gradients are sensitive to image distortions, while different local structures in a distorted image suffer different degrees of degradations. This motivates us to explore the use of global variation of gradient based local quality map for overall image quality prediction. We find that the pixel-wise gradient magnitude similarity (GMS) between the reference and distorted images combined with a novel pooling strategy-the standard deviation of the GMS map-can predict accurately perceptual image quality. The resulting GMSD algorithm is much faster than most state-of-the-art IQA methods, and delivers highly competitive prediction accuracy. MATLAB source code of GMSD can be downloaded at http://www4.comp.polyu.edu.hk/~cslzhang/IQA/GMSD/GMSD.htm. Wufeng Xue, Lei Zhang 0006, Xuanqin Mou, Alan C. Bovik |
IEEE Trans. Image Process. | 1 |
| 2013 | Learning without Human Scores for Blind Image Quality AssessmentabstractGeneral purpose blind image quality assessment (BIQA) has been recently attracting significant attention in the fields of image processing, vision and machine learning. State-of-the-art BIQA methods usually learn to evaluate the image quality by regression from human subjective scores of the training samples. However, these methods need a large number of human scored images for training, and lack an explicit explanation of how the image quality is affected by image local features. An interesting question is then: can we learn for effective BIQA without using human scored images? This paper makes a good effort to answer this question. We partition the distorted images into overlapped patches, and use a percentile pooling strategy to estimate the local quality of each patch. Then a quality-aware clustering (QAC) method is proposed to learn a set of centroids on each quality level. These centroids are then used as a codebook to infer the quality of each patch in a given image, and subsequently a perceptual quality score of the whole image can be obtained. The proposed QAC based BIQA method is simple yet effective. It not only has comparable accuracy to those methods using human scored images in learning, but also has merits such as high linearity to human perception of image quality, real-time implementation and availability of image local quality map. Wufeng Xue, Lei Zhang 0006, Xuanqin Mou |
CVPR | 1 |
| 2013 | Perceptual Fidelity Aware Mean Squared ErrorabstractHow to measure the perceptual quality of natural images is an important problem in low level vision. It is known that the Mean Squared Error (MSE) is not an effective index to describe the perceptual fidelity of images. Numerous perceptual fidelity indices have been developed, while the representatives include the Structural SIMilarity (SSIM) index and its variants. However, most of those perceptual measures are nonlinear, and they cannot be easily dopted as an objective function to minimize in various low level vision tasks. Can MSE be perceptual fidelity aware after some minor adaptation? In this paper we propose a simple framework to enhance the perceptual fidelity awareness of MSE by introducing an l2-norm structural error term to it. Such a Structural MSE (SMSE) can lead to very competitive image quality assessment (IQA) results. More surprisingly, we show that by using certain structure extractors, SMSE can be further turned into a Gaussian smoothed MSE (i.e., the Euclidean distance between the original and distorted images after Gaussian smooth filtering), which is much simpler to calculate but achieves rather better IQA performance than SSIM. The so called Perceptual-fidelity Aware MSE (PAMSE) can have great potentials in applications such as perceptual image coding and perceptual image restoration. Wufeng Xue, Xuanqin Mou, Lei Zhang 0006, Xiangchu Feng |
ICCV | 1 |
| 2013 | Edge Strength Similarity for Image Quality AssessmentabstractThe objective image quality assessment aims to model the perceptual fidelity of semantic information between two images. In this letter, we assume that the semantic information of images is fully represented by edge-strength of each pixel and propose an edge-strength-similarity-based image quality metric (ESSIM). Through investigating the characteristics of the edge in images, we define the edge-strength to take both anisotropic regularity and irregularity of the edge into account. The proposed ESSIM is considerably simple, however, it can achieve slightly better performance than the state-of-the-art image quality metrics as evaluated on six subject-rated image databases. Xuande Zhang, Xiangchu Feng, Weiwei Wang 0005, Wufeng Xue |
IEEE Signal Process. Lett. | 4 |
| 2011 | An image quality assessment metric based on Non-shift EdgeabstractIn this paper, we propose a novel metric for image quality assessment based on the ratio of Non-shift Edge (rNSE), whose elegance lies in succinctness and effectiveness. In this metric, an image is filtered by the LOG operator, who acts like the classical receptive field, and the edge points are detected as the zero-crossings of the filtered image. Then the binary Non-shift Edge (NSE) map is derived to represent the strong edge structure remained in the distorted image. The perceptual quality is calculated by the ratio of NSE. The performance of rNSE in the scale-threshold plane shows similar frequency and threshold selectivity. Comparing with the existing well-designed metrics, the proposed rNSE performs equivalently in accuracy and consistency. Wufeng Xue, Xuanqin Mou |
ICIP | 1 |