EDBT 2026 Demo / reviewers in the wild / expert
Xiaowei Xu 0004
dblp:181/2733-4
· DBLP profile ↗
63ranked-venue papers
11as first author
43since 2021 · last 2025
0000-0002-1046-6379ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 3 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 15 since 2021Systems, architecture and hardware · 17 · 6 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MGS-EP: Mask Guided Segmentation of Regional Wall with Expert Prior Pre-DecouplingabstractCoronary artery disease (CAD) is a type of heart disease, where echocardiography can be used in the diagnosis. Due to the time-consuming and non-reproducibility of manual assessment, automatic evaluation methods are increasingly required where the regional wall segmentation is a crucial step. Currently, most studies prioritize designing sophisticated networks, yet overlooking the fact that the poor segmentation performance comes from the inherent fuzziness and low contrast of echocardiography. In this paper, a framework named MGS-EP is proposed. Inspired by clinical annotation practices where experts infer missing wall structures using anatomical knowledge, this expert intuition is formalized as the Expert Prior (EP). The proposed MGS-EP integrates EP constraints to compensate for image degradation, which consists of a pre-decoupler and a mask-guided segmentation (MGS) network. Where the pre-decoupler first models the EP, generating pseudo masks that encapsulate both complete topology and approximate spatial localization of the regional walls. These pseudo masks are subsequently concatenated with raw echocardiography to form a composite input for the MGS network, thereby enabling regional wall segmentation with topological integrity. The concatenation of pseudo masks with original echocardiography as input to the MGS network serves two purposes: mitigating regional walls' contour degradation due to fuzziness while embedding EP topological constraints. Experimental results demonstrate that, compared to the baseline nnU-Net, the MGS-EP enables topologically continuous segmentation, and can achieve an average improvement of 7.76% in Dice and a reduction of$\mathbf{1 1. 6 1}$pixels in Hausdorff Distance. Dawei Li 0012, Tienan Chen, Yongqiang Cui, Xiaowei Xu 0004, Yiyu Shi 0001 |
BIBM | 4 |
| 2025 | Attention Ensemble based Whole Heart and Great Vessel Segmentation for Surgical Planning of Total Anomalous Pulmonary Venous ConnectionsabstractTotal anomalous pulmonary venous connections (TAPVC) is a serious congenital heart disease, and surgery is the main treatment for TAPVC patients. Surgical repair of TAPVC is challenging, and recently, 3D visualization including 3D printing and virtual reality has been adopted in clinical practice to mitigate this problem. However, the whole heart and great vessels for 3D printing is manually segmented by experts, which is time-consuming, subject, and costly. Though whole heart and great vessel segmentation have been a hot topic in the community for decades, the existing works focus on general diseases, and There is no specific attention on TAPVC. In this paper, we collect the first whole heart and great vessel segmentation dataset for surgical planning of TAPVC. The dataset contains 92 samples, which is annotated by experienced radiologist with about 1 hour. Unlike existing works which doesn’t consider pulmonary veins and left atrium separately, 10 anatomies including left ventricle, right ventricle, left atrium, right atrium, aorta, pulmonary artery, myocardium, superior vena cava, inferior vena cava, and pulmonary veins are annotated. Then, we proposed attention ensemble based network, AE-Network for whole heart and great vessel segmentation dataset for surgical planning of TAPVC. AE-Network adopts a three-branch structure, where each branch network is based on the U-shaped network with attention mechanisms. For each of the three branch, spatial attention, channel attention, and pixel attention are used to fully exploit the context. Dynamic segmentation result fusion (DSRF) is then used to combine all the three branches to achieve accurate and robust segmentation. Experimental results show that AE-Network achieved an average Dice coefficient of 80.60%, significantly outperforming existing works. Our dataset and code will be open-sourced once the paper is accepted. Erlei Zhang, Jinglei Li, Xiaowei Xu 0004 |
IJCNN | 4 |
| 2025 | Quantization-based deep diversified ensemble for medical image segmentation
Qi Wang 0044, Yanchun Zhang, Weihong Han, Yangyang Mei, Yiyu Shi 0001, Jian Zhuang, Meiping Huang, Xiaowei Xu 0004 |
Eng. Appl. Artif. Intell. | 10 |
| 2025 | Domain knowledge based comprehensive segmentation of Type-A aortic dissection with clinically-oriented evaluation
Hailong Qiu, Meiping Huang, Jian Zhuang, Qing Lu 0001, Yiyu Shi 0001, Xiaomeng Li 0001, Wen Xie 0008, Guang Tong, Xiaowei Xu 0004 |
Medical Image Anal. | 10 |
| 2025 | Constrained multi-scale dense connections for biomedical image segmentation
Yanchun Zhang, Hailong Qiu, Xiaomeng Li 0001, Shanfeng Zhu, Meiping Huang, Jian Zhuang, Yiyu Shi 0001, Xiaowei Xu 0004 |
Pattern Recognit. | 10 |
| 2025 | A Benchmark Framework for the Right Atrium Cavity Segmentation From LGE-MRIsabstractThe right atrium (RA) is critical for cardiac hemodynamics but is often overlooked in clinical diagnostics. This study presents a benchmark framework for RA cavity segmentation from late gadolinium-enhanced magnetic resonance imaging (LGE-MRIs), leveraging a two-stage strategy and a novel 3D deep learning network, RASnet. The architecture addresses challenges in class imbalance and anatomical variability by incorporating multi-path input, multi-scale feature fusion modules, Vision Transformers, context interaction mechanisms, and deep supervision. Evaluated on datasets comprising 354 LGE-MRIs, RASnet achieves SOTA performance with a Dice score of 92.19% on a primary dataset and demonstrates robust generalizability on an independent dataset. The proposed framework establishes a benchmark for RA cavity segmentation, enabling accurate and efficient analysis for cardiac imaging applications. Open-source code (https://github.com/zjinw/RAS) and data (https://zenodo.org/records/15524472) are provided to facilitate further research and clinical adoption. Jieyun Bai, Jinwen Zhu, Zhiting Chen, Ziduo Yang, Yaosheng Lu, Lei Li 0020, Qince Li, Wei Wang 0169, Henggui Zhang, Kuanquan Wang, Jichao Zhao, Hua Lu 0022, Suining Li, Xiaoshen Zhang, Xiaowei Xu 0004, Yanfeng Tian, Víctor M. Campello, Karim Lekadir |
IEEE Trans. Medical Imaging | 18 |
| 2025 | Dynamic Subcluster-Aware Network for Few-Shot Skin Disease ClassificationabstractThis article addresses the problem of few-shot skin disease classification by introducing a novel approach called the subcluster-aware network (SCAN) that enhances accuracy in diagnosing rare skin diseases. The key insight motivating the design of SCAN is the observation that skin disease images within a class often exhibit multiple subclusters, characterized by distinct variations in appearance. To improve the performance of few-shot learning (FSL), we focus on learning a high-quality feature encoder that captures the unique subclustered representations within each disease class, enabling better characterization of feature distributions. Specifically, SCAN follows a dual-branch framework, where the first branch learns classwise features to distinguish different skin diseases, and the second branch aims to learn features, which can effectively partition each class into several groups so as to preserve the subclustered structure within each class. To achieve the objective of the second branch, we present a cluster loss to learn image similarities via unsupervised clustering. To ensure that the samples in each subcluster are from the same class, we further design a purity loss to refine the unsupervised clustering results. We evaluate the proposed approach on two public datasets for few-shot skin disease classification. The experimental results validate that our framework outperforms the state-of-the-art methods by around 2%-5% in terms of sensitivity, specificity, accuracy, and F1-score on the SD-198 and Derm7pt datasets. Shuhan Li, Xiaomeng Li 0001, Xiaowei Xu 0004, Kwang-Ting Cheng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | CardiacNet: Learning to Reconstruct Abnormalities for Cardiac Disease Assessment from Echocardiogram Videos
Jiewen Yang, Yiqun Lin, Bin Pu, Jiarong Guo, Xiaowei Xu 0004, Xiaomeng Li 0001 |
ECCV (23) | 5 |
| 2024 | Contrastive Learning with Synthetic Positives
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Yiyu Shi 0001 |
ECCV (37) | 4 |
| 2024 | Breast Ultrasound Computer-Aided Diagnosis Using Structure-Aware Triplet Path NetworksabstractBreast ultrasound (BUS) is an effective imaging modality for breast cancer diagnosis. The structural characteristics of breast lesions play an important role in computer-aided diagnosis. In this paper, a novel structure-aware triplet path network (SATPN) was designed to integrate classification and image reconstruction tasks to achieve accurate diagnosis on BUS images. Specifically, we enhanced clinically-approved structure characteristics of breast lesion by converting original BUS images to BI-RADS-oriented feature maps (BFMs) with a distance-transformation coupled Gaussian filter. Then, the converted BFMs were used as the inputs of the SATPN, which performed a supervised lesion classification task and two separate unsupervised stacked convolutional auto-encoder tasks for benign and malignant image reconstruction. We trained the SATPN with an alternative learning strategy by balancing image reconstruction error and classification label prediction error. The lesion label was determined by weighted voting of reconstruction error and label prediction error. We compared the performance of the SATPN with five deep learning methods using the original images and BFMs as inputs. Experimental results on two BUS datasets showed that SATPN performed the best among the six networks, with classification accuracy around 96%. These findings indicate that SATPN is promising for effective ultrasound computer-aided diagnosis of breast lesions. Erlei Zhang, Xiaowei Xu 0004, Zhicheng Zhang 0005, Jinglei Li |
ICASSP | 3 |
| 2024 | Fusion of Machine Learning and Deep Neural Networks for Pulmonary Arteries and Veins Segmentation in Lung Cancer Surgery Planning
Limin Zheng, Xiaowei Xu 0004 |
ICPR (12) | 6 |
| 2024 | Data-Algorithm-Architecture Co-Optimization for Fair Neural Networks on Skin Lesion Dataset
Junhuan Yang, James Alaina, Xiaowei Xu 0004, Yiyu Shi 0001, Jingtong Hu, Weiwen Jiang, Lei Yang 0018 |
MICCAI (10) | 5 |
| 2024 | HOCM-Net: 3D coarse-to-fine structural prior fusion based segmentation network for the surgical planning of hypertrophic obstructive cardiomyopathy
Hailong Qiu, Yanchun Zhang, Weihong Han, Yiyu Shi 0001, Meiping Huang, Jian Zhuang, Huiming Guo, Xiaowei Xu 0004 |
Expert Syst. Appl. | 12 |
| 2024 | TinyML Design Contest for Life-Threatening Ventricular Arrhythmia DetectionabstractThe first ACM/IEEE TinyML Design Contest (TDC) held at the 41st International Conference on Computer-Aided Design (ICCAD) in 2022 is a challenging, multimonth, research and development competition. TDC’22 focuses on real-world medical problems that require the innovation and implementation of artificial intelligence/machine learning (AI/ML) algorithms on implantable devices. The challenge problem of TDC’22 is to develop a novel AI/ML-based real-time detection algorithm for life-threatening ventricular arrhythmia (VA) over low-power microcontrollers utilized in implantable cardioverter-defibrillators (ICDs). The dataset contains more than 38000 5-s intracardiac electrograms (IEGMs) segments over eight different types of rhythm from 90 subjects. The dedicated hardware platform is NUCLEO-L432KC manufactured by STMicroelectronics. TDC’22, which is open to multiperson teams world-wide, attracted more than 150 teams from over 50 organizations. This article first presents the medical problem, dataset, and evaluation procedure in detail. It further demonstrates and discusses the designs developed by the leading teams as well as representative results. This article concludes with the direction of improvement for the future TinyML design for health monitoring applications. Zhenge Jia, Dawei Li 0012, Liqi Liao, Xiaowei Xu 0004, Lichuan Ping, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | Automatic Segmentation of Aortic and Mitral Valves for Heart Surgical Planning of Hypertrophic Obstructive Cardiomyopathy
Limin Zheng, Lu Qing, Jian Zhuang, Xiaowei Xu 0004 |
ACML | 6 |
| 2023 | Quantization through Search: A Novel Scheme to Quantize Convolutional Neural Networks in Finite Weight SpaceabstractQuantization has become an essential technique in compressing deep neural networks for deployment onto resource-constrained hardware. It is noticed that, the hardware efficiency of implementing quantized networks is highly coupled with the actual values to be quantized into, and therefore, with given bit widths, we can smartly choose a value space to further boost the hardware efficiency. For example, using weights of only integer powers of two, multiplication can be fulfilled by bit operations. Under such circumstances, however, existing quantization-aware training methods are either not suitable to apply or unable to unleash the expressiveness of very low bit-widths. For the best hardware efficiency, we revisit the quantization of convolutional neural networks and propose to address the training process from a weight-searching angle, as opposed to optimizing the quantizer functions as in existing works. Extensive experiments on CIFAR10 and ImageNet classification tasks are examined with implementations onto well-established CNN architectures, such as ResNet, VGG, and MobileNet, etc. It is shown the proposed method can achieve a lower accuracy loss than the state of arts, and/or improving implementation efficiency by using hardware-friendly weight values at the same time. Qing Lu 0001, Weiwen Jiang, Xiaowei Xu 0004, Jingtong Hu, Yiyu Shi 0001 |
ASP-DAC | 3 |
| 2023 | Enhance Regional Wall Segmentation by Style Transfer for Regional Wall Motion Assessment
Yiyu Shi 0001, Jian Zhuang, Meiping Huang, Hongwen Fei, Boyang Li 0003, Qing Lu 0001, Erlei Zhang, Xiaowei Xu 0004 |
BMVC | 10 |
| 2023 | GraphEcho: Graph-Driven Unsupervised Domain Adaptation for Echocardiogram Video SegmentationabstractEchocardiogram video segmentation plays an important role in cardiac disease diagnosis. This paper studies the unsupervised domain adaption (UDA) for echocardiogram video segmentation, where the goal is to generalize the model trained on the source domain to other unlabelled target domains. Existing UDA segmentation methods are not suitable for this task because they do not model local information and the cyclical consistency of heartbeat. In this paper, we introduce a newly collected CardiacUDA dataset and a novel GraphEcho method for cardiac structure segmentation. Our GraphEcho comprises two innovative modules, the Spatial-wise Cross-domain Graph Matching (SCGM) and the Temporal Cycle Consistency (TCC) module, which utilize prior knowledge of echocardiogram videos, i.e., consistent cardiac structure across patients and centers and the heartbeat cyclical consistency, respectively. These two modules can better align global and local features from source and target domains, leading to improved UDA segmentation results. Experimental results showed that our GraphEcho outperforms existing state-of-the-art UDA segmentation methods. Our collected dataset and code will be publicly released upon acceptance. This work will lay a new and solid cornerstone for cardiac structure segmentation from echocardiogram videos. Code and dataset are available at : https://github.com/xmedlab/GraphEcho Jiewen Yang, Xinpeng Ding, Xiaowei Xu 0004, Xiaomeng Li 0001 |
ICCV | 4 |
| 2023 | SEDSkill: Surgical Events Driven Method for Skill Assessment from Thoracoscopic Surgical Videos
Xinpeng Ding, Xiaowei Xu 0004, Xiaomeng Li 0001 |
MICCAI (9) | 2 |
| 2023 | MPBD-LSTM: A Predictive Model for Colorectal Liver Metastases Using Time Series Multi-phase Contrast-Enhanced CT Scans
Weixiang Weng, Xiaowei Xu 0004, Yiyu Shi 0001 |
MICCAI (6) | 4 |
| 2023 | Additional Positive Enables Better Representation Learning for Medical Images
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (1) | 4 |
| 2023 | GL-Fusion: Global-Local Fusion Network for Multi-view Echocardiogram Video Segmentation
Jiewen Yang, Xinpeng Ding, Xiaowei Xu 0004, Xiaomeng Li 0001 |
MICCAI (4) | 4 |
| 2023 | Hardware-aware neural architecture search for stochastic computing-based neural networks on tiny devices
Yuhong Song, Edwin H.-M. Sha, Qingfeng Zhuge, Rui Xu 0013, Xiaowei Xu 0004, Bingzhe Li, Lei Yang 0018 |
J. Syst. Archit. | 5 |
| 2023 | A clinically applicable AI system for diagnosis of congenital heart diseases based on computed tomography images
Xiaowei Xu 0004, Qianjun Jia, Haiyun Yuan, Hailong Qiu, Yuhao Dong, Wen Xie 0008, Zeyang Yao, Zhiqaing Nie, Xiaomeng Li 0001, Yiyu Shi 0001, James Zou 0001, Meiping Huang, Jian Zhuang |
Medical Image Anal. | 1 |
| 2023 | Less Is More: Surgical Phase Recognition From Timestamp SupervisionabstractSurgical phase recognition is a fundamental task in computer-assisted surgery systems. Most existing works are under the supervision of expensive and time-consuming full annotations, which require the surgeons to repeat watching videos to find the precise start and end time for a surgical phase. In this paper, we introduce timestamp supervision for surgical phase recognition to train the models with timestamp annotations, where the surgeons are asked to identify only a single timestamp within the temporal boundary of a phase. This annotation can significantly reduce the manual annotation cost compared to the full annotations. To make full use of such timestamp supervisions, we propose a novel method called uncertainty-aware temporal diffusion (UATD) to generate trustworthy pseudo labels for training. Our proposed UATD is motivated by the property of surgical videos, i.e., the phases are long events consisting of consecutive frames. To be specific, UATD diffuses the single labelled timestamp to its corresponding high confident (i.e., low uncertainty) neighbour frames in an iterative way. Our study uncovers unique insights of surgical phase recognition with timestamp supervision: 1) timestamp annotation can reduce 74% annotation time compared with the full annotation, and surgeons tend to annotate those timestamps near the middle of phases; 2) extensive experiments demonstrate that our method can achieve competitive results compared with full supervision methods, while reducing manual annotation costs; 3) less is more in surgical phase recognition, i.e., less but discriminative pseudo labels outperform full but containing ambiguous frames; 4) the proposed UATD can be used as a plug-and-play method to clean ambiguous labels near boundaries between phases, and improve the performance of the current surgical phase recognition methods. Code and annotations obtained from surgeons are available at https://github.com/xmed-lab/TimeStamp-Surgical. Xinpeng Ding, Xinjian Yan, Wei Zhao 0029, Jian Zhuang, Xiaowei Xu 0004, Xiaomeng Li 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2022 | ImageALCAPA: A 3D Computed Tomography Image Dataset for Automatic Segmentation of Anomalous Left Coronary Artery from Pulmonary ArteryabstractAnomalous left coronary artery from pulmonary artery (ALCAPA) is a serious cardiac anomaly, and surgical repair is the main treatment for ALCAPA patients in clinical practice. Recently, 3D printing has been widely adopted in the surgical planning of ALCAPA, which can give surgeons an intuitive structure of the heart especially the coronary arteries. However, before 3D printing is conducted, experienced radiologists need to manually segment the coronary arteries on computed tomography angiography (CTA) images, which is time-consuming, tedious and biased. On the other hand, automatic coronary artery segmentation with normal structures has been extensively studied in the community, but cannot be effectively applied to ALCAPA due to the significant variation of coronary artery structure in ALCAPA. In this paper, we propose ImageALCAPA, the first 3D CTA image dataset of ALCAPA. The proposed dataset contains 30 ALCAPA CTA images, which is of decent size compared with existing medical imaging datasets. We further propose a baseline method that performs multi-task 2D- 3D ensemble for automatic segmentation of ALCAPA. It is shown by experiment that our baseline method outperforms popular existing works on coronary artery segmentation. However, as the highest average Dice Similarity Coefficient of coronary arteries is merely 65%, there is still much room for improvement. To facilitate further research on this challenging problem, our dataset and codes are released to the public [1]. An Zeng, Chenxi Mi, Dan Pan 0001, Qing Lu 0001, Xiaowei Xu 0004 |
BIBM | 5 |
| 2022 | RT-DNAS: Real-Time Constrained Differentiable Neural Architecture Search for 3D Cardiac Cine MRI Segmentation
Qing Lu 0001, Xiaowei Xu 0004, Shunjie Dong, Cong Hao, Lei Yang 0018, Cheng Zhuo, Yiyu Shi 0001 |
MICCAI (5) | 2 |
| 2022 | FairPrune: Achieving Fairness Through Pruning for Dermatological Disease Diagnosis
Yawen Wu, Dewen Zeng, Xiaowei Xu 0004, Yiyu Shi 0001, Jingtong Hu |
MICCAI (1) | 3 |
| 2022 | Combining multi-view ensemble and surrogate lagrangian relaxation for real-time 3D biomedical image segmentation on the edge
Shanglin Zhou, Xiaowei Xu 0004, Mikhail A. Bragin |
Neurocomputing | 2 |
| 2022 | A Quasi-digital QPSK Modulator Design for Biomedical DevicesabstractFor the biomedical transceiver, the data transmission is often asymmetric. At the downlink, the transceiver only needs to receive a simple command to control the operation of the external device, and the receiving data rate is low, about hundreds of Kb/s. However, data collected by external devices such as temperature sensors, pressure sensors, or cameras are often very large, which results in a transmitting data rate of several Mb/s. Therefore, a high energy-efficient modulator is needed. Compared with conventional digital modulator, analog modulator circuits have demonstrated superior energy efficiency at high data rates. This article presents a quasi-digital quadrature phase-shift keying (QPSK) modulator design realized by pure analog circuits which follows a logic design flow. The simulation results show that the system can generate a stable carrier of 64 MHz that meets intra-body communications (IBCs) requirements with a data transmission rate of 10 Mb/s. When the signal-to-noise ratios (SNRs) of the Gaussian channel is 14 dB, it can still maintain a bit error rate (BER) below 10 4 . Dawei Li 0012, Yang Zhou 0030, Shaopin Chen, Xiaowei Xu 0004 |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2022 | VisualNet: An End-to-End Human Visual System Inspired Framework to Reduce Inference Latency of Deep Neural NetworksabstractAcceleration of deep neural network (DNN) inference has gained increasing attention recently with the wide adoption of DNNs for practical applications. For computer vision tasks where inputs are images, existing works mostly focus on improving the throughput of inference for multiple images. However, in many real-time applications, it is critical to reduce the latency of a single image inference, which is more complicated than improving the throughput because of the inherent data dependencies. On the other hand, from human brain's perspective, the complexity in our visual surroundings is first encoded as a pattern of light on a two dimensional array of photoreceptors, with little direct resemblance to the original input or the ultimate percept. Within just a few hundred microns of retinal thickness, this initial signal encoded by our photoreceptors must be transformed into an adequate representation of the entire visual scene. Inspired by how the retina helps human brain incept new information efficiently, we present an end-to-end structured framework built using any existing convolutional neural network (CNN) as the backbone. The proposed framework, called VisualNet, can create task parallelism for the backbone during the inference of a single image. Experiments using a number of neural networks for the ImageNet classification task and the CIFAR-10 classification task on GPUs and CPUs show that the proposed VisualNet reduces the latency of the regular network it builds on by up to 80.6% when both are fully parallelized with state-of-the-art acceleration libraries. At the same time, VisualNet can achieve similar or slightly higher accuracy. Jinjun Xiong, Song Bian 0001, Zheyu Yan, Meiping Huang, Jian Zhuang, Takashi Sato 0001, Xiaowei Xu 0004, Yiyu Shi 0001 |
IEEE Trans. Computers | 9 |
| 2021 | C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion TransferabstractHuman video motion transfer (HVMT) aims to synthesize videos that one person imitates other persons' actions. Although existing GAN-based HVMT methods have achieved great success, they either fail to preserve appearance details due to the loss of spatial consistency between synthesized and exemplary images, or generate incoherent video results due to the lack of temporal consistency among video frames. In this paper, we propose Coarse-to-Fine Flow Warping Network (C2F-FWN) for spatial-temporal consistent HVMT. Particularly, C2F-FWN utilizes coarse-to-fine flow warping and Layout-Constrained Deformable Convolution (LC-DConv) to improve spatial consistency, and employs Flow Temporal Consistency (FTC) Loss to enhance temporal consistency. In addition, provided with multi-source appearance inputs, C2F-FWN can support appearance attribute editing with great flexibility and efficiency. Besides public datasets, we also collected a large-scale HVMT dataset named SoloDance for evaluation. Extensive experiments conducted on our SoloDance dataset and the iPER dataset show that our approach outperforms state-of-art HVMT methods in terms of both spatial and temporal consistency. Source code and the SoloDance dataset are available at https://github.com/wswdx/C2F-FWN. Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang |
AAAI | 2 |
| 2021 | "One-Shot" Reduction of Additive Artifacts in Medical ImagesabstractMedical images may contain various types of artifacts with different patterns and mixtures, which depend on many factors such as scan setting, machine condition, patients’ characteristics, surrounding environment, etc. However, existing deep-learning-based artifact reduction methods are restricted by their training set with specific predetermined artifact types and patterns. As such, they have limited clinical adoption. In this paper, we introduce One-Shot medical image Artifact Reduction (OSAR), which exploits the power of deep learning but without using pre-trained general networks. Specifically, we train a light-weight image-specific artifact reduction network using data synthesized from the input image at test-time. Without requiring any prior large training data set, OSAR can work with almost any medical images that contain varying additive artifacts which are not in any existing data sets. In addition, Computed Tomography (CT) and Magnetic Resonance Imaging (MRI) are used as vehicles and show that the proposed method can reduce artifacts better than state-of-the-art both qualitatively and quantitatively using shorter test time. Yen-Jung Chang, Shao-Cheng Wen, Xiaowei Xu 0004, Meiping Huang, Haiyun Yuan, Jian Zhuang, Yiyu Shi 0001, Tsung-Yi Ho |
BIBM | 4 |
| 2021 | Invited: Hardware-aware Real-time Myocardial Segmentation Quality Control in Contrast EchocardiographyabstractAutomatic myocardial segmentation of contrast echocardio-graphy has shown great potential in the quantification of myocardial perfusion parameters. Segmentation quality control is an important step to ensure the accuracy of segmentation results for quality research as well as its clinical application. Usually, the segmentation quality control happens after the data acquisition. At the data acquisition time, the operator could not know the quality of the segmentation results. On-the-fly segmentation quality control could help the operator to adjust the ultrasound probe or retake data if the quality is unsatisfied, which can greatly reduce the effort of time-consuming manual correction. However, it is infeasible to deploy state-of-the-art DNN-based models because the segmentation module and quality control module must fit in the limited hardware resource on the ultrasound machine while satisfying strict latency constraints. In this paper, we propose a hardware-aware neural architecture search framework for automatic myocardial segmentation and quality control of contrast echocardiography. We explicitly incorporate the hardware latency as a regularization term into the loss function during training. The proposed method searches the best neural network architecture for the segmentation module and quality prediction module with strict latency. Dewen Zeng, Yukun Ding, Haiyun Yuan, Meiping Huang, Xiaowei Xu 0004, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
DAC | 5 |
| 2021 | Pyramid U-Net for Retinal Vessel SegmentationabstractRetinal blood vessel can assist doctors in diagnosis of eyerelated diseases such as diabetes and hypertension, and its segmentation is particularly important for automatic retinal image analysis. However, it is challenging to segment these vessels structures, especially the thin capillaries from the color retinal image due to low contrast and ambiguousness. In this paper, we propose pyramid U-Net for accurate retinal vessel segmentation. In pyramid U-Net, the proposed pyramid-scale aggregation blocks (PSABs) are employed in both the encoder and decoder to aggregate features at higher, current and lower levels. In this way, coarse-to-fine context information is shared and aggregated in each block thus to improve the location of capillaries. To further improve performance, two optimizations including pyramid inputs enhancement and deep pyramid supervision are applied to PSABs in the encoder and decoder, respectively. For PSABs in the encoder, scaled input images are added as extra inputs. While for PSABs in the decoder, scaled intermediate outputs are supervised by the scaled segmentation labels. Extensive evaluations show that our pyramid U-Net outperforms the current state-of-the-art methods on the public DRIVE and CHASE-DB1 datasets. Yanchun Zhang, Xiaowei Xu 0004 |
ICASSP | 3 |
| 2021 | ObjectAug: Object-level Data Augmentation for Semantic Image SegmentationabstractEffective training of deep neural networks (DNNs) usually requires labeling a large dataset, which is time and labor intensive. Recently, various data augmentation strategies like regional dropout and mix strategies have been proposed, which are effective as the augmented dataset can guide the model to attend on less discriminative parts. However, these strategies operate only at the image level, where the objects and the background are coupled. Thus, the boundaries are not well augmented due to the fixed semantic scenario. In this paper, we propose ObjectAug to perform object-level augmentation for semantic image segmentation. Our method first decouples the image into individual objects and the background using semantic labels. Second, each object is augmented individually with commonly used augmentation methods (e.g., scaling, shifting, and rotation). Third, the pixel artifacts brought by object augmentation are further restored using image inpainting. Finally, the augmented objects and background are assembled as an augmented image. In this way, the boundaries can be fully explored in the various semantic scenarios. In addition, ObjectAug can support category-aware augmentation that gives various possibilities to objects in each category, and can be easily combined with existing image-level augmentation methods to further boost the performance. Comprehensive experiments are conducted on both natural image and medical image datasets. Experiment results demonstrate that our ObjectAug can effectively improve segmentation performance. Yanchun Zhang, Xiaowei Xu 0004 |
IJCNN | 3 |
| 2021 | Semi-supervised Contrastive Learning for Label-Efficient Medical Image Segmentation
Xinrong Hu, Dewen Zeng, Xiaowei Xu 0004, Yiyu Shi 0001 |
MICCAI (2) | 3 |
| 2021 | EchoCP: An Echocardiography Dataset in Contrast Transthoracic Echocardiography for Patent Foramen Ovale Diagnosis
Zhihe Li, Meiping Huang, Jian Zhuang, Shanshan Bi, Yiyu Shi 0001, Hongwen Fei, Xiaowei Xu 0004 |
MICCAI (6) | 9 |
| 2021 | Positional Contrastive Learning for Volumetric Medical Image Segmentation
Dewen Zeng, Yawen Wu, Xinrong Hu, Xiaowei Xu 0004, Haiyun Yuan, Meiping Huang, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
MICCAI (2) | 4 |
| 2021 | Quantization of Deep Neural Networks for Accurate Edge ComputingabstractDeep neural networks have demonstrated their great potential in recent years, exceeding the performance of human experts in a wide range of applications. Due to their large sizes, however, compression techniques such as weight quantization and pruning are usually applied before they can be accommodated on the edge. It is generally believed that quantization leads to performance degradation, and plenty of existing works have explored quantization strategies aiming at minimum accuracy loss. In this paper, we argue that quantization, which essentially imposes regularization on weight representations, can sometimes help to improve accuracy. We conduct comprehensive experiments on three widely used applications: fully connected network for biomedical image segmentation, convolutional neural network for image classification on ImageNet, and recurrent neural network for automatic speech recognition, and experimental results show that quantization can improve the accuracy by 1%, 1.95%, 4.23% on the three applications respectively with 3.5x-6.4x memory reduction. Hailong Qiu, Jian Zhuang, Chutong Zhang, Yu Hu 0002, Qing Lu 0001, Yiyu Shi 0001, Meiping Huang, Xiaowei Xu 0004 |
ACM J. Emerg. Technol. Comput. Syst. | 10 |
| 2021 | Multi-Cycle-Consistent Adversarial Networks for Edge Denoising of Computed Tomography ImagesabstractAs one of the most commonly ordered imaging tests, the computed tomography (CT) scan comes with inevitable radiation exposure that increases cancer risk to patients. However, CT image quality is directly related to radiation dose, and thus it is desirable to obtain high-quality CT images with as little dose as possible. CT image denoising tries to obtain high-dose-like high-quality CT images (domain Y ) from low dose low-quality CT images (domain X ), which can be treated as an image-to-image translation task where the goal is to learn the transform between a source domain X (noisy images) and a target domain Y (clean images). Recently, the cycle-consistent adversarial denoising network (CCADN) has achieved state-of-the-art results by enforcing cycle-consistent loss without the need of paired training data, since the paired data is hard to collect due to patients’ interests and cardiac motion. However, out of concerns on patients’ privacy and data security, protocols typically require clinics to perform medical image processing tasks including CT image denoising locally (i.e., edge denoising). Therefore, the network models need to achieve high performance under various computation resource constraints including memory and performance. Our detailed analysis of CCADN raises a number of interesting questions that point to potential ways to further improve its performance using the same or even fewer computation resources. For example, if the noise is large leading to a significant difference between domain X and domain Y , can we bridge X and Y with a intermediate domain Z such that both the denoising process between X and Z and that between Z and Y are easier to learn? As such intermediate domains lead to multiple cycles, how do we best enforce cycle- consistency? Driven by these questions, we propose a multi-cycle-consistent adversarial network (MCCAN) that builds intermediate domains and enforces both local and global cycle-consistency for edge denoising of CT images. The global cycle-consistency couples all generators together to model the whole denoising process, whereas the local cycle-consistency imposes effective supervision on the process between adjacent domains. Experiments show that both local and global cycle-consistency are important for the success of MCCAN, which outperforms CCADN in terms of denoising quality with slightly less computation resource consumption. Xiaowei Xu 0004, Jinglan Liu, Yukun Ding, Hailong Qiu, Haiyun Yuan, Jian Zhuang, Wen Xie 0008, Yuhao Dong, Qianjun Jia, Meiping Huang, Yiyu Shi 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2021 | DAC-SDC Low Power Object Detection Challenge for UAV ApplicationsabstractThe 55th Design Automation Conference (DAC) held its first System Design Contest (SDC) in 2018. SDC'18 features a lower power object detection challenge (LPODC) on designing and implementing novel algorithms based object detection in images taken from unmanned aerial vehicles (UAV). The dataset includes 95 categories and 150k images, and the hardware platforms include Nvidia's TX2 and Xilinx's PYNQ Z1. DAC-SDC'18 attracted more than 110 entries from 12 countries. This paper presents in detail the dataset and evaluation procedure. It further discusses the methods developed by some of the entries as well as representative results. The paper concludes with directions for future improvements. Xiaowei Xu 0004, Xinyi Zhang 0001, Bei Yu 0001, Xiaobo Sharon Hu, Chris Rowen, Jingtong Hu, Yiyu Shi 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | GAC-GAN: A General Method for Appearance-Controllable Human Video Motion TransferabstractHuman video motion transfer has a wide range of applications in multimedia, computer vision, and graphics. Recently, due to the rapid development of Generative Adversarial Networks (GANs), there has been significant progress in the field. However, almost all existing GAN-based works are prone to address the mapping from human motions to video scenes, with scene appearances encoded individually in the trained models. Therefore, each trained model can only generate videos with a specific scene appearance, and new models are required to be trained to generate new appearances. Besides, existing works lack the capability of appearance control. For example, users have to provide video records of wearing new clothes or performing in new backgrounds to enable clothes or background changing in their synthetic videos, which greatly limits the application flexibility. In this paper, we propose General Appearance-Controllable GAN (GAC-GAN), a general method for appearance-controllable human video motion transfer. To enable general-purpose appearance synthesis, we propose to include appearance information in the conditioning inputs. Thus, once trained, our model can generate new appearances by altering the input appearance information. To achieve appearance control, we first obtain the appearance-controllable conditioning inputs, and then utilize a two-stage GAC-GAN to generate the corresponding appearance-controllable outputs, where we utilize an Appearance-Consistency GAN (ACGAN) loss, and a shadow extraction module for output foreground, and background appearance control respectively. We further build a solo dance dataset containing a large number of dance videos for training, and evaluation. Experimental results on our solo dance dataset, and iPER dataset show that our proposed GAC-GAN can not only support appearance-controllable human video motion transfer but also achieve higher video quality than state-of-art methods. Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang |
IEEE Trans. Multim. | 2 |
| 2020 | Do Noises Bother Human and Neural Networks In the Same Way? A Medical Image Analysis PerspectiveabstractDeep learning had already demonstrated its power in medical images, including denoising, classification, segmentation, etc. All these applications are proposed to automatically analyze medical images beforehand, which brings more information to radiologists during clinical assessment for accuracy improvement. Recently, many medical denoising methods had shown their significant artifact reduction result and noise removal both quantitatively and qualitatively. However, those existing methods are developed around human-vision, i.e., they are designed to minimize the noise effect that can be perceived by human eyes. In this paper, we introduce an application-guided denoising framework, which focuses on denoising for the following neural networks. In our experiments, we apply the proposed framework to different datasets, models, and use cases. Experimental results show that our proposed framework can achieve a better result than human-vision denoising network. Shao-Cheng Wen, Zihao Liu 0015, Wujie Wen, Xiaowei Xu 0004, Yiyu Shi 0001, Tsung-Yi Ho, Qianjun Jia, Meiping Huang, Jian Zhuang |
BIBM | 5 |
| 2020 | Constrained Multi-scale Dense Connections for Accurate Biomedical Image SegmentationabstractBiomedical image segmentation plays a critical role in clinical diagnosis and medical intervention. Recently, a variety of deep neural networks have boosted the biomedical image segmentation performance with a large margin, which adopts dense connections to explore rich representations in multiple scales. In multi-scale dense connections, features from all or most scales are fused or iteratively aggregated. In this paper, we propose constrained multi-scale dense connections (CMDC) for accurate biomedical image segmentation, which only fuse features from the nearest scales containing the most relevant appearance or semantic information. Based on CMDC, we further construct constraint multi-scale dense networks (CMD-Net) by applying CMDC to existing segmentation networks. Experiments across various architectures (including FCN-8s, U-Net, and DeepLabV3) and datasets (including GlaS, CRAG, KID, and ECS) demonstrate that CMD-Net not only outperforms existing schemes on both accuracy and efficiency but also can be easily generalized to a variety of segmentation networks. In addition, CMD-Net achieves state-of-the-art performance on two instance segmentation datasets, GlaS and CRAG. Yanchun Zhang, Shanfeng Zhu, Xiaowei Xu 0004 |
BIBM | 4 |
| 2020 | Towards Cardiac Intervention Assistance: Hardware-aware Neural Architecture Exploration for Real-Time 3D Cardiac Cine MRI SegmentationabstractReal-time cardiac magnetic resonance imaging (MRI) plays an increasingly important role in guiding various cardiac interventions. In order to provide better visual assistance, the cine MRI frames need to be segmented on-the-fly to avoid noticeable visual lag. In addition, considering reliability and patient data privacy, the computation is preferably done on local hardware. State-of-the-art MRI segmentation methods mostly focus on accuracy only, and can hardly be adopted for real-time application or on local hardware. In this work, we present the first hardware-aware multi-scale neural architecture search (NAS) framework for real-time 3D cardiac cine MRI segmentation. The proposed framework incorporates a latency regularization term into the loss function to handle realtime constraints, with the consideration of underlying hardware. In addition, the formulation is fully differentiable with respect to the architecture parameters, so that stochastic gradient descent (SGD) can be used for optimization to reduce the computation cost while maintaining optimization quality. Experimental results on ACDC MICCAI 2017 dataset demonstrate that our hardware-aware multi-scale NAS framework can reduce the latency by up to 3.5× and satisfy the real-time constraints, while still achieving competitive segmentation accuracy, compared with the state-of-the-art NAS segmentation framework. Dewen Zeng, Weiwen Jiang, Xiaowei Xu 0004, Haiyun Yuan, Meiping Huang, Jian Zhuang, Jingtong Hu, Yiyu Shi 0001 |
ICCAD | 4 |
| 2020 | BUNET: Blind Medical Image Segmentation Based on Secure UNET
Song Bian 0001, Xiaowei Xu 0004, Weiwen Jiang, Yiyu Shi 0001, Takashi Sato 0001 |
MICCAI (2) | 2 |
| 2020 | Orchestrating Medical Image Compression and Remote Segmentation Networks
Zihao Liu 0015, Sicheng Li 0001, Yen-Kuang Chen, Tao Liu 0023, Qi Liu 0017, Xiaowei Xu 0004, Yiyu Shi 0001, Wujie Wen |
MICCAI (4) | 6 |
| 2020 | ICA-UNet: ICA Inspired Statistical UNet for Real-Time 3D Cardiac Cine MRI Segmentation
Xiaowei Xu 0004, Jinjun Xiong, Qianjun Jia, Haiyun Yuan, Meiping Huang, Jian Zhuang, Yiyu Shi 0001 |
MICCAI (6) | 2 |
| 2020 | ImageCHD: A 3D Computed Tomography Image Dataset for Classification of Congenital Heart Disease
Xiaowei Xu 0004, Jian Zhuang, Haiyun Yuan, Meiping Huang, Jianzheng Cen, Qianjun Jia, Yuhao Dong, Yiyu Shi 0001 |
MICCAI (4) | 1 |
| 2020 | Binarizing Weights Wisely for Edge Intelligence: Guide for Partial Binarization of Deconvolution-Based GeneratorsabstractThis article explores the weight binarization of the deconvolution-based generator in a generative adversarial network (GAN) for memory saving and speedup of image construction on the edge. This article suggests that different from convolutional neural networks (including the discriminator) where all layers can be binarized, only some of the layers in the generator can be binarized without significant performance loss. Supported by theoretical analysis and verified by experiments, a direct metric based on the dimension of deconvolution operations is established, which can be used to quickly decide which layers in a generator can be binarized. Our results also indicate that both the generator and the discriminator should be binarized simultaneously for balanced competition and better performance during training. The experimental results on CelebA dataset with DCGAN and original loss functions suggest that directly applying state-of-the-art binarization techniques to all the layers of the generator will lead to 2.83× performance loss measured by sliced Wasserstein distance compared with the original generator, while applying them to selected layers only can yield up to 25.81× saving in memory consumption, and 1.96× and 1.32× speedup in inference and training, respectively, with little performance loss. Similar conclusions can also be drawn on other loss functions for different GANs. Jinglan Liu, Jiaxin Zhang 0014, Yukun Ding, Xiaowei Xu 0004, Meng Jiang 0001, Yiyu Shi 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2019 | SCNN: A General Distribution Based Statistical Convolutional Neural Network with Application to Video Object DetectionabstractVarious convolutional neural networks (CNNs) were developed recently that achieved accuracy comparable with that of human beings in computer vision tasks such as image recognition, object detection and tracking, etc. Most of these networks, however, process one single frame of image at a time, and may not fully utilize the temporal and contextual correlation typically present in multiple channels of the same image or adjacent frames from a video, thus limiting the achievable throughput. This limitation stems from the fact that existing CNNs operate on deterministic numbers. In this paper, we propose a novel statistical convolutional neural network (SCNN), which extends existing CNN architectures but operates directly on correlated distributions rather than deterministic numbers. By introducing a parameterized canonical model to model correlated data and defining corresponding operations as required for CNN training and inference, we show that SCNN can process multiple frames of correlated images effectively, hence achieving significant speedup over existing CNN models. We use a CNN based video object detection as an example to illustrate the usefulness of the proposed SCNN as a general network model. Experimental results show that even a nonoptimized implementation of SCNN can still achieve 178% speedup over existing CNNs with slight accuracy degradation. Jinjun Xiong, Xiaowei Xu 0004, Yiyu Shi 0001 |
AAAI | 3 |
| 2019 | Machine Vision Guided 3D Medical Image Compression for Efficient Transmission and Accurate Segmentation in the CloudsabstractCloud based medical image analysis has become popular recently due to the high computation complexities of various deep neural network (DNN) based frameworks and the increasingly large volume of medical images that need to be processed. It has been demonstrated that for medical images the transmission from local to clouds is much more expensive than the computation in the clouds itself. Towards this, 3D image compression techniques have been widely applied to reduce the data traffic. However, most of the existing image compression techniques are developed around human vision, i.e., they are designed to minimize distortions that can be perceived by human eyes. In this paper, we will use deep learning based medical image segmentation as a vehicle and demonstrate that interestingly, machine and human view the compression quality differently. Medical images compressed with good quality w.r.t. human vision may result in inferior segmentation accuracy. We then design a machine vision oriented 3D image compression framework tailored for segmentation using DNNs. Our method automatically extracts and retains image features that are most important to the segmentation. Comprehensive experiments on widely adopted segmentation frameworks with HVSMR 2016 challenge dataset show that our method can achieve significantly higher segmentation accuracy at the same compression rate, or much better compression rate under the same segmentation accuracy, when compared with the existing JPEG 2000 method. To the best of the authors' knowledge, this is the first machine vision guided medical image compression framework for segmentation in the clouds. Zihao Liu 0015, Xiaowei Xu 0004, Tao Liu 0023, Qi Liu 0017, Yanzhi Wang 0001, Yiyu Shi 0001, Wujie Wen, Meiping Huang, Haiyun Yuan, Jian Zhuang |
CVPR | 2 |
| 2019 | MSU-Net: Multiscale Statistical U-Net for Real-Time 3D Cardiac MRI Video Segmentation
Jinjun Xiong, Xiaowei Xu 0004, Meng Jiang 0001, Haiyun Yuan, Meiping Huang, Jian Zhuang, Yiyu Shi 0001 |
MICCAI (2) | 3 |
| 2019 | Whole Heart and Great Vessel Segmentation in Congenital Heart Disease Using Deep Neural Networks and Graph Matching
Xiaowei Xu 0004, Yiyu Shi 0001, Haiyun Yuan, Qianjun Jia, Meiping Huang, Jian Zhuang |
MICCAI (2) | 1 |
| 2019 | Optimal design of a low-power, phase-switching modulator for implantable medical applications
Dawei Li 0012, Xiaowei Xu 0004, Leibo Liu, Li Zhang 0021, Cheng Zhuo, Yiyu Shi 0001 |
Integr. | 2 |
| 2019 | MDA: A Reconfigurable Memristor-Based Distance Accelerator for Time Series Mining on Data CentersabstractThe rapid development of Internet-of-Things is yielding a huge volume of time series data, the real-time mining of which becomes a major load for data centers. The computation bottleneck in time series data mining is distance function, which is the fundamental element of many high data mining tasks. Recently various software optimization and hardware acceleration techniques have been proposed to tackle the challenge. However, each of these techniques is only designed or optimized for a specific distance function. To address this problem, in this paper we propose MDA, a high-throughput reconfigurable memristor-based distance accelerator for real-time and energy-efficient data mining with time series in data centers. Common circuit structure is extracted for efficiency, and the circuit can be configured to any specific distance functions. Particularly, we adopt the emerging device memristor for the design of MDA. Comprehensive experiments are presented with public available datasets to evaluate the performance of the proposed MDA. Experimental results show that compared with existing works, MDA has achieved a speedup of 3.5×-376× on performance and an improvement of 1-3 orders of magnitude on energy efficiency with little accuracy loss. Xiaowei Xu 0004, Feng Lin 0004, Wenyao Xu, Xin-Wei Yao 0001, Yiyu Shi 0001, Dewen Zeng, Yu Hu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Quantization of Fully Convolutional Networks for Accurate Biomedical Image SegmentationabstractWith pervasive applications of medical imaging in health-care, biomedical image segmentation plays a central role in quantitative analysis, clinical diagnosis, and medical intervention. Since manual annotation suffers limited reproducibility, arduous efforts, and excessive time, automatic segmentation is desired to process increasingly larger scale histopathological data. Recently, deep neural networks (DNNs), particularly fully convolutional networks (FCNs), have been widely applied to biomedical image segmentation, attaining much improved performance. At the same time, quantization of DNNs has become an active research topic, which aims to represent weights with less memory (precision) to considerably reduce memory and computation requirements of DNNs while maintaining acceptable accuracy. In this paper, we apply quantization techniques to FCNs for accurate biomedical image segmentation. Unlike existing literatures on quantization which primarily targets memory and computation complexity reduction, we apply quantization as a method to reduce overfitting in FCNs for better accuracy. Specifically, we focus on a state-of-the-art segmentation framework, suggestive annotation [26], which judiciously extracts representative annotation samples from the original training dataset, obtaining an effective small-sized balanced training dataset. We develop two new quantization processes for this framework: (1) suggestive annotation with quantization for highly representative training samples, and (2) network training with quantization for high accuracy. Extensive experiments on the MICCAI Gland dataset show that both quantization processes can improve the segmentation performance, and our proposed method exceeds the current state-of-the-art performance by up to 1%. In addition, our method has a reduction of up to 6.4x on memory usage. Xiaowei Xu 0004, Qing Lu 0001, Lin Yang 0003, Xiaobo Sharon Hu, Danny Ziyi Chen, Yu Hu 0002, Yiyu Shi 0001 |
CVPR | 1 |
| 2018 | Efficient Hardware Implementation of Cellular Neural Networks with Incremental Quantization and Early ExitabstractCellular neural networks (CeNNs) have been widely adopted in image processing tasks. Recently, various hardware implementations of CeNNs have emerged in the literature, with Field Programmable Gate Array (FPGA) being one of the most popular choices due to its high flexibility and low time-to-market. However, CeNNs typically involve extensive computations in a recursive manner. As an example, to simply process an image of 1,920 × 1,080 pixels requires 4--8 Giga floating point multiplications (for 3 × 3 templates and 50–100 iterations), which needs to be done in a timely manner for real-time applications. To address this issue, in this article, we propose a compressed CeNN framework for efficient FPGA implementations. It involves various techniques, such as incremental quantization and early exit, which significantly reduces computation demands while maintaining an acceptable performance. Particularly, incremental quantization quantizes the numbers in CeNN templates to powers of two, so that complex and expensive multiplications can be converted to simple and cheap shift operations, which only require a minimum number of registers and logical elements (LEs). While a similar concept has been explored in hardware implementations of Convolutional Neural Networks (CNNs), CeNNs have completely different computation patterns, which require different quantization and implementation strategies. Experimental results on FPGAs show that incremental quantization and early exit can achieve a speedup of up to 7.8× and 8.3×, respectively, compared with the state-of-the-art implementations, while with almost no performance loss with four widely adopted applications. We also discover that different from CNNs, the optimal quantization strategies of CeNNs depend heavily on the applications. We hope that our work can serve as a pioneer in the hardware optimization of CeNNs. Xiaowei Xu 0004, Qing Lu 0001, Yu Hu 0002, Chen Zhuo, Jinglan Liu, Yiyu Shi 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 1 |
| 2018 | A Multi-Level-Optimization Framework for FPGA-Based Cellular Neural Network ImplementationabstractCellular Neural Network (CeNN) is considered as a powerful paradigm for embedded devices. Its analog and mix-signal hardware implementations are proved to be applicable to high-speed image processing, video analysis, and medical signal processing with its efficiency and popularity limited by smaller implementation size and lower precision. Recently, digital implementations of CeNNs on FPGA have attracted researchers from both academia and industry due to its high flexibility and short time-to-market. However, most existing implementations are not well optimized to fully utilize the advantages of FPGA platform with unnecessary design and computational redundancy that prevents speedup. We propose a multi-level-optimization framework for energy-efficient CeNN implementations on FPGAs. In particular, the optimization framework is featured with three level optimizations: system-, module-, and design-space-level, with focus on computational redundancy and attainable performance, respectively. Experimental results show that with various configurations our framework can achieve an energy-efficiency improvement of 3.54× and up to 3.88× speedup compared with existing implementations with similar accuracy. Zhongyang Liu, Shaoheng Luo, Xiaowei Xu 0004, Yiyu Shi 0001, Cheng Zhuo |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2018 | Accelerating Dynamic Time Warping With Memristor-Based Customized FabricsabstractThe rapid development of Internet of Things is yielding a huge volume of time series data, the real-time mining of which becomes a major load for data centers. The computation bottleneck in time series mining is the distance measure, in which dynamic time warping (DTW) is one of the most widely used distance measures. Recently, various software optimization and hardware acceleration techniques have been proposed for DTW acceleration. However, the throughput and energy efficiency of DTW are still big concerns considering the ever-increasing volume of times series. In this paper, we propose a high-throughput and efficient memristor-based DTW architecture for real-time time series mining on data centers. Specifically, memristors have been adopted for both computation and configuration of the computing architecture. The computation flow in this architecture is fully presented in a continuous and asynchronous manner. To improve the computation efficiency, we propose an early lower bound algorithm by exploiting the predictability in the circuit characteristic. Experiments are performed with module evaluation and end-to-end evaluation including three popular applications: 1) similarity search; 2) classification; and 3) anomaly detection. Experimental results indicate that, compared to existing approaches, the speedup and energy efficiency improvement are 12x-43x and 51x-287x, respectively. Xiaowei Xu 0004, Feng Lin 0004, Aosen Wang, Xin-Wei Yao 0001, Qing Lu 0001, Wenyao Xu, Yiyu Shi 0001, Yu Hu 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | An Efficient Memristor-based Distance Accelerator for Time Series Data Mining on Data CentersabstractThe rapid development of Internet-of-Things (IoT) is yielding a huge volume of time series data, the real-time mining of which becomes a major load for data centers. The computation bottleneck in time series data mining is the distance function, which has been tackled by various software optimization and hardware acceleration techniques recently. However, each of these techniques is only designed or optimized for a specific distance function. To address this problem, in this paper we propose an efficient and reconfigurable memristor-based distance accelerator for real-time and energy-efficient data mining with time series on data centers. Common circuit structure is extracted to save chip areas, and the circuit can be configured to any specific distance functions. Experimental results show that compared with existing works, our work has achieved a speedup of 3.5x-376x on performance and an improvement of 1-3 orders of magnitude on energy efficiency. Xiaowei Xu 0004, Dewen Zeng, Wenyao Xu, Yiyu Shi 0001, Yu Hu 0002 |
DAC | 1 |
| 2017 | Edge segmentation: Empowering mobile telemedicine with compressed cellular neural networksabstractWith the need for increased care and welfare of the rapidly aging population, mobile telemedicine is becoming popular for providing remote health care to increase the quality of life. Recently, image analysis is being actively applied for medical diagnosis and treatment, in which image segmentation is of the fundamental importance for other image processing such as visualization and detection. However, given the tasks challenges in transmitting large volume of high-resolution images and the real-time constraints that are commonly present for mobile telemedicine, image segmentation is best done at the “edge”, i.e., locally so that only segmentation results are communicated. A powerful approach to medical image segmentation is cellular neural network (CeNN), which can achieve very high accuracy through proper training. However, CeNNs typically involve extensive computations in a recursive manner. As an example, to simply process an image of 1920×1080 pixels requires 4-8 Giga floating point multiplications (for 3×3 templates and 50-100 iterations), which needs to be done in a timely manner for real-time medical image segmentation. Such a demand is too high for most low power mobile computing platforms in IoTs, This paper presents a compressed CeNN framework for computation reduction in CeNNs, which is the first in the literature. It involves various techniques such as early exit and parameter quantization, which significantly reduces computation demands while maintaining an acceptable performance. Xiaowei Xu 0004, Qing Lu 0001, Jinglan Liu, Cheng Zhuo, Xiaobo Sharon Hu, Yiyu Shi 0001 |
ICCAD | 1 |