Yu Cao 0002

dblp:68/6563-2 · DBLP profile ↗
← Back
56ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0001-8624-1099ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Systems, architecture and hardware · 4 · 1 first-authorComputer networks · 2Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 One-stage Framework for Thyroid Nodule Detection with Mixup and Negative Sample Utilization
abstract
Ultrasound imaging plays a significant role in the early diagnosis of thyroid nodules. However, ambiguous nodule boundaries and their similarity to surrounding tissue cause challenges. These cause high false positive rates, reducing the diagnostic efficiency of ultrasound imaging. Although recent advancements in deep learning such as Faster R-CNN and YOLO, have shown progress, they still have limitations in handling imbalanced datasets and noise during real-time detection. To solve these challenges, our team proposes a novel deep learning framework integrating mixup data augmentation and negative sample utilization. This approach reduces false positives while maintaining a stable true positive detection rate. The method employs a one-stage-based object detection model as the baseline. The proposed framework based on that can improve the model’s robustness against background noise and optimize mixup parameters to enhance generalization. The experimental results demonstrate our approach significantly reduces false positive rates by 40% (from 50% to below 10%). In the meantime, it can stabilize high true positive rates above 80% in thyroid nodule detection during testing. This improvement highlights the framework can balance diagnostic sensitivity and specificity. The framework provides an efficient solution for detecting nodules in thyroid ultrasound imaging and has great potential to enhance diagnostic accuracy.
Qilei Chen, Zinan Xiong, Yu Cao 0002, Benyuan Liu
ICIP5
2023 A CNN-Based Disease Detection Framework for Wireless Capsule Endoscopy Videos
abstract
In colonoscopy, wireless capsule endoscopy (WCE) is widely used since it is more physically friendly for patients than standard endoscopy. WCE is a low-risk and effective clinical operation for the small intestine endoscopy, which is less accessible to standard endoscopy. However, reviewing WCE videos can be challenging as it is time-consuming and requires considerable expertise. WCE videos are usually captured at a low resolution and partial frames are filtered out due to hardware limitations. Additional challenges arise from the diversity and complexity of gastrointestinal (GI) diseases. Moreover, inadequate clinic attention can cause clinical errors. Consequently, physicians are often overburdened with work. This paper presents a convolutional neutral networks (CNN) based framework to assist clinicians to review the WCE videos. The framework combines both image classification and object detection results to provide comprehensive results for disease detection in WCE videos. For disease classification, ResNet-50 was selected when experiments were conducted on the dataset. For disease detection, YOLO-X is employed on the images labelled with bonding boxes. Enhanced with an offline hard example mining (offline-HEM) procedure and fine-tuning on hyper-parameters, this framework can achieve high sensitivity for disease instances while maintaining acceptable specificity for false positives in video tests.
Qilei Chen, Yani Yin, Guanghui Lian, Shuijiao Chen, Yu Cao 0002, Benyuan Liu
HealthCom8
2023 A Greedy Algorithm-Based Self-Training Pipeline for Expansion of Dental Caries Dataset
abstract
Dental caries, the most prevalent oral disease, poses a significant healthcare challenge. Deep Neural Network (DNN)-based object detection techniques offer promising solutions to improve the efficiency of dental caries diagnosis. It is widely acknowledged that the performance of DNN models heavily relies on the availability of sufficient and accurately labeled data. The collection and annotation of dental X-ray images encounter obstacles due to privacy concerns and the requirement for specialized expertise. Consequently, the limited access to labeled dental image datasets restricts the potential of DNNs in supporting oral and dental healthcare. Self-Training (ST) is a semi-supervised machine leaning approach that addresses this problem to a large extent. It repeats the procedures of training a model on the labeled dataset, and then applying it to generate pseudo labels on the unlabeled dataset, and further using the combined data with the original and pseudo labels to train new models. However, the latent errors of the pseudo labels can arise and even be amplified throughout the ST pipeline, which leads to a significant performance decline for DNN models. In this paper, we propose a Greedy algorithm-based Self-Training (Greedy-ST) pipeline to address this problem. At each iteration, the Greedy-ST selects an optimal confidence threshold to generate predictions as pseudo labels, and uses static fine-tuning (SFT) and dynamic fine-tuning (DFT) to refine them. Experimental results demonstrate that by utilizing the pseudo labels generated by the Greedy-ST pipeline, the selected baseline model achieves improved performance compared to using the pseudo labels generated by the vanilla ST approach.
Qilei Chen, Yu Cao 0002, Xinwen Fu, Benyuan Liu
HealthCom6
2023 MLMSA: Multi-Level and Multi-Scale Attention for Lesion Detection in Endoscopy
abstract
The advancement of deep learning techniques has significantly improved abnormality detection in gastrointestinal (GI) endoscopy. However, this imaging process comes with challenges due to the complex nature of GI abnormalities. The wide variety of abnormalities in terms of type, color, texture, shape, and scale of lesions makes it difficult to accurately detect them in different scenarios. Furthermore, the presence of multiple types of lesions within the same region create complex scenarios that complicate abnormality detection. Additionally, differentiating early-stage cancers from non-cancerous lesions is a significant challenge even for experienced professionals. The simultaneous identification of cancers, particularly early-stage ones, and non-cancerous lesions within the same region remains a challenging issue in GI endoscopy imaging. In this study, we discover that multiple types of lesions exhibit a scale-sensitive characteristic that can be leveraged by multi-level feature-based deep learning models. Hence, we propose the use of a multi-level and multi-scale attention (MLMSA) neck module in a deep learning network. The MLMSA module utilizes multiple levels of features extracted from the backbone network to generate processed multi-level features that assist the detection head. By integrating the MLMSA module into the deep learning framework, our goal is to enhance the detection and differentiation of lesions, particularly the early-stage cancer, thereby advancing the capabilities of GI endoscopy imaging. Our experiment results show that integrating the MLMSA module leads to a significant improvement in the detection of GI abnormalities, providing compelling evidence for the enhanced performance achieved through the utilization of the MLMSA module in our approach.
Shuijiao Chen, Qilei Chen, Yizhe Zhang 0001, Yu Cao 0002, Benyuan Liu
HealthCom7
2023 A Deep Learning Framework with Pruning RoI Proposal for Dental Caries Detection in Panoramic X-ray Images
Qilei Chen, Yu Cao 0002, Xinwen Fu, Benyuan Liu
ICONIP (3)6
2022 Deep Learning Assisted Mouth-Esophagus Passage Time Estimation During Gastroscopy
abstract
A gastroscopy involves examining the upper digestive system using a flexible tube equipped with a small camera. Generally, it is performed to determine the cause of digestive symptoms, such as vomiting blood, stomach pains, and difficulty swallowing. Though this procedure has been performed since the mid-19th century, and various measures have been implemented to make it easier and less invasive, it is still not risk-free. One of the major complications is esophagus perforation, and most of them happen during the insertion of the gastroscopy. Therefore, it is necessary to develop an effective method for evaluating the performance of the operator. One appropriate metric is the time interval between the mouth and esophagus during the intubation. In this paper, we propose a gastroscopy video processing system based on deep learning to automatically evaluate the mouth-esophagus passage time. In this system, a Convolutional Neural Network (CNN) based model is adopted to detect the mouth and esophagus, track the timestamps of the last appearance of the mouth and the first appearance of the esophagus, and calculate the interval between those appearances. Our system is capable of dealing with abnormal circumstances that can occur during a procedure, as well as reporting accurate results. Experiment results show that our best model achieves an accuracy of 88.92% on image dataset, and an accuracy of 99.86% on videos for the mouth-esophagus passage time.
Zinan Xiong, Qilei Chen, Yu Cao 0002, Benyuan Liu
ICTAI4
2022 Enhance Chest X-ray Classification with Multi-image Fusion and Pseudo-3D Reconstruction
abstract
Chest radiography (X-ray) is a critical imaging modality for the diagnosis of thorax diseases. Automated classification of chest X-rays has enormous potential to benefit diagnostic decision making. Most existing models are designed to consider a single-view chest X-ray image, mostly the frontal view, failing to utilize the complementary information provided by the lateral view when available. In clinical practice, radiologists often examine both frontal and lateral views to gather more information for diagnosis. Thus, a well-designed model should take advantage of the paired views to enhance the classification results. To this end, we present a novel dual-view deep learning framework to enhance the classification that uses an intermediate bi-directional fusion architecture to exploit the intrinsic spatial correlation registered between the two views. In particular, we design two different modules, a vertical alignment fusion and a 3D reconstruction fusion, that can effectively extract and fuse important features from both input streams to generate a pair of recalibrating signals, which are then used to amplify the pertinent parts while suppressing the irrelevant segments of the feature map in each stream. We evaluate our proposed methods extensively on two large public chest X-ray datasets, CheXpert and MIMIC-CXR-JPG. Experiment results show that our methods can effectively improve the classification results over fine-tuned baselines and other state-of-the-art methods, demonstrating the benefit of exploiting the spatial correlation between the paired views in the model. We also find that our fusion methods significantly outperform the early fusion and late fusion models, highlighting the need to capture the spatial correlation across different intermediate feature layers.
Jing Ni, Zubin Bhuyan, Qilei Chen, Xinzi Sun, Dechun Wang, Yu Cao 0002, Benyuan Liu
IJCNN6
2022 AFP-Mask: Anchor-Free Polyp Instance Segmentation in Colonoscopy
abstract
Colorectal cancer (CRC) is a common and lethal disease. Globally, CRC is the third most commonly diagnosed cancer in males and the second in females. The most effective way to prevent CRC is through using colonoscopy to identify and remove precancerous growths at an early stage. The detection and removal of colorectal polyps have been found to be associated with a reduction in mortality from colorectal cancer. However, the false negative rate of polyp detection during colonoscopy is often high even for experienced physicians. With recent advances in deep learning based object detection techniques, automated polyp detection shows great potential in helping physicians reduce false positive rate during colonoscopy. In this paper, we propose a novel anchor-free instance segmentation framework that can localize polyps and produce the corresponding instance level masks without using predefined anchor boxes. Our framework consists of two branches: (a) an object detection branch that performs classification and localization, (b) a mask generation branch that produces instance level masks. Instead of predicting a two-dimensional mask directly, we encode it into a compact representation vector, which allows us to incorporate instance segmentation with one-stage bounding-box detectors in a simple yet effective way. Moreover, our proposed encoding method can be trained jointly with object detector. Our experiment results show that our framework achieves a precision of 99.36% and a recall of 96.44% on public datasets, outperforming existing anchor-free instance segmentation methods by at least 2.8% in mIoU on our private dataset.
Dechun Wang, Shuijiao Chen, Xinzi Sun, Qilei Chen, Yu Cao 0002, Benyuan Liu
IEEE J. Biomed. Health Informatics5
2022 DMA-Net: DeepLab With Multi-Scale Attention for Pavement Crack Segmentation
abstract
Cracks are important indicators of pavement structural and operational conditions. Early pavement crack detection and treatments can help extend pavement service life, reduce fuel consumption, and improve safety and ride quality. Pavement distress surveys have traditionally been performed manually by visually inspecting the roads, which is labor-intensive and time-consuming. Therefore, computer-vision-based automated crack detection has great practical significance in pavement maintenance and traffic safety. Traditional image processing techniques are sensitive to noise in images and are thus likely to miss detecting some cracks due to the crack texture variety, complex lighting conditions, and various similar but irrelevant objects on the road. This paper adopts and enhances DeepLabv3+, a popular deep learning framework for semantic image segmentation, for road pavement crack detection. We propose a multi-scale attention module in the decoder of DeepLabv3+ to generate an attention mask and dynamically assign weights between high-level and low-level feature maps. Compared with fixed weights across different features, the dynamic weights strategy can assign more reasonable weights to different feature maps. Ablation experiments show that the attention mask can effectively help the model better combine multi-scale features and generate more accurate pavement crack segmentation results. The proposed method achieves state-of-the-art results on three benchmarks, including Crack500, DeepCrack, and FMA (Fitchburg Municipal Airport) datasets. We further test it on pavement crack images captured by smartphones, and the results show that it provides a viable approach to road pavement crack segmentation in practice with excellent performance.
Xinzi Sun, Yuanchang Xie, Liming Jiang 0004, Yu Cao 0002, Benyuan Liu
IEEE Trans. Intell. Transp. Syst.4
2021 Gastric Location Classification During Esophagogastroduodenoscopy Using Deep Neural Networks
abstract
Esophagogastroduodenoscopy (EGD) is a common procedure that visualizes the esophagus, stomach, and the duodenum by inserting a camera, attached to a long flexible tube, into the patient's mouth and down the stomach. A comprehensive EGD needs to examine all gastric locations, but since the camera is controlled manually, it is easy to miss some surface area and create diagnostic blind spots, which often result in life-costing oversights of early gastric cancer and other serious illnesses. In order to address this problem, we train a convolutional neural network to classify gastric locations based on the camera feed during an EGD, and based on the classifier and a triggering algorithm we propose, we build a video processing system that checks off each location as visited, allowing human operators to keep track of which locations they have visited and which they have not. Based on collected clinical patient reports, we consider six gastric locations, and we add a background class to our classifier to accomodate for the frames in EGD videos that do not resemble the six defined classes (including when the camera is outside of the patient body). Our best classifier achieves 98 % accuracy within the six gastric locations and 88 % accuracy including the background class, and our video processing system clearly checks off gastric locations in an expected order when being tested on recorded EGD videos. Lastly, we use class activation mapping to provide human-readable insight into how our trained classifier works.
Alexander Ding, Ying Li 0133, Qilei Chen, Yu Cao 0002, Benyuan Liu, Shuijiao Chen
BIBE4
2021 Lower Body Rehabilitation Dataset and Model Optimization
abstract
Human pose estimation has enabled numerous applications by classifying and tracking body movements. Although a few open datasets have emerged to facilitate the evaluation of pose detection methods, they are too generic to benefit do-main specific applications such as physical therapy which has quantitative clinical metrics and requires precise differentiation and measurement. To address this issue, we construct the first human keypoints detection dataset for physical therapy, in particular lower body rehabilitation. The dataset consists of 1,885,637 distinctive human poses for 31 lower body rehab exercises, which are performed by 20 actors under the guidance of a licensed physical therapist. Their motion are captured with state of art of motion tracking system in both 3D and 2D to establish the ground truth. Furthermore, to optimize a number of deep learning models applied on this unique dataset, we utilize Extremely Efficient Spatial Pyramid (EESP) and attention mechanism to reduce the models’ computational complexity. Our experiment results show that the optimized models achieve comparable performance with nearly 4x reduction in complexity.
Ying Li 0133, Zinan Xiong, Yan Luo 0001, Yu Cao 0002
ICME5
2021 Detection of Endoscope Withdrawal Time in Colonoscopy Videos
abstract
A colonoscopy is an endoscopic examination that visualizes the colon by inserting a camera, attached to a long, flexible tube, into the patient’s anus and up the rectum. One significant landmark of a colonoscopy is when the endoscope begins to withdraw after reaching the farthest point in its path, the cecum. In this paper, we demonstrate that the withdrawal point is closely related to the sighting of local anatomical features in the cecum, specifically the ileocecal valve. The withdrawal point allows us to determine the amount of time the endoscope spends withdrawing, which is an important indicator of the quality of a colonoscopy. In this paper, we present a colonoscopy video processing system that detects the withdrawal point in real time by using a convolutional neural network to classify between ileocecal valves and other images. The system then processes the raw classifier output to determine the withdrawal point. We collect a novel dataset of colonoscopy images and videos to train and evaluate the classifier, as well as to evaluate the video processing system. We explore a range of state-of-the-art classifier architectures, and our best model achieves 99.6% accuracy on the image-level dataset. Then, we provide human-readable insight into our classifier using class activation mapping and principle component analysis. Using this classifier and optimized parameters, our video processing system achieves 70.5% accuracy (± 10s) on the video-level dataset.
Ying Li 0133, Alexander Ding, Yu Cao 0002, Benyuan Liu, Shuijiao Chen
ICMLA3
2021 Dual-CLVSA: a Novel Deep Learning Approach to Predict Financial Markets with Sentiment Measurements
abstract
It is a challenging task to predict financial markets. The complexity of this task is mainly due to the interaction between financial markets and market participants, who are not able to keep rational all the time, and often affected by emotions such as fear and ecstasy. Based on the state-of-the-art approach particularly for financial market predictions, a hybrid convolutional LSTM Based variational sequence-to-sequence model with attention (CLVSA), we propose a novel deep learning approach, named dual-CLVSA, to predict financial market movement with both trading data and the corresponding social sentiment measurements, each through a separate sequence-to-sequence channel. We evaluate the performance of our approach with backtesting on historical trading data of SPDR SP 500 Trust ETF over eight years. The experiment results show that dual-CLVSA can effectively fuse the two types of data, and verify that sentiment measurements are not only informative for financial market predictions, but they also contain extra profitable features to boost the performance of our predicting system.
Hongwei Zhu 0002, Jiancheng Shen, Yu Cao 0002, Benyuan Liu
ICMLA4
2020 A Knowledge-Based Decision Support System for In Vitro Fertilization Treatment
abstract
In Vitro Fertilization (IVF) is the most widely used Assisted Reproductive Technology (ART). IVF usually involves controlled ovarian stimulation, oocyte retrieval, fertilization in the laboratory with subsequent embryo transfer. The first two steps correspond with females' follicular phase and ovulation in their menstrual cycle. Therefore, we refer to it as the treatment cycle in our paper. The treatment cycle is crucial because the stimulation medications in IVF treatment are applied directly on patients. In order to optimize the stimulation effects and lower the side effects of the stimulation medications, prompt treatment adjustments are in need. In addition, the quality and quantity of the retrieved oocytes have a significant effect on the outcome of the following procedures. To improve the IVF success rate, we propose a knowledge-based decision support system that can provide medical advice on the treatment protocol and medication adjustment for each patient visit during IVF treatment cycle. Our system is efficient in data processing and light-weighted which can be easily embedded into electronic medical record systems. Moreover, an oocyte retrieval oriented evaluation demonstrates that our system performs well in terms of accuracy of advice for the protocols and medications.
Jing Ni, Xinzi Sun, Zitao Liu 0005, Yu Cao 0002, Benyuan Liu
HealthCom8
2020 Colorectal Polyp Detection in Real-world Scenario: Design and Experiment Study
abstract
Colorectal polyps are abnormal tissues growing on the intima of the colon or rectum with a high risk of developing into colorectal cancer, the third leading cause of cancer death worldwide. Early detection and removal of colon polyps via colonoscopy have proved to be an effective approach to prevent colorectal cancer. Recently, various CNN-based computer-aided systems have been developed to help physicians detect polyps. However, these systems do not perform well in real-world colonoscopy operations due to the significant difference between images in a real colonoscopy and those in the public datasets. Unlike the well-chosen clear images with obvious polyps in the public datasets, images from a colonoscopy are often blurry and contain various artifacts such as fluid, debris, bubbles, reflection, specularity, contrast, saturation, and medical instruments, with a wide variety of polyps of different sizes, shapes, and textures. All these factors pose a significant challenge to effective polyp detection in a colonoscopy. To this end, we collect a private dataset that contains 7,313 images from 224 complete colonoscopy procedures. This dataset represents realistic operation scenarios and thus can be used to better train the models and evaluate a system's performance in practice. We propose an integrated system architecture to address the unique challenges for polyp detection. Extensive experiments results show that our system can effectively detect polyps in a colonoscopy with excellent performance in real time.
Xinzi Sun, Dechun Wang, Zinan Xiong, Yu Cao 0002, Benyuan Liu, Shuijiao Chen
ICTAI6
2020 Pseudo-Labeling for Small Lesion Detection on Diabetic Retinopathy Images
abstract
Diabetic retinopathy (DR) is a primary cause of blindness in working-age people worldwide. About 3 to 4 million people with diabetes become blind because of DR every year. Diagnosis of DR through color fundus images is a common approach to mitigate such problem. However, DR diagnosis is a difficult and time consuming task, which requires experienced clinicians to identify the presence and significance of many small features on high resolution images. Convolutional Neural Network (CNN) has proved to be a promising approach for automatic biomedical image analysis recently. In this work, we investigate lesion detection on DR fundus images with CNN-based object detection methods. Lesion detection on fundus images faces two unique challenges. The first one is that our dataset is not fully labeled, i.e., only a subset of all lesion instances are marked. Not only will these unlabeled lesion instances not contribute to the training of the model, but also they will be mistakenly counted as false negatives, leading the model move to the opposite direction. The second challenge is that the lesion instances are usually very small, making them difficult to be found by normal object detectors. To address the first challenge, we introduce an iterative training algorithm for the semi-supervised method of pseudo-labeling, in which a considerable number of unlabeled lesion instances can be discovered to boost the performance of the lesion detector. For the small size targets problem, we extend both the input size and the depth of feature pyramid network (FPN) to produce a large CNN feature map, which can preserve the detail of small lesions and thus enhance the effectiveness of the lesion detector. The experimental results show that our proposed methods significantly outperform the baselines.
Qilei Chen, Jing Ni, Yu Cao 0002, Benyuan Liu, Honggang Zhang 0003
IJCNN4
2020 Retinopathy of Prematurity Stage Diagnosis Using Object Segmentation and Convolutional Neural Networks
abstract
Retinopathy of Prematurity (ROP) is an eye disorder primarily affecting premature infants with lower weights. It causes proliferation of vessels in the retina and could result in vision loss and, eventually, retinal detachment, leading to blindness. While human experts can easily identify severe stages of ROP, the diagnosis of earlier stages, which are the most relevant to determining treatment choice, are much more affected by variability in subjective interpretations of human experts. In recent years, there has been a significant effort to automate the diagnosis using deep learning. This paper builds upon the success of previous models and develops a novel architecture, which combines object segmentation and convolutional neural networks (CNN) to construct an effective classifier of ROP stages 1-3 based on neonatal retinal images. Motivated by the fact that the formation and shape of a demarcation line in the retina is the distinguishing feature between earlier ROP stages, our proposed system first trains an object segmentation model to identify the demarcation line at a pixel level and adds the resulting mask as an additional "color" channel in the original image. Then, the system trains a CNN classifier based on the processed images to leverage information from both the original image and the mask, which helps direct the model's attention to the demarcation line. In a number of careful experiments comparing its performance to previous object segmentation systems and CNN-only systems trained on our dataset, our novel architecture significantly outperforms previous systems in accuracy, demonstrating the effectiveness of our proposed pipeline.
Alexander Ding, Qilei Chen, Yu Cao 0002, Benyuan Liu
IJCNN3
2020 Human pose estimation based in-home lower body rehabilitation system
abstract
In this paper, we design, develop and evaluate an in-home lower body rehabilitation system based on a novel lightweight human pose estimation model. To achieve that, we first create a lower body rehabilitation dataset of 500,000 images with each image annotated with the ground truth joint point locations. The dataset consists of 31 different types of lower body rehabilitation activities from twenty volunteers. After that, we design a lightweight but powerful neural network model, which runs on a smartphone, to estimate human pose. Furthermore, we develop a series of principles for evaluating in-home rehabilitation activities of patients in terms of the range of motion and duration of activities. For the concern of privacy, all the data collected from patients are encrypted, stored and processed locally on patients' own smartphones. Only the sanitized evaluation reports are uploaded and shared with the patients' primary doctors. Our model achieves 70.8 in AP score on the COCO val2017 set with only 4.7M parameters and 1.0 GFLOPs. Using our system, patients can perform lower body rehabilitation activities at home and obtain evaluation report without the presence of physical therapists. We believe our system can greatly facilitate in-home rehabilitation and reduce the cost for patients.
Ying Li 0133, Yu Cao 0002, Benyuan Liu, Joanna Tan, Yan Luo 0001
IJCNN3
2019 People Re-Identification by Multi-Branch CNN with Multi-Scale Features
abstract
People re-identification is a retrieval problem to find a person of interest among a gallery of person images from different cameras in various poses or view angles. How to get a strong feature representation for a person image plays an important role in performing people re-identification. In this paper, we present a novel end-to-end framework that extracts both global and local features with multiple scales to generate more discriminative representations. The model we design is a multi-branch network consisting of one global branch to obtain features of the whole input image from different convolutional layers and several local branches to obtain features from horizontal partitions in different granularities. Our method achieves state-of-the-art results on three challenging datasets (Market-1501, CUHK03 and DukeMTMC-reid).
Xinzi Sun, Qilei Chen, Yu Cao 0002, Benyuan Liu
ICIP4
2019 Multi-Stream Single Shot Spatial-Temporal Action Detection
abstract
We present a 3D Convolutional Neural Networks (CNNs) based single shot detector for spatial-temporal action detection tasks. Our model includes: (i) two short-term appearance and motion streams, with single RGB and optical flow image input separately, in order to capture the spatial and temporal information for the current frame; (ii) two long-term 3D ConvNet based stream, working on sequences of continuous RGB and optical flow images to capture the context from past frames. Our model achieves strong performance for action detection in video and can be easily integrated into any current two-stream action detection methods. We report a frame-mAP of 71.30% on the challenging UCF101-24 [1] actions dataset, achieving the state-of-the-art result of the one-stage methods. To the best of our knowledge, our work is the first system that combined 3D CNN and SSD in action detection tasks.
Yu Cao 0002, Benyuan Liu
ICIP2
2019 An Edge Computing Visual System for Vegetable Categorization
abstract
In self-service supermarket and retail industry, efforts to reduce customer wait time using automatic grocery item identification are challenged by low recognition accuracy, long response time and substantial requirement for equipment. In this paper, we propose a novel edge computing system named EdgeVegfru for vegetable and fruit image classification. While existing work on Vegfru dataset shows excellent performance, few of them have been deployed in real-world applications. We adopt an edge computing paradigm, design, implement and evaluate the whole system on the Android devices. The proposed deep learning model and quantization algorithm reduce the model size and inference time significantly. Our system has shown out-standing accuracy within limited time and computation resources, compared with other machine learning methods(such as Support Vector Machine(SVM), Random Forest(RF)), thus providing the potential path for automatic recognition and pricing in self-service retail stores.
Chang Liu 0033, Jing Ni, Yu Cao 0002, Benyuan Liu
ICMLA4
2019 An Effective CNN Approach for Diabetic Retinopathy Stage Classification with Dual Inputs and Selective Data Sampling
abstract
Diabetic retinopathy (DR) is a vision-threatening complication among the diabetic population and a leading cause of blindness for working-age adults. Early detection and timely treatment can reduce the occurrence of blindness due to DR. Computer-aided diagnosis have great potential to significantly improve the accuracy and speed in DR detection over the traditional manual diagnosis process. In this paper, we present a deep convolutional neural network for DR stage classification, trained and evaluated on a large dataset. Our model uses high-resolution retinal fundus images of both the left and right eyes as inputs to take advantage of more detailed retinal lesion information in images and strong correlation between both eyes. Selective data sampling (SeS) is applied in the training process to mitigate the data imbalance problem. Experiments show that our model outperforms the fine-tuned Inception-v3 model by every measure, achieving an accuracy of 87.2% and a Kappa score of 0.806 on the Kaggle dataset.
Jing Ni, Qilei Chen, Chang Liu 0033, Yu Cao 0002, Benyuan Liu
ICMLA5
2019 Predicting Futures Market Movement using Deep Neural Networks
abstract
Recently there have been many efforts to study the predictability of financial market trend using various machine learning approaches. In this paper we explore the idea of using deep neural networks to analyze and predict futures market movements. Our approach adopts deep long short term memory (LSTM) as the main model architecture and predicts futures market movement using augmented market trading data. Training and testing of our model is performed in a rolling fashion to ensure the validity and reliability of the prediction. We discuss the design trade-off of several configurations and variations of our model, and evaluate the impact of various parameter choices as well as how model and backtesting perform under different parameter settings. We further design and implement a complete trading platform to evaluate our approach. Backtesting and live paper trading of our model on this platform achieves promising returns. Moreover, a total return of 58.69% is obtained with live paper trading for a twelve-month period when taking into account of slippage and commissions, which demonstrates the effectiveness of our proposed approach.
Tong Sun 0007, Jing Ni, Yu Cao 0002, Benyuan Liu
ICMLA4
2019 Colorectal Polyp Segmentation by U-Net with Dilation Convolution
abstract
Colorectal cancer (CRC) is one of the most commonly diagnosed cancers and a leading cause of cancer deaths in the United States. Colorectal polyps that grow on the intima of the colon or rectum is an important precursor for CRC. Currently, the most common way for colorectal polyp detection and precancerous pathology is the colonoscopy. Therefore, accurate colorectal polyp segmentation during the colonoscopy procedure has great clinical significance in CRC early detection and prevention. In this paper, we propose a novel end-to-end deep learning framework for the colorectal polyp segmentation. The model we design consists of an encoder to extract multi-scale semantic features and a decoder to expand the feature maps to a polyp segmentation map. We improve the feature representation ability of the encoder by introducing the dilated convolution to learn high-level semantic features without resolution reduction. We further design a simplified decoder which combines multi-scale semantic features with fewer parameters than the traditional architecture. Furthermore, we apply three post processing techniques on the output segmentation map to improve colorectal polyp detection performance. Our method achieves state-of-the-art results on CVC-ClinicDB and ETIS-Larib Polyp DB.
Xinzi Sun, Dechun Wang, Yu Cao 0002, Benyuan Liu
ICMLA4
2019 Toward Secure, Privacy-Preserving, and Interoperable Medical Data Sharing via Blockchain
abstract
In the era of cloud computing and big data analysis, how to efficiently share and utilize medical information scattered across various care providers has become a critical problem. This paper proposes a new framework for sharing medical data in a secure and privacy-preserving way. This framework holistically integrates multi-authority attribute based encryption, blockchain and smart contract, as well as software defined networking to define and enforce sharing policies. Specifically in our framework, patients' medical records are encrypted and stored in hospital databases, where strict access controls are enforced with attribute based encryption coupled with privacy level classification. Our framework leverages blockchain technology to connect scattered private databases from participating hospitals for efficient and secure data provision, smart contracts to enable the business logic of clinical data usage, and software defined networking to revoke sharing privileges. The performance evaluation of our prototype demonstrates that the associated computation costs are reasonable in practice.
Yan Luo 0001, Yu Cao 0002, Jomol Mathew
ICPADS5
2019 Mini Lesions Detection on Diabetic Retinopathy Images via Large Scale CNN Features
abstract
Diabetic retinopathy (DR) is a diabetes complication that affects eyes. DR is a primary cause of blindness in working-age people and it is estimated that 3 to 4 million people with diabetes are blinded by DR every year worldwide. Early diagnosis have been considered an effective way to mitigate such problem. The ultimate goal of our research is to develop novel machine learning techniques to analyze the DR images generated by the fundus camera for automatically DR diagnosis. In this paper, we focus on identifying small lesions on DR fundus images. The results from our analysis, which include the lesion category and their exact locations in the image, can be used to facilitate the determination of DR severity (indicated by DR stages). Different from traditional object detection for natural images, lesion detection for fundus images have unique challenges. Specifically, the size of a lesion instance is usually very small, compared with the original resolution of the fundus images, making them diffcult to be detected. We analyze the lesion-vs-image scale carefully and propose a large-size feature pyramid network (LFPN) to preserve more image details for mini lesion instance detection. Our method includes an effective region proposal strategy to increase the sensitivity. The experimental results show that our proposed method is superior to the original feature pyramid network (FPN) method and Faster RCNN.
Qilei Chen, Xinzi Sun, Yu Cao 0002, Benyuan Liu
ICTAI4
2019 AFP-Net: Realtime Anchor-Free Polyp Detection in Colonoscopy
abstract
Colorectal cancer (CRC) is a common and lethal disease. Globally, CRC is the third most commonly diagnosed cancer in males and the second in females. For colorectal cancer, the best screening test available is the colonoscopy. During a colonoscopic procedure, a tiny camera at the tip of the endoscope generates a video of the internal mucosa of the colon. The video data are displayed on a monitor for the physician to examine the lining of the entire colon and check for colorectal polyps. Detection and removal of colorectal polyps are associated with a reduction in mortality from colorectal cancer. However, the miss rate of polyp detection during colonoscopy procedure is often high even for very experienced physicians. The reason lies in the high variation of polyp in terms of shape, size, textural, color and illumination. Though challenging, with the great advances in object detection techniques, automated polyp detection still demonstrates a great potential in reducing the false negative rate while maintaining a high precision. In this paper, we propose a novel anchor free polyp detector that can localize polyps without using predefined anchor boxes. To further strengthen the model, we leverage a Context Enhancement Module and Cosine Ground truth Projection. Our approach can respond in real time while achieving state-of-the-art performance with 99.36% precision and 96.44% recall.
Dechun Wang, Xinzi Sun, Yu Cao 0002, Benyuan Liu
ICTAI6
2019 An Efficient Spatial-Temporal Polyp Detection Framework for Colonoscopy Video
abstract
Recent computer-aided polyp detection systems showed its effectiveness to decrease the polyp miss rate in colonoscopy operations, which is helpful to reduce colorectal cancer mortality. However, traditional polyp detection approaches suffer from the following drawbacks: low precision and sensitivity caused by the variance of polyp's appearance, and the system may not be able to detect polyps in real time due to the high computation complexity of the detection algorithms. To alleviate those problems, we introduce a real-time detection framework that incorporates spatial and temporal information extracted from colonoscopy videos. Our framework consists of the following three components: 1) we adopt Single Shot MultiBox Detector (SSD) to generate the proposal bounding boxes in each video frame. 2) Simultaneously, we compute optical flow from neighboring frames to extract temporal information and generate another group of polyp proposals with the temporal detection network. 3) At last, the final result is generated by a fusion module that connects the end of both streams. Experimental results on ETIS-LARIB dataset demonstrate that our proposed approach reaches the state-of-the-art performance on polyp localization with real-time performance.
Xinzi Sun, Dechun Wang, Yu Cao 0002, Benyuan Liu
ICTAI5
2019 CLVSA: A Convolutional LSTM Based Variational Sequence-to-Sequence Model with Attention for Predicting Trends of Financial Markets
abstract
Financial markets are a complex dynamical system. The complexity comes from the interaction between a market and its participants, in other words, the integrated outcome of activities of the entire participants determines the markets trend, while the markets trend affects activities of participants. These interwoven interactions make financial markets keep evolving. Inspired by stochastic recurrent models that successfully capture variability observed in natural sequential data such as speech and video, we propose CLVSA, a hybrid model that consists of stochastic recurrent networks, the sequence-to-sequence architecture, the self- and inter-attention mechanism, and convolutional LSTM units to capture variationally underlying features in raw financial trading data. Our model outperforms basic models, such as convolutional neural network, vanilla LSTM network, and sequence-to-sequence model with attention, based on backtesting results of six futures from January 2010 to December 2017. Our experimental results show that, by introducing an approximate posterior, CLVSA takes advantage of an extra regularizer based on the Kullback-Leibler divergence to prevent itself from overfitting traps.
Tong Sun 0007, Benyuan Liu, Yu Cao 0002, Hongwei Zhu 0002
IJCAI4
2019 A Deep Reinforcement Learning Approach to Multi-Component Job Scheduling in Edge Computing
abstract
The following topics are dealt with: learning (artificial intelligence); mobile computing; wireless sensor networks; feature extraction; Internet of Things; optimisation; graph theory; social networking (online); telecommunication network topology; pattern clustering.
Zhi Cao 0009, Honggang Zhang 0003, Yu Cao 0002, Benyuan Liu
MSN3
2018 Financial Markets Prediction with Deep Learning
abstract
Financial markets are difficult to predict due to its complex systems dynamics. Although there have been some recent studies that use machine learning techniques for financial markets prediction, they do not offer satisfactory performance on financial returns. We propose a novel one-dimensional convolutional neural networks (CNN) model to predict financial market movement. The customized one-dimensional convolutional layers scan financial trading data through time, while different types of data, such as prices and volume, share parameters (kernels) with each other. Our model automatically extracts features instead of using traditional technical indicators and thus can avoid biases caused by selection of technical indicators and pre-defined coefficients in technical indicators. We evaluate the performance of our prediction model with strictly backtesting on historical trading data of six futures from January 2010 to October 2017. The experiment results show that our CNN model can effectively extract more generalized and informative features than traditional technical indicators, and achieves more robust and profitable financial performance than previous machine learning approaches.
Tong Sun 0007, Benyuan Liu, Yu Cao 0002
ICMLA4
2018 A New Deep Learning-Based Food Recognition System for Dietary Assessment on An Edge Computing Service Infrastructure
abstract
Literature has indicated that accurate dietary assessment is very important for assessing the effectiveness of weight loss interventions. However, most of the existing dietary assessment methods rely on memory. With the help of pervasive mobile devices and rich cloud services, it is now possible to develop new computer-aided food recognition system for accurate dietary assessment. However, enabling this future Internet of Things-based dietary assessment imposes several fundamental challenges on algorithm development and system design. In this paper, we set to address these issues from the following two aspects: (1) to develop novel deep learning-based visual food recognition algorithms to achieve the best-in-class recognition accuracy; (2) to design a food recognition system employing edge computing-based service computing paradigm to overcome some inherent problems of traditional mobile cloud computing paradigm, such as unacceptable system latency and low battery life of mobile devices. We have conducted extensive experiments with real-world data. Our results have shown that the proposed system achieved three objectives: (1) outperforming existing work in terms of food recognition accuracy; (2) reducing response time that is equivalent to the minimum of the existing approaches; and (3) lowering energy consumption which is close to the minimum of the state-of-the-art.
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma, Songqing Chen
IEEE Trans. Serv. Comput.2
2017 TX-CNN: Detecting tuberculosis in chest X-ray images using convolutional neural network
abstract
In Low and Middle-Income Countries (LMICs), efforts to eliminate the Tuberculosis (TB) epidemic are challenged by the persistent social inequalities in health, the limited number of local healthcare professionals, and the weak healthcare infrastructure found in resource-poor settings. The modern development of computer techniques has accelerated the TB diagnosis process. In this paper, we propose a novel method using Convolutional Neural Network(CNN) to deal with unbalanced, less-category X-ray images. Our method improves the accuracy for classifying multiple TB manifestations by a large margin. We explore the effectiveness and efficiency of shuffle sampling with cross-validation in training the network and find its outstanding effect in medical images classification. We achieve an 85.68% classification accuracy in a large TB image dataset, surpassing any state-of-art classification accuracy in this area. Our methods and results show a promising path for more accurate and faster TB diagnosis in LMICs healthcare facilities.
Chang Liu 0033, Yu Cao 0002, Marlon Fernandes de Alcântara, Benyuan Liu, Maria J. Brunette, Jesús Peinado, Walter H. Curioso
ICIP2
2017 Identify Non-fatigue State to Fatigue State Using Causality Measure During Game Play
Yuying Zhu 0005, Yi-Ning Wu, Hui Su, Sanqing Hu, Tong Cao, Yu Cao 0002
ICONIP (4)7
2017 Improved Multimodal Representation Learning with Skip Connections
abstract
Multimodal Deep Boltzmann Machines (DBMs) have demonstrated huge successes in multimodal representation learning tasks. During inference, DBMs function as Recurrent Neural Nets (RNNs) because of the intractable distributions. To learn the parameters, optimizations can alternatively be operated on these surrogate RNNs with "truncated message passing". As a consequence, the gradient will propagate through a long chain without any local guidance which can potentially affects the optimization procedure. In this paper, we address this problem by adding skip connections during back-propagation while keeping the forward propagation (inference) untouched. With skip connections, we implicitly assign local "targets" for the states of intermediate inference loops to approach. Applied to different training criteria on different data sets, we demonstrate the proposed algorithms can consistently help to train better models while at a lower cost of training time. Experimental results show that our algorithms can achieve state-of-the-art performance on the Multimedia Information Retrieval (MIR) Flickr data set.
Yu Cao 0002, Benyuan Liu, Yan Luo 0001
ACM Multimedia2
2016 DeepFood: Deep Learning-Based Food Image Recognition for Computer-Aided Dietary Assessment
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma
ICOST2
2016 PNN for EEG-based Emotion Recognition
abstract
The effort to integrate emotions into human-computer interaction (HCI) system has attracted broad attentions. Automatic emotion recognition enables the HCI to become more intelligent and user friendly. Although numerous studies have been performed in this field, emotion recognition is still an extremely challenging task, especially in real-world practice usage. In this work, probabilistic neural network (PNN), with advantage of simple, efficient, and easy to train, was employed to recognize emotions elicited by watching music videos from scalp EEG. The publicly available DEAP emotion database was used to validate our algorithms. The powers of 4 frequency bands of EEG were extracted as features. The results show that the mean classification accuracy of PNN is 81.21% for valence(≥5 and <;5) and 81.26% for arousal(≥5 and <;5) across 32 subjects, similar with the results of SVM. In addition, they demonstrate that higher frequency bands (beta and gamma) play more important role in emotion classification than lower ones (theta and alpha). For the purpose of practical emotion recognition system, we proposed a ReliefF-based channel selection algorithm to reduce the number of used channels for convenience in practical usage. The results show that while using PNN, the 98% of the maximum classification accuracy can be obtained with only 9 (for valence) and 8 (for arousal) best channels, however, 19 (for valence) and 14 (for arousal) channels are needed while using SVM.
Sanqing Hu, Yu Cao 0002, Robert Kozma 0001
SMC4
2016 Shortcomings/Limitations of Blockwise Granger Causality and Advances of Blockwise New Causality
abstract
Multivariate blockwise Granger causality (BGC) is used to reflect causal interactions among blocks of multivariate time series. In particular, spectral BGC and conditional spectral BGC are used to disclose blockwise causal flow among different brain areas in various frequencies. In this paper, we demonstrate that: 1) BGC in time domain may not necessarily disclose true causality and 2) due to the use of the transfer function or its inverse matrix and partial information of the multivariate linear regression model, both of spectral BGC and conditional spectral BGC have shortcomings and/or limitations, which may inevitably lead to misinterpretation. We then, in time and frequency domains, develop two new multivariate blockwise causality methods for the linear regression model called blockwise new causality (BNC) and spectral BNC, respectively. By several examples, we confirm that BNC measures are more reasonable and sensitive to reflect true causality or trend of true causality than BGC or conditional BGC. Finally, for electroencephalograph data from an epilepsy patient, we analyze event-related potential causality and demonstrate that both of the BGC and BNC methods show significant causality flow in frequency domain, but the spectral BNC method yields satisfactory and convincing results, which are consistent with an event-related time-frequency power spectrum activity. The spectral BGC method is shown to generate misleading results. Thus, we deeply believe that our new blockwise causality definitions as well as our previous NC definitions may have wide applications to reflect true causality among two blocks of time series or two univariate time series in economics, neuroscience, and engineering.
Sanqing Hu, Xinxin Jia, Wanzeng Kong, Yu Cao 0002
IEEE Trans. Neural Networks Learn. Syst.5
2016 Comparison Analysis: Granger Causality and New Causality and Their Applications to Motor Imagery
abstract
In this paper we first point out a fatal drawback that the widely used Granger causality (GC) needs to estimate the autoregressive model, which is equivalent to taking a series of backward recursive operations which are infeasible in many irreversible chemical reaction models. Thus, new causality (NC) proposed by Hu et al. (2011) is theoretically shown to be more sensitive to reveal true causality than GC. We then apply GC and NC to motor imagery (MI) which is an important mental process in cognitive neuroscience and psychology and has received growing attention for a long time. We study causality flow during MI using scalp electroencephalograms from nine subjects in Brain-computer interface competition IV held in 2008. We are interested in three regions: Cz (central area of the cerebral cortex), C3 (left area of the cerebral cortex), and C4 (right area of the cerebral cortex) which are considered to be optimal locations for recognizing MI states in the literature. Our results show that: 1) there is strong directional connectivity from Cz to C3/C4 during left- and right-hand MIs based on GC and NC; 2) during left-hand MI, there is directional connectivity from C4 to C3 based on GC and NC; 3) during right-hand MI, there is strong directional connectivity from C3 to C4 which is much clearly revealed by NC than by GC, i.e., NC largely improves the classification rate; and 4) NC is demonstrated to be much more sensitive to reveal causal influence between different brain regions than GC.
Sanqing Hu, Wanzeng Kong, Yu Cao 0002, Robert Kozma 0001
IEEE Trans. Neural Networks Learn. Syst.5
2015 Automatic Eating Detection using head-mount and wrist-worn accelerometers
abstract
Automatic Eating Detection (AED) provides an important tool to help users regulate their dietary behavior for many health applications, such as weight management. In this paper we propose an AED solution using a head-mount and a wrist-worn accelerometers that are commonly available in commercial wearable devices. Experimental results, using Google Glass and Pebble Watch, validated that the proposed approach is highly effective to detect head motion from chewing and to detect hand-to-mouth (HtM) gestures when eating, resulting in 89.5% to 95.1% detection accuracy. Further we combined the features from both devices to achieve 97% cross-person eating detection accuracy and the average error when predicting duration of eating meals was only 105 seconds.
Xu Ye, Yu Cao 0002
HealthCom3
2015 FAST: A fog computing assisted distributed analytics system to monitor fall for stroke mitigation
abstract
Fog computing is a recently proposed computing paradigm that extends Cloud computing and services to the edge of the network. The new features offered by fog computing (e.g., distributed analytics and edge intelligence), if successfully applied for pervasive health monitoring applications, has great potential to accelerate the discovery of early predictors and novel biomarkers to support smart care decision making in a connected health scenarios. While promising, how to design and develop real-word fog computing-based pervasive health monitoring system is still an open question. As a first step to answer this question, in this paper, we employ pervasive fall detection for stroke mitigation as a case in study. There are four major contributions in this paper: (1) to investigate and develop a set of new fall detection algorithms, including new fall detection algorithms based on acceleration magnitude values and non-linear time series analysis techniques, as well as new filtering techniques to facilitate fall detection process; (2) to design and employ a real-time fall detection system employing fog computing paradigm, which distribute the analytics throughout the network by splitting the detection task between the edge devices (e.g., smartphones attached to the user) and the server (e.g., servers in the cloud); (3) we carefully exam the special needs and constraints of stroke patients and propose patient-centered design that is minimal intrusive to patients. This type of patient-centered design is currently lacking in most of the existing work; and (4) our experiments with real-word data show that our proposed system achieves the high sensitivity (low missing rate) while it also achieves the high specificity (low false alarm rate). At the same time, the response time and energy consumption of our system are close to the minimum of the existing approaches.
Yu Cao 0002, Songqing Chen, Donald Brown
NAS1
2015 HeteroSpark: A heterogeneous CPU/GPU Spark platform for machine learning algorithms
abstract
Analytics algorithms on big data sets require tremendous computational capabilities. Spark is a recent development that addresses big data challenges with data and computation distribution and in-memory caching. However, as a CPU only framework, Spark cannot leverage GPUs and a growing set of GPU libraries to achieve better performance and energy efficiency. We present HeteroSpark, a GPU-accelerated heterogeneous architecture integrated with Spark, which combines the massive compute power of GPUs and scalability of CPUs and system memory resources for applications that are both data and compute intensive. We make the following contributions in this work: (1) we integrate the GPU accelerator into current Spark framework to further leverage data parallelism and achieve algorithm acceleration; (2) we provide a plug-n-play design by augmenting Spark platform so that current Spark applications can choose to enable/disable GPU acceleration; (3) application acceleration is transparent to developers, therefore existing Spark applications can be easily ported to this heterogeneous platform without code modifications. The evaluation of HeteroSpark demonstrates up to 18× speedup on a number of machine learning applications.
Yan Luo 0001, Yu Cao 0002
NAS4
2015 Multi-level sample importance ranking based progressive transmission strategy for time series body sensor data
abstract
Body sensors have gained increasing interest during the past several years. With more applications deployed, it is imperative to ensure the success of data analysis, which largely depends on data transmission reliability as well as the importance of samples received. Traditional approaches focus on improving data reliability through various schemes such as prioritization of MAC access. In this paper, we analyzed the characteristics of time series body sensor data and propose to rank sample importance based on a multi-level approach. With this approach, samples are grouped into five levels, indicating their importance with regard to data analysis. Then, a progressive transmission strategy is designed to transmit samples in order of their importance so that the overall received data quality is maximized. Preliminary simulation results indicate that as much as 40-60% bandwidth saving can be achieved while meeting the requirements of data analysis algorithms.
Ming Li 0007, Yu Cao 0002, B. Prabhakaran 0001
WOWMOM2
2014 Accelerator of Stacked Convolutional Independent Subspace Analysis for Deep Learning-Based Action Recognition
abstract
Action recognition has been a research challenge in multimedia computing and machine vision. Recent advances in deep learning combined with stacked convolutional Independent Subspace Analysis (ISA) has achieved a better performance superior to all previously published results on several public available data sets. Unfortunately, one major issue in large-scale deployment of this new deep learning-based approach is the unacceptable latency of training with high-dimension data. In this paper, we propose a new hardware accelerator that can reduce the training time substantially for deep learning-based action recognition. Specifically, our proposed approach focuses on accelerating the convolutional stacked ISA algorithm, the core components of the deep learning-based action recognition algorithms. We design parallel pipelines, data parallelisms and look-up table to speed up the algorithm. With an embedded heterogeneous platform consisting of a general purpose processor and a FPGA, we are able to achieve up to 10X speedup for stacked ISA training compared to a software-only implementation.
Yan Luo 0001, Yu Cao 0002
FCCM3
2014 Causality from Cz to C3/C4 or between C3 and C4 revealed by granger causality and new causality during motor imagery
abstract
Interaction between different brain regions has received wide attention recently. Granger causality (GC) is one of the most popular methods to explore causality relationship between different brain regions. In 2011, Hu et. al [1] pointed out shortcomings and/or limitations of GC by using a large of number of illustrative examples and showed that GC is only a causality definition in the sense of Granger and does not reflect real causality at all, and meanwhile proposed a new causality (NC) which is shown to be more reasonable and understandable than GC by those examples. Motor imagery (MI) is an important mental process in cognitive neuroscience and cognitive psychology and has received growing attention for a long time. However, there is few work about causality flow so far during MI based on scalp EEG. In this paper, we use scalp EEG to study causality flow during MI. The scalp EEGs are from 9 subjects in BCI competition IV held in 2008 [2] and provided by Graz University of Technology. We are interested in three regions: Cz (the centre of cerebral cortex), C3 (the left of cerebral cortex) and C4 (the right of cerebral cortex) which are considered to be optimal locations for recognizing MI states in literature. We apply GC and NC to scalp EEG and find that i) there is strong directional connectivity from Cz to C3/C4 during left hand and right hand MI based on GC and NC. ii) During left hand MI, there is directional connectivity from C4 to C3 based on GC and NC. iii) During right hand MI, there is strong directional connectivity from C3 to C4 which is much clearly revealed by NC method than by GC method. iv) Our results suggest that NC method in time and frequency domains is demonstrated to be much better to reveal causal influence between different brain regions than GC method. Thus, we deeply believe that NC method will shed new light on causality analysis in economics and neuroscience.
Sanqing Hu, Wanzeng Kong, Yu Cao 0002
IJCNN5
2014 Smartphone-Based Walking Speed Estimation for Stroke Mitigation
abstract
Each year, 15 million people suffer stroke worldwide. Among them, 5 million die and another 5 million are permanently disabled. Stroke recovery is a lifelong process. Clinical research have shown that gait velocity (a.k.a, walking speed) is a very powerful indicator of function and prognosis after stroke. In this paper, we focus on developing new algorithms to estimate the walking speed using pervasive devices, such as smartphone. While there are existing techniques to measure walking speed using inertial sensors, very little research has specifically involved smartphone, due to some unique challenges caused by pervasive devices, such as placement of the sensor and sensor drifting. We propose new practical algorithms based on high pass filter, integration of accelerator's reminder, and feedback loop. We evaluate our proposed approach with real world data and present a through analysis on the results. The experimental results have indicated that proposed approach is a promising practical approach for gait speed estimation.
Jeffrey Cox, Yu Cao 0002, Degui Xiao
ISM2
2013 Motor imagery classification based on joint regression model and spectral power
Sanqing Hu, Qiangqiang Tian, Yu Cao 0002, Wanzeng Kong
Neural Comput. Appl.3
2012 Multimedia data semantics: guest editors introduction
Raphaël Troncy, B. Prabhakaran 0001, Yu Cao 0002
Multim. Tools Appl.3
2011 Medical multimedia analysis and retrieval
abstract
Advances in sensor technology, processing speed, high-speed networking, and the massive digital storages are being incorporated into today's healthcare practice. Tremendous amounts of medical multimedia data are captured and recorded in digital format during the daily clinical practice, medical research, and education. Intelligent medical knowledge discovery and retrieval from medical multimedia data is very useful and highly desirable. Motivated by the huge potential benefits, we have organized the ACM International Workshop on Medical Multimedia Analysis and Retrieval (MMAR2011), held in conjunction with The Annual ACM International Conference on Multimedia (ACM Multimedia 2011). In this paper, we first introduce the background and overview of the workshop. Then we introduce the workshop review process, followed with a brief introduction of the accepted papers in this workshop, as well as the conclusion.
Yu Cao 0002, Jayashree Kalpathy-Cramer, Devrim Ünay
ACM Multimedia1
2010 Introduction to the special issue on "data semantics for multimedia systems"
Mei-Ling Shyu, Yu Cao 0002, Ming Li 0007, Mathias Lux, Jie Bao 0001
Multim. Tools Appl.2
2008 Audio-visual event classification via spatial-temporal-audio words
abstract
In this paper, we propose a generative model-based approach for audio-visual event classification. This approach is based on a new unsupervised learning method using an extended probabilistic latent semantic analysis (pLSA) model. We represent each video clip as a collection of spatial-temporal-audio words, which are generated by fusing the visual and audio features using the pLSA model. Each audio-visual event class is treated as the latent topic in this model. The probability distributions of the spatial-temporal-audio words are learnt from training examples, which include a sequence of videos that represent different types of audio-visual events. Experimental results show the effectiveness of the proposed approach.
Yu Cao 0002, Sung Baang, Shih-Hsi Liu, Ming Li 0007, Sanqing Hu
ICPR1
2008 Medical Video Event Classification Using Shared Features
abstract
Advances in video technology are being incorporated into today’s medical research and education. Medical videos contain important medical events, such as diagnostic or therapeutic operations. Automatic discovery and classification of these events are highly desirable and very useful. In this paper, we present a novel method for multi-class educational medical video event categorization. Our method employs a learning procedure based on boosted decision stumps. There are two key contributions in this paper. The first contribution is that the proposed multi-class boosting algorithms utilize the common features which can be shared among different video event categories. Compared with the class-specific features, the entire set of shared features can provide more efficient and reliable representation to classify multiple video event categories. The second key contribution of this paper is the adaption of the spacetime interest point detection techniques for feature extraction on both the spatial dimension and the temporal dimension. Experimental results have shown that the proposed approach is a very promising strategy for solving the multi-class video event classification problem.
Yu Cao 0002, Shih-Hsi Liu, Ming Li 0007, Sung Baang, Sanqing Hu
ISM1
2008 A Semantics- and Data-Driven SOA for Biomedical Multimedia Systems
abstract
Due to the problems of heterogeneous data and platforms, abundant functional and QoS requirements, high data size, and tangling correlation between data/contents and software functionalities, developing large-scale biomedical multimedia database systems is a challenging task. This paper presents a semantics- and data-driven service-oriented architecture (SOA) to take the interoperability and scalability advantages of conventional SOA and solve the aforementioned problems. By establishing data ontology with respect to data properties, contents, QoS, and biomedical regulations and expanding service ontology to describe more functional and QoS specifications supported by services, appropriate services for processing biomedical multimedia data may be discovered, performed, tuned up or replaced as needed. Additionally, six transmission services are introduced to support dynamic adaptation under specific requirements.
Shih-Hsi Liu, Yu Cao 0002, Ming Li 0007, Pranay Kilaru, Thell Smith, Shaen Toner
ISM2
2005 Automatic measurement of quality metrics for colonoscopy videos
abstract
Colonoscopy is the accepted screening method for detection of colorectal cancer or its precursor lesions, colorectal polyps. Indeed, colonoscopy has contributed to a decline in the number of colorectal cancer related deaths. However, not all cancers or large polyps are detected at the time of colonoscopy, and methods to investigate why this occurs are needed. We present a new computer-based method that allows automated measurement of a number of metrics that likely reflect the quality of the colonoscopic procedure. The method is based on analysis of a digitized video file created during colonoscopy, and produces information regarding insertion time, withdrawal time, images at the time of maximal intubation, the time and ratio of clear versus blurred or non-informative images, and a first estimate of effort performed by the endoscopist. As these metrics can be obtained automatically, our method allows future quality control in the day-to-day medical practice setting on a large scale. In addition, our method can be adapted to other healthcare procedures. Last but not least, our method may be useful to assess progress during colonoscopy training, or as part of endoscopic skills assessment evaluations.
Sae Hwang, Jung-Hwan Oh 0001, JeongKyu Lee, Yu Cao 0002, Wallapak Tavanapong, Danyu Liu, Johnny S. Wong, Piet C. de Groen
ACM Multimedia4
2004 A framework for parsing colonoscopy videos for semantic units
abstract
Colonoscopy is an important screening procedure for colorectal cancer. During this procedure, the endoscopist visually inspects the colon. Currently, there is no content-based analysis and retrieval system that automatically analyzes videos captured from colonoscopic procedures and provides a user-friendly and efficient access to important content. Such a system will be valuable for endoscopic research and education. The first necessary step for the analysis is parsing for semantic units. Since the characteristics of colonoscopy videos differ from those of videos studied in the literature, we introduce a new video parsing framework that includes: (i) a new scene definition and a new video parsing paradigm; (ii) a novel scene segmentation algorithm using audio analysis and finite state automata to recognize scenes and associated boundaries. Our experimental results show average precision and recall of 95% and 81%, respectively, for parsing scenes. The framework is extensible to videos captured from other endoscopic procedures such as upper gastrointestinal endoscopy, enteroscopy, cystoscopy, and laparoscopy.
Yu Cao 0002, Wallapak Tavanapong, Johnny S. Wong, Jung-Hwan Oh 0001, Piet C. de Groen
ICME1
2004 Parsing and browsing tools for colonoscopy videos
abstract
Colonoscopy is an important screening tool for colorectal cancer. During a colonoscopic procedure, a tiny video camera at the tip of the endoscope generates a video signal of the internal mucosa of the colon. The video data are displayed on a monitor for real-time analysis by the endoscopist. We call videos captured from colonoscopic procedures colonoscopy videos. Because these videos possess unique characteristics, new types of semantic units and parsing techniques are required. In this paper, we define new semantic units called operation shots, each is a segment of visual and audio data that correspond to a therapeutic or biopsy operation. We introduce a new spatio-temporal analysis technique to detect operation shots. Our experiments on colonoscopy videos demonstrate that the technique does not miss any meaningful operation shots and incurs a small number of false operation shots. Our prototype parsing software implements the operation shot detection technique along with our other techniques previously developed for colonoscopy videos. Our browsing tool enables users to quickly locate operation shots of interest. The proposed technique and software are useful (1) for post-procedure reviews and analyses for causes of complications due to biopsy or therapeutic operations, (2) for developing an effective content-based retrieval system for colonoscopy videos to facilitate endoscopic research and education, and (3) for development of a systematic approach to assess endoscopists' procedural skills.
Yu Cao 0002, Dalei Li, Wallapak Tavanapong, Jung-Hwan Oh 0001, Johnny S. Wong, Piet C. de Groen
ACM Multimedia1