VLDB 2026 Research / reviewers in the wild / expert
Guitao Cao
dblp:09/2151
· DBLP profile ↗
67ranked-venue papers
4as first author
50since 2021 · last 2026
0000-0002-4059-4806ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 1 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 26 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 11 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic prototype with discriminative representation for rapid adaptation in new organ segmentation
Xinyue Zhang 0010, Guitao Cao, Wenming Cao 0001 |
Pattern Recognit. | 4 |
| 2025 | A Preliminary Exploration of Children Autism Spectrum Disorder Detection Based on Environmental Variables
Guitao Cao, Qiaoyun Liu, Xidong Xi |
ICIC (28) | 2 |
| 2025 | Counterfactual Thinking Driven Emotion Regulation for Image Sentiment RecognitionabstractImage sentiment recognition (ISR) facilitates the practical application of affective computing on rapidly growing social platforms. Nowadays, region-based ISR methods that use affective regions to guide emotion prediction have gained significant attention. However, existing methods lack a causality-based mechanism to guide affective region generation and effective tools to quantitatively evaluate their quality. Inspired by the psychological theory of Emotion Regulation, we propose a counterfactual thinking driven emotion regulation network (CTERNet), which simulates the Emotion Regulation Theory by modeling the entire process of ISR based on human causality-driven mechanisms. Specifically, we first use multi-scale perception for feature extraction to simulate the stage of situation selection. Next, we combine situation modification, attentional deployment, and cognitive change into a counterfactual thinking based cognitive reappraisal module, which learns both affective regions (factual) and other potential affective regions (counterfactual). In the response modulation stage, we compare the factual and counterfactual outcomes to encourage the network to discover the most emotionally representative regions, thereby quantifying the quality of affective regions for ISR tasks. Experimental results demonstrate that our method outperforms or matches the state-of-the-art approaches, proving its effectiveness in addressing the key challenges of region-based ISR. Xinyue Zhang 0010, Zhaoxia Wang 0001, Guitao Cao |
IJCAI | 4 |
| 2025 | A Novel Framework for Inverse Problems: Fixed-Point Iteration Using Consistency ModelsabstractInverse problems play a crucial role in science and engineering, especially in the field of computer vision, where tasks such as deblurring, super-resolution, and colorization can be formally modeled as inverse problems. Consistency models excel in generation speed while maintaining high quality, making them a promising family of generative models. However, existing sampling methods struggle to achieve high-quality results when applying consistency models to image inverse problems. To address this limitation, we propose the Consistency Inverse Reconstruction Sampling (CIRS) framework, which incorporates two modes: CIRS-Hybrid and CIRS-Pure. In CIRS-Hybrid, the posterior formula of inverse problems is utilized by estimating the prior term using a diffusion denoiser and the likelihood term with a consistency model, enabling reconstruction under dual-model guidance. To overcome the complexities of dual-model tuning and inefficiencies caused by employing a diffusion denoiser, we introduce CIRS-Pure, which relies solely on a consistency model. By eliminating the iterative noise addition and denoising steps, the iterative procedure is transformed into a fixed-point iteration, achieving efficient and high-quality restoration. Extensive experiments demonstrate that CIRS-Pure outperforms state-of-the-art methods in zero-shot image restoration tasks such as image deblurring and colorization while achieving competitive performance in super-resolution. Xinke Wang, Guitao Cao |
IJCNN | 2 |
| 2025 | CLMP: Cross Learning Multi-head Prediction Guided Source-Free Domain Adaptation for Medical Image SegmentationabstractUnsupervised Domain Adaptation is crucial for transferring segmentation models trained on the source domain to the target domain. However, in medical image segmentation, the unavailability of source domain data makes Source-Free Domain Adaptation (SFDA) an attractive alternative, allowing model adaptation without relying on source domain data. Current SFDA methods predominantly rely on pseudo-labels, yet obtaining high-quality pseudo-labels continues to pose significant challenges. In this paper, we propose a novel framework for medical image segmentation, termed Cross-Learning Multi-head Prediction-guided (CLMP) SFDA. This framework leverages consistency constraints within a multi-head prediction architecture to enhance pseudo-label accuracy. Specifically, we deploy a multi-head prediction model designed to generate robust pseudo-labels. To minimize discrepancies between the predictions of different heads, we propose a method called Pixel-Level Synergistic Consistency Loss (PSCL). Additionally, we implement dual forward propagation to counteract model degradation during self-supervised training and integrate Chebyshev uncertainty estimation to selectively filter out unreliable pseudo-labels. Moreover, we introduce a foreground-background statistical weighting module to tackle class imbalance effectively. Experiments conducted on two public medical image datasets demonstrate an average performance improvement of over 6% across all metrics compared to existing state-of-the-art methods, underscoring the effectiveness of the CLMP framework. Wenqiang Lin, Guitao Cao |
IJCNN | 4 |
| 2025 | Dual Knowledge-Aware Guidance for Source-Free Domain Adaptive Fundus Image Segmentation
Chunwei Wu, Guitao Cao |
MICCAI (6) | 4 |
| 2025 | C2BA: Cross-Domain Consistency and Bidirectional Alignment for Cross-Modal Domain-Incremental LearningabstractIn Cross-Modal Domain-Incremental learning, the primary challenge lies in learning from varying data distributions and maintaining its performance on prior domains. However, existing methods often overlook the importance of shared knowledge across domains and the interaction between modalities is still insufficient. To address these issues, we propose Cross-Domain Consistency and Bidirectional Alignment (C2BA), a novel framework that enhances the model's generalization ability and improves the cross-modal integration in VLMs through two key components. We design a Cross-domain Global Consistency Constraint (CGCC) to stabilize domain-invariant representations during incremental training, preventing excessive shifts of shared distributions toward new domains. In addition, we design a Bidirectional Cross-Modal Attention (BCMA) module, which enables effective interaction between visual and textual features through a bidirectional attention mechanism, thereby reducing cross-modal discrepancies. Experiments on three benchmark datasets demonstrate that our method outperforms state-of-the-art exemplar-free and even exemplar-based approaches, achieving superior generalization and cross-modal interaction. Xidong Xi, Guitao Cao |
SMC | 4 |
| 2025 | Tagging knowledge concepts for math problems based on multi-label text classification
Yuzhuo Wu, Guitao Cao, Liangyu Chen 0001 |
Expert Syst. Appl. | 4 |
| 2025 | Potential Knowledge Extraction Network for Class-Incremental Learning
Xidong Xi, Guitao Cao, Wenming Cao 0001, Yan Li 0063 |
Neurocomputing | 2 |
| 2025 | Uncertainty guided semi-supervised few-shot segmentation with prototype level fusion
Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
Neural Networks | 4 |
| 2025 | NLA-GNN: Non-local information aggregated graph neural network for heterogeneous graph embedding
Siheng Wang, Guitao Cao, Wenming Cao 0001, Yan Li 0063 |
Pattern Recognit. | 2 |
| 2025 | Adaptive Meta-Path Selection Based Heterogeneous Spatial Enhancement for circRNA-Disease Associations PredictionabstractCircular RNAs (circRNAs) play a crucial role in human biological processes as miRNA sponges, regulating gene expression and affecting disease manifestations. Establishing heterogeneous nodal feature relationships through meta-paths can effectively enhance the predictive capability of models. Previous research primarily constructed associations between meta-paths manually, and excessive noise made it difficult to capture highly correlated hidden features. Equally important is learning more about feature distributions, which is key to improving the generalization ability of algorithms. To address these challenges, we propose an adaptive meta-path selection method named AdaMH. The core scheme introduces an adaptive path selection method that automatically identifies highly relevant heterogeneous meta-paths during iterative training rounds. Considering the sparsity of data distribution, we introduce controlled random noise into the data through graph contrastive learning to ensure an even distribution of features. Subsequently, a multi-head attention mechanism is utilized to capture relationships in the high-dimensional heterogeneous feature space, enhancing feature representation capability. Comparing with state-of-the-art (SOTA) algorithms, AdaMH is the only one that surpasses a performance threshold of 0.95 across seven evaluation metrics. Guitao Cao, Wenming Cao 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | HyperEditor: Achieving Both Authenticity and Cross-Domain Capability in Image Editing via HypernetworksabstractEditing real images authentically while also achieving cross-domain editing remains a challenge. Recent studies have focused on converting real images into latent codes and accomplishing image editing by manipulating these codes. However, merely manipulating the latent codes would constrain the edited images to the generator's image domain, hindering the attainment of diverse editing goals. In response, we propose an innovative image editing method called HyperEditor, which utilizes weight factors generated by hypernetworks to reassign the weights of the pre-trained StyleGAN2's generator. Guided by CLIP's cross-modal image-text semantic alignment, this innovative approach enables us to simultaneously accomplish authentic attribute editing and cross-domain style transfer, a capability not realized in previous methods. Additionally, we ascertain that modifying only the weights of specific layers in the generator can yield an equivalent editing result. Therefore, we introduce an adaptive layer selector, enabling our hypernetworks to autonomously identify the layers requiring output weight factors, which can further improve our hypernetworks' efficiency. Extensive experiments on abundant challenging datasets demonstrate the effectiveness of our method. Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
AAAI | 3 |
| 2024 | CausVSR: Causality Inspired Visual Sentiment Recognition
Xinyue Zhang 0010, Zhaoxia Wang 0001, Jing Xiang, Chunwei Wu, Guitao Cao |
IJCAI | 6 |
| 2024 | HWSformer: History Window Serialization Based Transformer for Semantic Enrichment Driven Stock Market PredictionabstractAfter the Transformer model demonstrated excellent performance in natural language processing (NLP) tasks and computer vision tasks, people have started to explore the use of Transformer models in the field of time series prediction. Because of the significant role of the stock market in the global economy, stock market prediction is of paramount importance for investors. Stock indices forecasting is one of the fields of stock market forecasting and researchers have also set their sights on Transformer. However, with limited semantic information available in time series data and the unique characteristics of the self-attention mechanism, the Transformer model has not gained widespread adoption in stock indices forecasting. In this paper, we propose a history window serialization based Transformer model (HWSformer) specifically designed for predicting stock price indices. Our innovation is to introduce the historical window serialization layer to solve the problem of limited semantic richness in time series data, which affects the validity of self-attention. Additionally, in order to capture the original distribution accurately and retain the valuable non-stationary information, we incorporate the Reversible Instance Normalization (RevIN) method. We conducted experiments on 12 stock price index datasets collected from multiple countries and demonstrated that HWSformer outperforms traditional Transformer models by approximately 20% and varying degrees of improvement compared to other recent variants of Transformers. Yisheng Hu, Guitao Cao, Dawei Cheng |
IJCNN | 2 |
| 2024 | Self-supported Prototype Rectification for Few-shot Medical Image SegmentationabstractFew-shot semantic segmentation aims to quickly adapt to pixel-wise predictions for novel classes with only a few labeled images. Recent works rely on prototypical learning, where prototypes obtained from support images are applied to the segmentation of query images. However, there are inherent intra-class appearance differences between support images and query images, and the prototypes extracted from a small number of support images contain limited deep semantic information, which makes it difficult to accurately guide the segmentation of query images. To alleviate this problem, we propose a Self-Supported Prototype Rectification Network. Specifically, we introduce a Pseudo Mask Generation (PMG) module to generate a pseudo query mask by means of many-to-many prototype matching. We design a Prototype Rectification (PR) module with a learnable parameter λ to balance self-supported rectified prototype between support prototype obtained from support image and query prototype extracted from query features with pseudo query mask. Furthermore, we introduce a prototype-based multi-class segmentation approach mitigate the issue of confusion area prediction among different organs for query images in multi-organ segmentation scenario. Our method outperforms other SOTAs on two widely used datasets: CHAOST2 and MS-CMR. Zhaoxu Li, Guitao Cao |
IJCNN | 3 |
| 2024 | DLE: Document Illumination Correction with Dynamic Light EstimationabstractDocument images captured through mobile devices in natural environments are often affected by various types of illumination degradation. The degradation diminishes the clarity and readability of document images, thereby complicating their application to OCR downstream tasks. Existing methods typically address only one or a limited number of degradation types and do not consider the diversity of image degradation types. Additionally, these methods typically involve a pre-trained fixed sub-network to estimate background light or shadows, which lacks flexibility and adaptability. To overcome these challenges, this study proposes a novel framework named DLE, which comprises a two-loop generative adversarial network and a multi-modal discriminator. Specifically, to improve the quality of image representation, a mask extractor is embedded before the image input generator. This forces the model to focus on the distinct features in the image, enhancing the representation of illumination anomalous and degraded regions. The mask extractor generates a luminance mask to evaluate the difference in illumination between the input and target images. Subsequently, the consistency loss computation incorporates a dynamic optimization of the mask extractor, strengthening its ability to estimate the illumination degradation part. Moreover, a pre-trained visual-language model is introduced into the multi-modal discriminator, leveraging its robust cross-modal alignment capability to improve the semantic consistency of the generated images with the preset input text. Extensive experiments demonstrate that our approach achieves the SOTA performance in terms of edit distance (ED) and character error rate (CER). Jiahao Quan, Chunwei Wu, Guitao Cao |
SMC | 4 |
| 2024 | Enhanced Slicing Prototype and Hybrid Metric Transformer for Few-shot Medical Image ClassificationabstractAs one of the most popular neural network modules, Transformer plays a key role in many fundamental deep learning models such as few-shot medical image segmentation, which aims to segment the target objects in query under the condition of a few annotated support images. Most previous works strive to mine more semantically effective information from the support to match with the corresponding objects in query. The traditional models generally input the whole image into the deep neural network to obtain the feature representation, and use only one measurement method to improve efficiency. If the objects in them show large intra-class diversity, the discrepancy gap between query and support images is ignored. To solve this problem, we propose an enhanced slicing prototype and multidimensional metric mechanism to address the inefficiency of existing few-shot learning methods in medical image classification. Instead of whole image is input into the deep neural network, our proposed model segments the image into slices, and then use the self-attention mechanism to generate enhanced feature vectors based on transformer. And then, a hybrid metric is used to measure similarity between features by calculating the distance between the support set and query set slice prototypes to improve efficiency. Experiments demonstrate that our model has better classification effect on mini-MedMNIST, which is a few-shot medical image dataset constructed from MedMNIST dataset. Guitao Cao |
SMC | 3 |
| 2024 | MVESF: Multi-View Enhanced Semantic Fusion for Controllable Text-to-Image GenerationabstractText-to-image(T2I) synthesis aims to generate semantically consistent images with texts. Currently, existing methods merely use the shallow semantics of the text to crudely guide image generation. They cannot fully integrate rich textual semantics with image features, leading to an inability to control image generation through text finely. To address this issue, we propose a GAN-based method named MVESF (Multi-View Enhanced Semantic Fusion), which enhances the semantic fusion of text and images from multiple perspectives for fine-grained controllable text-to-image synthesis. In multi-view, we introduce Multi-domain Semantic Guidance, Local Semantic Attention, and Visual-textual Consistency Loss to enhance the semantic fusion of text and images in image generation, image discrimination, and image supervision, respectively. Our method promotes the consistent alignment between text and images, allowing for fine-grained variations in the generated images when subtle changes in the input text without affecting unrelated regions. Extensive experimental results have demonstrated the effectiveness of our approach. Guitao Cao, Xinke Wang, Jiahao Quan |
SMC | 2 |
| 2024 | KD loss: Enhancing discriminability of features with kernel trick for object detection in VHR remote sensing images
Xi Chen 0004, Liyue Li, Qingli Li, Honggang Qi, Ying Wen 0003, Guitao Cao, Philip L. H. Yu |
Eng. Appl. Artif. Intell. | 9 |
| 2024 | Resource-aware Montgomery modular multiplication optimization for digital signal processing
Qiqi Tao, Liying Li 0002, Junlong Zhou, Guitao Cao, Dan Meng 0001 |
J. Syst. Archit. | 4 |
| 2024 | A Trustworthy Counterfactual Explanation Method With Latent Space SmoothingabstractDespite the large-scale adoption of Artificial Intelligence (AI) models in healthcare, there is an urgent need for trustworthy tools to rigorously backtrack the model decisions so that they behave reliably. Counterfactual explanations take a counter-intuitive approach to allow users to explore "what if" scenarios gradually becoming popular in the trustworthy field. However, most previous work on model's counterfactual explanation cannot generate in-distribution attribution credibly, produces adversarial examples, or fails to give a confidence interval for the explanation. Hence, in this paper, we propose a novel approach that generates counterfactuals in locally smooth directed semantic embedding space, and at the same time gives an uncertainty estimate in the counterfactual generation process. Specifically, we identify low-dimensional directed semantic embedding space based on Principal Component Analysis (PCA) applied in differential generative model. Then, we propose latent space smoothing regularization to rectify counterfactual search within in-distribution, such that visually-imperceptible changes are more robust to adversarial perturbations. Moreover, we put forth an uncertainty estimation framework for evaluating counterfactual uncertainty. Extensive experiments on several challenging realistic Chest X-ray and CelebA datasets show that our approach performs consistently well and better than the existing several state-of-the-art baseline approaches. Yan Li 0063, Chunwei Wu, Xiao Lin 0012, Guitao Cao |
IEEE Trans. Image Process. | 5 |
| 2023 | Calibrated Uncertainty-Guided Multi-task Framework for Medical Image SegmentationabstractMedical image segmentation is a crucial part of computer-aided diagnosis. Due to the enormous cost of labeling medical images, researchers have turned to exploring semi-supervised learning. However, the lack of supervisory information makes it difficult to accurately segment the fuzzy regions (e.g., complex edges or corners of organs). In this paper, we propose a novel method called Multi-task Consistency Segmentation Network based on Calibrated Uncertainty (CU-MCSNet). This model incorporates calibrated uncertainty to guide the network’s learning process. In addition, the model consists of two tasks: i) semantic segmentation as the primary task and ii) signed distance regression as the auxiliary task. To enhance the accuracy of edge segmentation, we propose the Edge Calibration Network for the primary task. This network integrates essential spatial and channel features, employing gradient complementation to hinder the accumulation of defective information and supply pertinent data to fuzzy regions. We also use the inter-task consistency loss to explore the underlying information of the images. In the multi-task domain, it is tough to balance each task manually, and we note that homoscedastic uncertainty focuses on inter-task variation. However, its numerical estimation may still be subject to bias. Therefore, we propose an adaptive loss-balancing strategy based on calibrated homoscedastic uncertainty. Extensive experiments show that our proposed method achieves state-of-the-art performance. Chunwei Wu, Guitao Cao |
BIBM | 4 |
| 2023 | FedPerturb: Covert Poisoning Attack on Federated Learning via Partial PerturbationabstractFederated learning breaks through the barrier of data owners by allowing them to collaboratively train a federated machine learning model without compromising the privacy of their own data. However, Federation Learning also faces the threat of poisoning attacks, especially from the client model updates, which may impair the accuracy of the global model. To defend against the poisoning attacks, previous work aims to identify the malicious updates in high dimensional spaces. However, we find that the distances in high dimensional spaces cannot identify the changes in a small subset of dimensions, and the small changes may affect the global models severely. Based on this finding, we propose an untargeted poisoning attack under the federated learning setting via the partial perturbations on a small subset of the carefully selected model parameters, and present two attack object selection strategies. We experimentally demonstrate that the proposed attack scheme achieves high attack success rate on five state-of-the-art defense schemes. Furthermore, the proposed attack scheme remains effective at low malicious client ratios and still circumvents three defense schemes with a malicious client ratio as low as 2%. Tongsai Jin, Zhihui Fu, Dan Meng 0001, Jun Wang 0020, Guitao Cao |
ECAI | 6 |
| 2023 | DCNet: Weakly Supervised Saliency Guided Dual Coding Network for Visual Sentiment RecognitionabstractVisual sentiment recognition is a challenging task with scientific significance in probing vision-processing mechanisms. Recent approaches mainly focused on using overall images or precise annotations to learn emotional representations, yet neglected to capture abstract semantics from regional information, or led to a heavy annotation burden. In this paper, we propose an end-to-end weakly supervised framework, called Dual Coding Network (DCNet), which models a dual coding process for both shallow features and high-level regional information. On the one hand, with the help of the fine-grained module (FG), visual features (e.g. texture features) are utilized to enhance the learning of distinguished representation. On the other hand, the DCNet innovatively leverages saliency information to imitate the neural decoding of perceived visual sentiment contents in human brain activity. Specifically, the saliency information guides the generation of sentiment-specific pseudo affective maps (SAMG), which serve as weak annotations. Then the DCNet couples fine-grained features with pseudo affective maps, and obtains semantic vectors for final sentiment prediction. Extensive experiments show that the proposed DCNet outperforms the state-of-the-art performance on five benchmark datasets. Xinyue Zhang 0010, Jing Xiang, Hanxiu Zhang, Chunwei Wu, Guitao Cao |
ECAI | 6 |
| 2023 | Making Adversarial Attack Imperceptible in Frequency Domain: A Watermark-based FrameworkabstractWith the development of multimedia communication technology, the image information stored in electronic devices faces increasing privacy risks and requires processing for protection. However, it is found that adversarial perturbations added to images for semantic information protection may corrupt the frequency domain watermarks added for copyright statement. With such challenges, we propose an Adversarial Frequency domain Watermarking (AFW) framework to protect images from both copyright and semantic content. Specifically, the AFW framework constructs the images as adversarial examples by embedding crafted adversarial watermarks in the frequency domain, followed by an optimization algorithm to improve the visual quality. Notably, AFW can generally integrate with existing watermark and attack methods. Extensive experiments on five network models and the ImageNet dataset demonstrate that the AFW framework can achieve information hiding and adversarial attacking goals under visual quality assurance. Hanxiu Zhang, Guitao Cao, Xinyue Zhang 0010, Jing Xiang, Chunwei Wu |
ICME | 2 |
| 2023 | Class-Balanced Universal Perturbations for Adversarial TrainingabstractUniversal attack generates image-agnostic perturbation called universal adversarial perturbation (UAP), which can be added to all samples in the data distribution to fool the classifier. However, a universal perturbation will likely mislead the classifier to identify most adversarial examples as the same label, resulting in the imbalance of attack strength between classes. In this paper, we propose class-balanced UAPs that enlarge the dispersion of the predicted labels for adversarial examples. To ensure attack strength and balance simultaneously, we design a novel diversity objective containing probability calibration and penalty regularizer, which fully considers the predicted label distribution between samples and the predicted probability distribution within samples. Furthermore, we apply class-balanced attacks in adversarial training to defend against universal perturbations since the class-balanced UAP provides diverse perturbation directions. We correspondingly reformulate adversarial training from the min-max optimization problem into a new two-stage framework. Experiments on several benchmark datasets demonstrate that the class-balanced attack achieves better performance than the universal attack, while adversarial training with class-balanced UAP achieves state-of-the-art results in clean accuracy and robustness to universal perturbations. Kexue Ma, Guitao Cao, Mengqian Xu, Chunwei Wu, Wenming Cao 0001 |
IJCNN | 2 |
| 2023 | Discriminative Feature Mining and Alignment for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain by reducing the cross-domain distribution discrepancy. Existing UDA methods mainly learn domain-invariant features by directly aligning the marginal distribution of source and target domains. However, they ignore mining the discriminative information of target data and aligning the cross-domain discriminative features, which may lead to performance degradation. To tackle these two issues simultaneously, we propose a Discriminative Feature Mining and Alignment (DFMA) algorithm for UDA. Specifically, DFMA advances a three-stage Instance-Class-Pseudo (ICP) strategy consisting of instance-level contrastive learning, class-level contrastive learning and pseudo-labeling methods to mine the discriminative structure of target data. Then we conduct the cross-domain discriminative feature alignment by integrating the adversarial-based method which designs a novel conditional domain discriminator, and the discrepancy-based method which leverages the first-order and second-order statistical information of features. Furthermore, we build a reconstruction network to enhance the class-level feature alignment. Extensive experiments on several standard UDA benchmark datasets validate the superiority of our proposed DFMA. Jing Xiang, Guitao Cao, Xinyue Zhang 0010, Hanxiu Zhang, Chunwei Wu |
IJCNN | 2 |
| 2023 | Chaos to Order: A Label Propagation Perspective on Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA), where only a pre-trained source model is used to adapt to the target distribution, is a more general approach to achieving domain adaptation in the real world. However, it can be challenging to capture the inherent structure of the target features accurately due to the lack of supervised information on the target domain. By analyzing the clustering performance of the target features, we show that they still contain core features related to discriminative attributes but lack the collation of semantic information. Inspired by this insight, we present Chaos to Order (CtO), a novel approach for SFDA that strives to constrain semantic credibility and propagate label information among target subpopulations. CtO divides the target data into inner and outlier samples based on the adaptive threshold of the learning state, customizing the learning strategy to fit the data properties best. Specifically, inner samples are utilized for learning intra-class structure thanks to their relatively well-clustered properties. The low-density outlier samples are regularized by input consistency to achieve high accuracy with respect to the ground truth labels. In CtO, by employing different learning strategies to propagate the labels from the inner local to outlier instances, it clusters the global samples from chaos to order. We further adaptively regulate the neighborhood affinity of the inner samples to constrain the local semantic credibility. In theoretical and empirical analyses, we demonstrate that our algorithm not only propagates from inner to outlier but also prevents local clustering from forming spurious clusters. Empirical evidence demonstrates that CtO outperforms the state of the arts on three public benchmarks: Office-31, Office-Home, and VisDA. Chunwei Wu, Guitao Cao, Yan Li 0063, Xidong Xi, Wenming Cao 0001 |
ACM Multimedia | 2 |
| 2023 | Geometric Contrastive Learning for Heterogeneous Graphs EncodingabstractHeterogeneous graphs can represent many network structures in the real world, and research on heterogeneous graph data has attracted more attention. Most existing approaches require additional label information to obtain meaningful node representations. However, labeling heterogeneous graphs is tedious and time-consuming, and low-quality labels will harm the efficiency of models. In addition, the number of nodes increases exponentially with the distance from the root node, but the linearly expanding Euclidean space is difficult to match this growth rate. In this paper, we propose a unified framework that leverages the label irrelevance of contrastive learning and the unique expressive ability of hyperbolic space to encode heterogeneous graphs. Specifically, we use the contrast mechanism to obtain semantic information in an unsupervised way. Meanwhile, we design hyperbolic encoders that are more suitable for graph structure to learn the latent information from heterogeneous graphs efficiently. Experiments on four real-world heterogeneous graph data sets demonstrate the competitive efficacy of the proposed method. Siheng Wang, Guitao Cao, Chunwei Wu |
SMC | 2 |
| 2023 | OpenNet: Incremental Learning for Autonomous Driving Object Detection with Balanced LossabstractAutomated driving object detection has always been a challenging task in computer vision due to environmental uncertainties. These uncertainties include significant differences in object sizes and encountering the class unseen. It may result in poor performance when traditional object detection models are directly applied to automated driving detection. Because they usually presume fixed categories of common traffic participants, such as pedestrians and cars. Worsely, the huge class imbalance between common and novel classes further exacerbates performance degradation. To address the issues stated, we propose OpenNet to moderate the class imbalance with the Balanced Loss, which is based on Cross Entropy Loss. Besides, we adopt an inductive layer based on gradient reshaping to fast learn new classes with limited samples during incremental learning. To against catastrophic forgetting, we employ normalized feature distillation. By the way, we improve multi-scale detection robustness and unknown class recognition through FPN and energy-based detection, respectively. The Experimental results upon the CODA dataset show that the proposed method can obtain better performance than that of the existing methods. Zezhou Wang, Guitao Cao, Xidong Xi, Jiangtao Wang 0009 |
SMC | 2 |
| 2023 | A novel inference paradigm based on multi-view prototypes for one-shot semantic segmentation
Guitao Cao, Wenming Cao 0001 |
Appl. Intell. | 2 |
| 2022 | Attention 3D Fully Convolutional Neural Network for False Positive Reduction of Lung Nodule Detection
Guitao Cao, Beichen Zheng |
ICONIP (6) | 1 |
| 2022 | Mix-up Consistent Cross Representations for Data-Efficient Reinforcement LearningabstractDeep reinforcement learning (RL) has achieved re-markable performance in sequential decision-making problems. However, it is a challenge for deep RL methods to extract task-relevant semantic information when interacting with limited data from the environment. In this paper, we propose Mix-up Consistent Cross Representations (MCCR), a novel self-supervised auxiliary task, which aims to improve data efficiency and encourage representation prediction. Specifically, we calculate the contrastive loss between low-dimensional and high-dimensional representations of different state observations to boost the mutual information between states, thus improving data efficiency. Furthermore, we employ a mixed strategy to generate intermediate samples, increasing data diversity and the smoothness of representations prediction in nearby timesteps. Experimental results show that MCCR achieves competitive results over the state-of-the-art approaches for complex control tasks in DeepMind Control Suite, notably improving the ability of pretrained encoders to generalize to unseen tasks. Guitao Cao, Yan Li 0063, Chunwei Wu, Xidong Xi |
IJCNN | 2 |
| 2022 | Feature Pyramid Vision Transformer for MedMNIST Classification DecathlonabstractMedMNIST is a medical dataset proposed to block the need for medical knowledge, but there is currently no model that can generalize well on all its sub-datasets. Owing to the inadequacy of long-range relation modeling, models based on convolutional neural networks (CNNs) cannot fully learn the information of images. Besides, relying only on high-level features limits the generalization effect as well. All of these remain challenges for MedMNIST Classification Decathlon. In this paper, we proposed Feature Pyramid Vision Transformer (FPViT), a strong alternative for MedMNIST Classification Decathlon. Our FPViT exhibits enhanced feature learning and modeling capabilities, which merits both residual network (ResNet) and Vision Transformer (ViT). Transformers in our model take the features extracted by ResNet as sequences to capture global contexts which compensate for the lack of locality of convolution operations. Moreover, the feature pyramid designed in our model effectively utilizes the multi-scale feature maps from basic layers of ResNet. These multi-scale features from low-level to high level enable our model to have better adaptability. And, the final prediction is based on the multi-scale ViT and the original ResNet heads. Through experiments, our FPViT can achieve superior classification and generalization on MedMNIST than state-of-the-art methods. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
IJCNN | 3 |
| 2022 | Adversarial Discriminative Feature Separation for Generalization in Reinforcement LearningabstractImporving the generalization ability of an agent is an important and challenging task in deep reinforcement learning (RL). Procedually generated environment is an important benchmark for testing generalization in deep RL. In this benchmark, each game consists of multiple levels, each level is an algorithmically created environment instance with a unique configuration of its factors of variation. Existing methods (e.g., regularization, data augmentation) for improving the generalization of RL agent do not learn well the invariant representation among multiple levels. Besides, existing methods for learning invariant representations in RL using adversarial training can only learn invariant information across two levels. To solve this problem, we propose Adversarial Discriminative Feature Separate (ADFS). First, ADFS design a new discriminator for distinguishing whether two observations belong to the same level. Thus, the policy encoder is encouraged to learn invariant information between multiple levels. Second, it separates the representation of observation into level-invariant features and level-discriminative features, so that correction of the optimization direction of the discriminator. The discriminative features are learned by reducing the similarity of specific features intra-levels and increasing that of inter-levels, respectively. Experimental results demonstrate that our method is quite competitive with existing state-of-the-art methods on Procgen Benchmark. Chunwei Wu, Xidong Xi, Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
IJCNN | 5 |
| 2022 | PassAugment: Pass Nodes Importance in Graph Data Augmentation for Graph ClassificationabstractData augmentation has been widely introduced into graph-based tasks to improve the generalizability of models. Based on empirical hypothesis, we show that the node in a graph has different importance, the important one is critical for classification task while the unimportant one hurts the performance. However, there are few works in data augmentation addressing the information propagation of nodes with different importance. In this work, we propose a novel graph data augmentation algorithm for graph classification task, called PassAugment, aiming to pass these importance in graph data augmentation. After distinguishing the importance of all nodes in each graph using the saliency map, we design a data augmentation approach including two strategies: (i) randomly adding edges between the important nodes and the other nodes to globally improve the effective information passing, and (ii) randomly removing edges between the unimportant nodes and their neighbors to locally reduce the ineffective information passing. More importantly, our proposed approach as a standalone module can be combined with many GNNs architectures. Experimental results on graph classification task show that our approach consistently improves the accuracy and achieves or closely matches the state-of-the-art performance. Xiaohu Li, Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
SMC | 4 |
| 2022 | CBPGM: A Cache Based Piecewise Geometric Model IndexabstractRecent works on learned indexes have changed the way we look at the decades-old field of Database Management System indexing. However, they are limited to too many hyperparameters, long model construction time, and not taking full advantage of CPU cache and hardware acceleration. In this paper, we propose a Cache Based Piecewise Geometric Model (CBPGM) Index to address these issues with only one hyperparameter and effectively combines a sampling approach to reduce training dataset size that accelerates the construction procedure and aligns models and data to the CPU cache line to improve search performance. Experimental results show that the CBPGM index can improve the construction speed up to 8$\times$ and the query speed by 30% compared with the PGM index. Xiaopei Xu, Guitao Cao |
SMC | 2 |
| 2022 | Multi-strategy mutual learning network for deformable medical image registration
Wenming Cao 0001, Ye Duan, Guitao Cao, Deliang Lian |
Neurocomputing | 4 |
| 2022 | Implicit user relationships across sessions enhanced graph for session-based recommendationabstractSession-based recommendation aims to predict users’ next preference based on the sequence of their own history preferences in a short period. Most state-of-the-art methods model the session as a graph using graph neural networks (GNN) to capture the dynamic transitions between items within sessions. However, the complex and hidden correlations between different sessions are not adequately addressed, especially during the testing stage. We argue that session-based recommendation tasks can be improved by exploiting the correlations between different sessions in both training and testing. To this end, we propose a novel three-GNN-based recommendation framework to exploit the intra- and inter-session item correlations and the session-session correlations. The first one is a graph-based multi-layer perceptron to learn the inter-session item representations in a contrastive learning scheme guided by a contrastive loss function. The second one is a multi-relation graph attention network for intra-session item representations. The two item embeddings are combined through a position attention scheme to form the session representation, which is modulated and enhanced by the third extra session GNN by capturing the session-session correlations. The three levels of correlations are used in the training and testing stages in a joint prediction manner. To alleviate the data sparsity issue faced by the session GNN, we expand the session items by incorporating the neighboring items in the global item graph built from the entire training sessions. We have evaluated our method on multiple benchmark datasets. The results have shown that our joint recommendation method based on session correlation has significantly improved the recommendation accuracy over the state-of-the-art by more than 10%. Wenming Cao 0001, Yishan Liu, Guitao Cao, Zhiquan He |
Inf. Sci. | 3 |
| 2022 | Adaptive weight multi-channel center similar deep hashing
Xinghua Liu 0004, Guitao Cao, Qiubin Lin, Wenming Cao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Client Scheduling and Resource Management for Efficient Training in Heterogeneous IoT-Edge Federated LearningabstractFederated learning (FL) offers a promising paradigm that empowers numerous Internet of Things (IoT) devices to implement distributed learning on the premise of ensuring user privacy and data security. However, since FL adopts a synchronous distributed training mode, the heterogeneity of participating IoT devices and limited communication resources make FL encounter serious issues of low training efficiency in actual deployment. In this article, we propose an excellent FL policy for the heterogeneous IoT-edge FL system to improve distributed training efficiency. Specifically, first, by borrowing the idea of clustering, we explore an iterative self-organizing data analysis techniques algorithm (ISODATA)-based heterogeneous-aware client scheduling strategy to alleviate the issue of low training efficiency incurred by the heterogeneity of clients. Subsequently, to tackle the challenge of limited communication resources in FL, we first analyze the characteristics of the optimal resource block allocation solution theoretically and then introduce a mixed-integer linear programming (MILP)-based strategy to judiciously allocate resource blocks for scheduled clients. Comprehensive experimental results demonstrate that, compared with benchmarking strategies, our proposed FL policy can achieve up to 55.22% accuracy improvement in a relaxed time scenario, and attain up to$3.62\times $acceleration for reaching the specific expected accuracy. Yangguang Cui, Kun Cao 0001, Guitao Cao, Meikang Qiu, Tongquan Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2021 | DCFG: Discovering Directional CounterFactual Generation for Chest X-raysabstractWhile Deep Neural Networks (DNNs) are achieving state-of-the-art performance on medical domains across a variety of tasks, the need for explainability of model predictions in these high-stakes tasks is still lacking. Current for the explainability in model predictions potentially relies on the supervised counterfactual generation that is time-consuming and direction uncontrollable. Yet, the counterfactual generation needs to be easy to implement and have a controllable direction. In light of this trend, we propose an approach for the unsupervised latent direction search of black-box models that are steerable to the user by enabling the user to effectively explore counterfactual generation in a directional way, without relying on domain- or data-specific assumptions. To identify these explainable directions, we use Principal Component Analysis (PCA), a general manifold learning framework to extract low-dimensional subspaces based on a local noise injection of the pre-trained generative model, so that a small perturbation in the subspaces would provide enough change in the resulting data. With experiments on three real-world CXR datasets involving 6 tasks, we find that our approach is capable of learning explainable predictions that discard unrelated confounding factors. Moreover, our method enables practitioners to edit directions to better understand which features are used for predictions. Yan Li 0063, Chunwei Wu, Xidong Xi, Guitao Cao, Wenming Cao 0001 |
BIBM | 5 |
| 2021 | Shape-aware Multi-task Learning for Semi-supervised 3D Medical Image SegmentationabstractSemi-supervised learning has achieved many successes in medical image segmentation since it reduces the costs of manually annotating by leveraging abundant unlabeled data. However, these semi-supervised methods lack attention to ambiguous regions (e.g., some edges or corners around the targets), which may lead to meaningless and unreliable guidance. In this paper, we propose a novel semi-supervised segmentation method called Shape-aware Multi-task Learning (SMTL) to address the above issue. Our multi-task framework includes three tasks namely i) the main task for segmentation ii) one auxiliary task for signed distance regression iii) another auxiliary task for contour detection. The multi-task framework jointly predicts probabilistic segmentation maps, signed distance maps (SDMs) and edge maps to collect complementary information in the existing target label. Specifically, these two auxiliary tasks explicitly enforce shape-priors on the segmentation output to generate more accurate masks. Moreover, we design a region-attention-based adversarial learning strategy that enforces the consistency of two auxiliary tasks prediction distributions on the unlabeled and labeled data to make a meaningful and reliable guidance. We evaluate our SMTL on the datasets of the 2018 Atrial Segmentation Challenge and the 2017 Liver Tumor Segmentation Challenge. The results demonstrate that our SMTL achieves improvements and outperforms the state-of-the-art semi-supervised methods. Yan Li 0063, Xiaohu Li, Guitao Cao |
BIBM | 4 |
| 2021 | Sample Efficient Lung Segmentation Using Group Structured Conditional Variational Data ImputationabstractPatients infected with COVID-19 can lead to their Chest X-rays (CXRs) with opacifications rendered regions, which may produce incomplete lung segmentation in automated image analysis models. To tackle this issue, we propose a Group structured Conditional Variational data Imputation model to capture the missing data accurately with conditional distribution, where the high-dimensional probability distribution is narrow down to a small latent space to account for unobserved features. This work particularly arises in the fight against COVID-19 that effectively modeling a segmentation of plausible can be presented to a subsequent automated risk scoring and treatment. We train this model with limited CXRs data to demonstrate the abilities on the task of data imputation and proved to be effective though with relatively small datasets. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
ICME | 2 |
| 2021 | Generating Adversarial Examples by Distributed Upsampling
Shunkai Zhou, Yueling Zhang, Guitao Cao, Jiangtao Wang 0009 |
ICONIP (1) | 3 |
| 2021 | Modeling nodule growth via spatial transformation for follow-up prediction and diagnosisabstractLung cancer is the leading cause of cancer deaths worldwide with its mortality rate higher than that of other leading cancers. Nodules in the lungs can grow quite large without any obvious symptoms until the condition has reached a certain stage. Early detection and diagnosis of growing nodules can lay a good foundation for further treatment and potentially improve lung cancer survival rate. In this paper a unified framework is proposed for visual prediction and diagnosis of follow-up lung nodules. Future nodule growth is predicted by modeling the nodule growth between consecutive Computed Tomography (CT) scans via spatial transformation using convolutional network. Nodule classification is made based on the predicted nodule growth and previous diagnosis. Experiments are conducted on a longitudinal follow-up dataset of 615 LDCT scans of 153 lung nodules in early stages from 125 patients. Quantitative and qualitative results demonstrate the effectiveness of the proposed method. Jiyu Sheng, Yan Li 0063, Guitao Cao |
IJCNN | 3 |
| 2021 | Embedding Bottleneck Gated Recurrent Unit Network for Radar Signal RecognitionabstractRadar signal recognition plays an significant role in civil applications. Corresponding to two types of intentional modulation signal and unintentional fingerprint signal, radar signal recognition has two kinds of tasks—automatic modulation classification and radar emitter identification. In this paper, we propose a Embedding Bottleneck Gated Recurrent Unit (EBGRU) network that can handle these two tasks separately. The EBGRU consists of three main processing steps. Firstly, the normalized signal pulses are trained in pulse embedding network containing several embedding methods: Pulse2Vec, GloVeP and EPMo, during which we regard the radar signal pulses as radar signal-linguistic sequences for the first time. Then, pulses embeddings are added to original pulses and are sampled to form latent representations of pulses through information bottleneck. Finally, the gated recurrent unit network is utilized to predict radar signal labels. Experiment results show that the proposed method has reached 95.33% on simulated modulation signals and 94.67% at real intercepted emitter signals with relatively less network parameters. Yannan Wang, Guitao Cao, Danning Su |
IJCNN | 2 |
| 2021 | Debiased Prototype Network for Adversarial Domain AdaptationabstractDomain adaptation is an important and challenging task. Existing adversarial domain adaptation methods explore the relationship between the source and target domains, with the knowledge learned in the source domain supporting the target domain task. The quality of the knowledge will affect the task performance of the transfer, i.e., the higher the quality of the knowledge, the better the transfer task performance. To obtain better domain-invariant knowledge, we extract domain-invariant semantic information over the unit sphere via the prototype network. With the help of geometric constraints from the hypersphere, the features can be more tightly clustered with the estimated prototype (representatives of each class). Adaptation is achieved by adversarial learning to align the domain distribution, which enhances the transferability of the learned features and obtains the basic prototype. Since the basic prototypes dominantly computed from the source domain are biased against the expected domain-invariant prototype, a debiased method is further proposed to obtain the domain-invariant prototypes. Specifically, our method diminishes the intra- and inter- class bias to achieve the class-level alignment. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several domain adaptation benchmark datasets. Our code is available at https://github.com/Chunweiwu-source/DPN. Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
IJCNN | 2 |
| 2021 | IAE-ClusterGAN: A new Inverse autoencoder for Generative Adversarial Attention Clustering network
Chao Ling, Guitao Cao, Wenming Cao 0001 |
Neurocomputing | 2 |
| 2020 | A Dynamic Group Equivariant Convolutional Networks for Medical Image AnalysisabstractGroup equivariant Convolutional Neural Networks (G-CNNs) has led to big empirical success in the medical domain, one fundamental assumption is that equivariance provides a powerful inductive bias for medical images. By leveraging concepts from group representation theory, we can generalize vanilla Convolutional Neural Networks (CNNs) to G-CNN. Currently, although embedding an arbitrary equivariance to CNNs can learn powerful disentangled representations in a higher dimensional domain, they lack explicit means to learn meaningful relationships among the equivariant convolutional kernels. In this paper, we propose a generalization of the dynamic convolutional method, named as dynamic group equivariant convolution, to strengthen the relationships and increase model capability by aggregating multiple group convolutional kernels via attention. Meanwhile, we generalize attention to an equivariant one to preserve equivariant of dynamic group convolution. In our approach, this leads to a flexible framework that enables a dynamic convolutional in G-CNNs by means of a dynamic routing layer expansions. We demonstrate that breast tumor classification is substantial improvements when compared to a recent baseline architecture. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
BIBM | 2 |
| 2019 | Emotion Recognition from Children Speech Signals Using Attention Based Time Series Deep LearningabstractChildren's emotions expression concentrates in the acoustic aspects such as the tones and timbres of the voice instead of the semantics, and there are a lot of lengthy fragments in their speech. This paper proposes an emotion recognition model using the time series deep learning technology, named attention based Bi-directional Long Short-Term Memory (CNN-BiLSTM) to extract the emotional features. After preprocessing the speech signal, the forty-dimensional Mel Frequency Cepstral Coefficients (MFCC) related parameters are extracted, including the dynamic and static features. And these frequency domain features are enhanced by convolutional neural networks (CNNs) as the emotional features of children's speech recognition. BiLSTM is used to solve the problem of poor performance of long-term dependent learning features, and attention mechanism is used for only a few frames contain emotional features in the children speech signal. Compared with the related speech emotion recognition models such as LSTM-CNN and 2D-CNN-LSTM, our proposed speech emotion recognition model improves the accuracy up to 71.6% on the FAU-AIBO children's speech emotion database. Guitao Cao, Yunming Tang, Jiyu Sheng, Wenming Cao 0001 |
BIBM | 1 |
| 2019 | Automatic Epileptic Seizure Detection via Attention-Based CNN-BiRNNabstractEpileptic seizure detection with multi-channel electroencephalography (EEG) signals is a commonly used method, but it is tedious and error-prone to manually detect seizures through EEG signals. In this work, we propose an end-to-end deep neural network called attention-based CNN-BiRNN for automatic seizure detection. Attention-based CNN-BiRNN mainly consists of three parts: the multi-scale convolution model, the attention model, and the multi-stream bidirectional recurrent model. Original signals are firstly sent to the multi-scale convolution model to extract multi-scale features. Then the attention model exploits the differences among channels for seizure detection. Afterwards, the robust temporal features are obtained by the multi-stream bidirectional recurrent model, and are further fed into a fully connected layer for classification. Moreover, a channel dropout method is proposed, for the model training stage, to obtain inconspicuous characteristics from all the channels of a certain EEG signal. The results on the dataset of CHB-MIT demonstrate that our approach outperforms state-of-the-art approaches in terms of both sensitivity and specificity. Furthermore, with the channel dropout method, our approach is shown to have a powerful ability of handling EEG signals with missing channels and different channels. Chengbin Huang, Weiting Chen, Guitao Cao |
BIBM | 3 |
| 2019 | Stacking-based deep neural network for Facial Expression RecognitionabstractWe present a scalable stacking-based deep neural network(S-DNN) for facial expression recognition. The network is a congregate of basic learning models in series to synthesize a deep neural network with feedforward network architecture. Thur, choosing trainable learning modules is the core to effectively build S-DNN in an end-to-end manner. Inspired by the manifold learning archetype, we implement a Patch Discriminative Analysis(PDA) as a basic learning model, followed by hashing and block histogram on the top, which sample image in a low discriminative space, and finding an efficient representation of the training data. As those self-learnable models trained, a low dimensional discriminative feature is implicitly learned, which proves to be useful in facial expression recognition. Experimental results on the facial expression dataset(CK+) show that the proposed model is superior to its counterparts, capable of achieving state-of-the-art performance. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
BIBM | 2 |
| 2019 | Density Map Estimation for Crowded Chicken
Tianze Rong, Guitao Cao |
ICIG (3) | 3 |
| 2018 | 3D Convolutional Neural Networks Fusion Model for Lung Nodule Detection onClinical CT Scans
Guitao Cao, Tiantian Huang, Wenming Cao 0001 |
BIBM | 1 |
| 2018 | Adaptive Discrete Hypergraph MatchingabstractThis paper addresses the problem of hypergraph matching using higher-order affinity information. We propose a solver that iteratively updates the solution in the discrete domain by linear assignment approximation. The proposed method is guaranteed to converge to a stationary discrete solution and avoids the annealing procedure and ad-hoc post binarization step that are required in several previous methods. Specifically, we start with a simple iterative discrete gradient assignment solver. This solver can be trapped in an -circle sequence under moderate conditions, where is the order of the graph matching problem. We then devise an adaptive relaxation mechanism to jump out this degenerating case and show that the resulting new path will converge to a fixed solution in the discrete domain. The proposed method is tested on both synthetic and real-world benchmarks. The experimental results corroborate the efficacy of our method. Junchi Yan, Yin Li 0003, Guitao Cao |
IEEE Trans. Cybern. | 4 |
| 2017 | Action recognition based on depth image sequenceabstractHuman action recognition is the process of labeling image sequences with action labels. Robust solutions to this problem have applications in domains such as medical care, human-computer interaction and virtual training. The task is challenging for feature extraction due to variations in motion performance, recording settings and inter-personal differences. To meet these challenges, we propose two types of feature extraction methods based on the Kinect depth image sequences in this paper. One is assuming that there exists even distribute position lines in the three-dimensional space of frame difference, it will be active when the moving object touches them. The other is mapping the 16 successive frame sequences to a single image by Speed Time Mapping (STM) or Time Depth Mapping (STDM), obtaining 36-dimensiona spatial-temporal features in this image. These features are fed into Support Vector Machine (SVM) to identify the action categories. The experiments compare their performance and demonstrate the effectiveness of STDM. Liangcan Liao, Guitao Cao, Wenming Cao 0001 |
BIBM | 2 |
| 2016 | Classification of tongue images based on doublet and color space dictionaryabstractRecently, pathological diagnosis plays a crucial role in many areas of medicine, and some researchers have proposed many models and algorithms for improving classification accuracy by extracting excellent feature or modifying the classifier. They have also achieved excellent results on pathological diagnosis using tongue images. However, pixel values can't express intuitive features of tongue images and different classifiers for training samples have different adaptability. Accordingly, this paper presents a robust approach to infer the pathological characteristics by observing tongue images. Our proposed method makes full use of the local information and similarity of tongue images. Firstly, tongue images in RGB color space are converted to Lab. Then, we compute tongue statistics information. In the calculation process, Lab space dictionary is created at first, through it, we compute statistic value for each dictionary value. After that, a method based on Doublets is taken for feature optimization. At last, we use XGBOOST classifier to predict the categories of tongue images. We achieve classification accuracy of 95.39% using statistics feature and the improved classifier, which is helpful for TCM (Traditional Chinese Medicine) diagnosis. Guitao Cao, Ye Duan, Liping Tu, Jiatuo Xu, Dong Xu 0002 |
BIBM | 1 |
| 2016 | Prediction of neonatal amplitude-integrated EEG based on LSTM methodabstractAmplitude-integrated EEG (aEEG) is becoming more and more useful in the monitoring of clinically ill neonates. If there is a method that can predict neonatal aEEG signals, doctors can forecast the possible abnormality of neonates' brain functions in advance and give early intervention. However, no such research on the prediction of aEEG signals has been found in the literature. In this paper, we combine aEEG signals with Long-Short Time Memory (LSTM) model and propose a method to predict aEEG signals based on LSTM. All of the aEEG signals after preprocessing were used as the input of the LSTM, a type of recurrent neural networks which can process long term signals with high accuracy. To assess the method, several experiments were conducted on 276 neonatal aEEG tracings including 217 normal cases and 59 abnormal ones. Experimental results show that the predicted aEEG signals are very close to the real aEEG signals. Our LSTM-based method might therefore help predict neonatal brain disorders in NICUs. Lizhe Liu, Weiting Chen, Guitao Cao |
BIBM | 3 |
| 2016 | A deep tongue image features analysis model for medical applicationabstractWith the improvement of people's living standards, there is no doubt that people are paying more and more attention to their health. However, shortage of medical resources is a critical global problem. As a result, an intelligent prognostics system has a great potential to play important roles in computer aided diagnosis. Numerous papers reported that tongue features have been closely related to a human's state. Among them, the majority of the existing tongue image analyses and classification methods are based on the low-level features, which may not provide a holistic view of the tongue. Inspired by a deep convolutional neural network (CNN), we propose a deep tongue image feature analysis system to extract unbiased features and reduce human labor for tongue diagnosis. With the unbalanced sample distribution, it is hard to form a balanced classification model based on feature representations obtained by existing low-level and high-level methods. Our proposed deep tongue image feature analysis model learns high-level features and provide more classification information during training time, which may result in higher accuracy when predicting testing samples. We tested the proposed system on a set of 267 gastritis patients, and a control group of 48 healthy volunteers (labeled according to Western medical practices). Test results show that the proposed deep tongue image feature analysis model can classify a given tongue image into healthy and diseased state with an average accuracy of 91.49%, which demonstrates the relationship between human body's state and its deep tongue image features. Guitao Cao, Ye Duan, Minghua Zhu, Liping Tu, Jiatuo Xu, Dong Xu 0002 |
BIBM | 2 |
| 2016 | Facial expression recognition based on LLENetabstractFacial expression recognition plays an important role in lie detection, and computer-aided diagnosis. Many deep learning facial expression feature extraction methods have a great improvement in recognition accuracy and robutness than traditional feature extraction methods. However, most of current deep learning methods need special parameter tuning and ad hoc fine-tuning tricks. This paper proposes a novel feature extraction model called Locally Linear Embedding Network (LLENet) for facial expression recognition. The proposed LLENet first reconstructs image sets for the cropped images. Unlike previous deep convolutional neural networks that initialized convolutional kernels randomly, we learn multi-stage kernels from reconstructed image sets directly in a supervised way. Also, we create an improved LLE to select kernels, from which we can obtain the most representative feature maps. Furthermore, to better measure the contribution of these kernels, a new distance based on kernel Euclidean is proposed. After the procedure of multi-scale feature analysis, feature representations are finally sent into a linear classifier. Experimental results on facial expression datasets (CK+) show that the proposed model can capture most representative features and thus improves previous results. Dan Meng 0001, Guitao Cao, Zhihai He, Wenming Cao 0001 |
BIBM | 2 |
| 2016 | Automated human physical function measurement using constrained high dispersal network with SVM-linearabstractPhysical measurement have been becoming increasingly helpful in monitoring the humans health status. Manual measurement of physical status is time consuming and may result in misdiagnosing, so an automatic method for identification the status of physical is urgently needed. This paper presents a novel feature extraction method based on using constrained high dispersal network for depth images and coped with Support Vector Machines (SVM) to measure human physical function. The proposed method can catch the most representative features of depth images belonging to different actions and statuses. We analyze the representation efficiency of hand-crafted features (HOG features, and LBP features), deep learning features (CNN features, and PCANet features) and our proposed deep learning features separately in order to validate the efficiency and accuracy of our proposed method. The results show superior performance of 85.19% on 3840 samples (three actions, each with four different statuses, and every status contains sixteen sequences) when the proposed deep features combined with SVM. Dan Meng 0001, Guitao Cao, Weiting Chen, Wenming Cao 0001 |
BIBM | 2 |
| 2016 | Automatic fall detection of human in video using combination of featuresabstractThe problem of automatically fall detection of older people living alone is a popular research topic since falls are one of the major health hazards among the aging population aged 65 and above and the population of them in China is more than 100 million. In this paper, we present an automatic human fall detection framework based on video surveillance which can improve safety of elders in indoor environments. First, a vision component was used to detect and extract moving people in videos from static cameras. Then, we combine Histograms of Oriented Gradients(HOG),Local Binary Pattern(LBP)and feature extracted by the Deep Learning Framework Caffe to form a new augmented feature and the feature is named HLC. We use HLC to represent a person's motion state in a frame of a video sequence. Because the process of fall is a sequence of movements, we use HLC features which were extracted from continuous frames of a video sequence to implement the fall detection. With the help of the HLC feature, we achieve an average fall detection result of 93.7% sensitivity and 92.0% specificity on three different datasets. Guitao Cao, Dan Meng 0001, Weiting Chen, Wenming Cao 0001 |
BIBM | 2 |
| 2016 | Task-Driven Progressive Part Localization for Fine-Grained Object RecognitionabstractThe problem of fine-grained object recognition is very challenging due to the subtle visual differences between different object categories. In this paper, we propose a task-driven progressive part localization (TPPL) approach for fine-grained object recognition. Most existing methods follow a two-step approach that first detects salient object parts to suppress the interference from background scenes and then classifies objects based on features extracted from these regions. The part detector and object classifier are often independently designed and trained. In this paper, our major finding is that the part detector should be jointly designed and progressively refined with the object classifier so that the detected regions can provide the most distinctive features for final object recognition. Specifically, we develop a part-based SPP-net (Part-SPP) as our baseline part detector. We then establish a TPPL framework, which takes the predicted boxes of Part-SPP as an initial guess, and then examines new regions in the neighborhood using a particle swarm optimization approach, searching for more discriminative image regions to maximize the objective function and the recognition performance. This procedure is performed in an iterative manner to progressively improve the joint part detection and object classification performance. Experimental results on the Caltech-UCSD-200-2011 dataset demonstrate that our method outperforms state-of-the-art fine-grained categorization methods both in part localization and classification, even without requiring a bounding box during testing. Zhihai He, Guitao Cao, Wenming Cao 0001 |
IEEE Trans. Multim. | 3 |
| 2016 | Animal Detection From Highly Cluttered Natural Scenes Using Spatiotemporal Object Region Proposals and Patch VerificationabstractIn this paper, we consider the animal object detection and segmentation from wildlife monitoring videos captured by motion-triggered cameras, called camera-traps. For these types of videos, existing approaches often suffer from low detection rates due to low contrast between the foreground animals and the cluttered background, as well as high false positive rates due to the dynamic background. To address this issue, we first develop a new approach to generate animal object region proposals using multilevel graph cut in the spatiotemporal domain. We then develop a cross-frame temporal patch verification method to determine if these region proposals are true animals or background patches. We construct an efficient feature description for animal detection using joint deep learning and histogram of oriented gradient features encoded with Fisher vectors. Our extensive experimental results and performance comparisons over a diverse set of challenging camera-trap data demonstrate that the proposed spatiotemporal object proposal and patch verification framework outperforms the state-of-the-art methods, including the recent Faster-RCNN method, on animal object detection accuracy by up to 4.5%. Zhi Zhang 0005, Zhihai He, Guitao Cao, Wenming Cao 0001 |
IEEE Trans. Multim. | 3 |
| 2011 | Data-driven based 3-D fuzzy logic controller design using nearest neighborhood clustering and linear support vector regressionabstractThree-dimensional fuzzy logic controller (3-D FLC) is a novel FLC developed for spatially distributed parameter systems. In this study, we are concerned with data-based 3-D FLC design. A nearest neighborhood clustering algorithm is employed to extract fuzzy rules from input-output data pairs, and then an optimization algorithm based on geometric similarity measure is used to reduce the obtained rule base. The consequent parameters are estimated using linear support vector regression. Finally, a catalytic packed-bed reactor is taken as an application to demonstrate the effectiveness of the 3-D FLC. Xianxia Zhang, Chenkun Qi, Guitao Cao |
FUZZ-IEEE | 5 |