VLDB 2026 Research / reviewers in the wild / expert
Wenming Cao 0001
dblp:90/3695-1 · also Wen-Ming Cao 0001, Wen-ming Cao 0001
· DBLP profile ↗
123ranked-venue papers
21as first author
94since 2021 · last 2026
0000-0002-8174-6167ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 11 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 4 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Computer networks · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical joint contrastive learning with knowledge distillation for self-supervised 3D skeleton-based action recognition
Wenming Cao 0001, Xinpeng Yin |
Expert Syst. Appl. | 1 |
| 2026 | Beyond single-view analysis: An interaction-aware multi-Period encoder network for multi-person motion prediction
Xinkai Chen, Wenming Cao 0001, Jianqi Zhong |
Expert Syst. Appl. | 2 |
| 2026 | Super-joint guided local relation groups with fuzzy mask for self-supervised skeleton action recognition
Songling Liu, Xinpeng Yin, Wenming Cao 0001, Jianhua Ji |
Neurocomputing | 3 |
| 2026 | Fine-grained behavior interaction-aware network for efficient multi-person motion forecasting
Wenming Cao 0001, Jianqi Zhong |
Multim. Syst. | 1 |
| 2026 | Bias-driven prediction update network for long-term 3D human motion prediction
Jianqi Zhong, Junyuan Tang, Tianming Xie, Wenming Cao 0001 |
Multim. Syst. | 6 |
| 2026 | Dual-decoder collaborative learning with multi-hybrid view augmentation for self-supervised 3D action recognition
Wenming Cao 0001, Yingfei Wu, Xinpeng Yin |
Pattern Recognit. | 1 |
| 2026 | Dynamic prototype with discriminative representation for rapid adaptation in new organ segmentation
Xinyue Zhang 0010, Guitao Cao, Wenming Cao 0001 |
Pattern Recognit. | 5 |
| 2026 | Geometric algebra enhanced twin-branch contrastive network design for few-shot hyperspectral image classification
Rujun Zhang, Wenming Cao 0001, Qifan Liu |
Pattern Recognit. | 2 |
| 2026 | AsymFormer: Asymmetric Interaction-Joints Dynamics Transformer for Multi-Person 3D Pose ForecastingabstractMulti-person pose forecasting has garnered increasing attention owing to its critical role in applications such as human-robotic interaction and autonomous systems. Existing methods have achieved notable progress by integrating interaction-aware mechanisms to capture relational dynamics among individuals. However, these approaches often suffer from two key limitations: (1) the implicit assumption that all individuals within a scene are socially relevant, and (2) the uniform treatment of all body joints, regardless of their activity levels or importance in motion representation. To address these challenges, we propose a novel Transformer-based architecture, termed AsymFormer, which is designed to asymmetrically model both human-level interactions and joint-level dynamics. At the human level, we introduce a Selective Interaction-Aware Module (SIAM), which leverages a discriminator pre-trained on ground-truth interaction labels to evaluate the social relevance of each individual based on their historical motion trajectories. This design enables the model to explicitly focus on meaningful interactions while filtering out socially irrelevant agents. At the joint level, we develop a Targeted Joints Dynamic Encoder (TJDE) to emphasize high-activity joints, thereby avoiding the dilution of informative motion signals caused by the equal treatment of all joints. This selective focus enhances the discriminative power of motion representations. Extensive experiments conducted on multiple benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches in multi-person 3D pose forecasting. Wenming Cao 0001, Xinkai Chen, Jianqi Zhong |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2026 | Constant-Invariant Information Guided Augmented Spatiotemporal Networks for Skeleton-Based Action RecognitionabstractContrastive learning has achieved remarkable success in unsupervised skeleton-based action recognition. For contrast learning, multiple contrast pairs have been found to enhance the learning quality of the model. However, existing methods often only mine invariant information within corresponding contrastive pairs, ignoring the constant-invariant information across different contrastive pairs. To address this limitation, this paper introduces a novel Constant-invariant Information Guided Augmented Spatiotemporal Networks (CI-AS), composed of constant-invariant contrastive learning and an augmented spatiotemporal network. The constant-invariant contrastive learning leverages the feature differences among multiple positive pairs, prompting the model to learn more constant-invariant information. It also significantly increases the number of negative samples without adding extra parameters. The augmented spatiotemporal network integrates two multi-contrastive strategies: different domains and varying augmentation levels, thus maximizing the potential of constant-invariant contrastive learning. Additionally, a domain augmentation module is introduced to strengthen the interaction between features of different augmentation strengths, enabling the model to incorporate information from diverse domains and augmentation levels during comparison. Experiments on three large-scale skeleton datasets demonstrate that our method outperforms several state-of-the-art approaches in multiple benchmarks. Our code is available at https://github.com/aoyuleo/CI-AS. Wenming Cao 0001, Aoyu Zhang, Xinpeng Yin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | D2Vformer: A Flexible Time-Series Prediction Model Based on Time-Position EmbeddingabstractExisting time-series forecasting methods often struggle to adapt to dynamic scenarios and lack flexibility in prediction. They typically require retraining the model when the prediction length or position changes. Moreover, these methods still face challenges in effectively capturing and utilizing time-position embeddings (PEs). To address these limitations, this article proposes a novel model called D2Vformer. Unlike conventional prediction methods that rely on fixed-length predictors, D2Vformer can directly handle scenarios with arbitrary prediction lengths. In addition, it significantly reduces training resource consumption and proves highly effective in real-world dynamic environments. In D2Vformer, the Date2Vec (D2V) module is devised to leverage timestamp information and feature sequences to generate time PEs. Subsequently, D2Vformer introduces an innovative fusion module that leverages an attention mechanism to capture the mapping between input and target time PEs, thereby enabling flexible prediction. Extensive experiments on six datasets demonstrate that D2V outperforms other time-PE methods, while D2Vformer surpasses state-of-the-art approaches in both fixed-length and arbitrary-length prediction tasks. The code for D2Vformer is available at: https://github.com/TeamofHaoWang/D2Vformer. Xiaobao Song, Hao Wang 0075, Liwei Deng 0004, Hongbo Qiu, Wenming Cao 0001, Andrew Chi-Sing Leung |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Vision transformer-based generalized zero-shot learning with data criticizing
Yucuan Liang, Zhenqi Zhang, Wenming Cao 0001 |
Appl. Intell. | 4 |
| 2025 | Linearformer: Tri-Net Multi-Layer DVF Medical Image RegistrationabstractABSTRACT In medical imaging, accurate registration is crucial for reliable analysis. While transformer models demonstrate potential, their application to large datasets like OASIS is constrained by substantial memory requirements, quadratic complexity and the challenge of managing complex deformations. To overcome these challenges, Linearformer is introduced, an efficient transformer‐based model with Linear‐ProbSparse self‐attention for optimised time and memory, along with TNM DVF, a Pyramid‐based framework for unsupervised non‐rigid registration. Evaluated on OASIS and LPBA40 brain MRI datasets, the model outperforms state‐of‐the‐art methods in Dice score and Jacobian metrics, surpassing TransMatch by 0.6% and 1.9% on the two datasets while maintaining a comparable voxel folding percentage. Zhiyue Yan, Wenming Cao 0001 |
Expert Syst. J. Knowl. Eng. | 3 |
| 2025 | Enhancing 3D medical image registration with cross attention, residual skips, and cascade attentionabstractAt the core of Deep Learning-based Deformable Medical Image Registration (DMIR) lies a strong foundation. Essentially, this network compares features in two images to identify their mutual correspondence, which is necessary for precise image registration. In this paper, we use three novel techniques to increase the registration process and enhance the alignment accuracy between medical images. First, we propose cross attention over multi-layers of pairs of images, allowing us to take out the correspondences between them at different levels and improve registration accuracy. Second, we introduce a skip connection with residual blocks between the encoder and decoder, helping information flow and enhancing overall performance. Third, we propose the utilization of cascade attention with residual block skip connections, which enhances information flow and empowers feature representation. Experimental results on the OASIS data set and the LPBA40 data set show the effectiveness and superiority of our proposed mechanism. These novelties contribute to the enhancement of 3D DMIR-based on unsupervised learning with potential implications in clinical practice and research. Zhiquan He, Wenming Cao 0001 |
Intell. Data Anal. | 3 |
| 2025 | Session-based recommendation with quaternion-enhanced attention calculationabstractA sound session-based Recommendation system is an important thing for users and enterprises. Session data from anonymous users is used to infer subsequent behavior and make high-quality recommendations. Graph neural networks have been extensively used to represent and learn information about the graph structure of session data in many existing studies, which have achieved significant progress. However, certain limitations still exist. This part of the model mainly focuses on how to enhance the expressiveness of the model while ignoring the quality of embedding. And previous studies have shown that complex deep learning models do not always have an edge over relatively simple algorithms in terms of prediction accuracy. To complement this aspect, this paper proposes a model for session recommendation based on quaternion-enhanced attention computation. Distinct from Euclidean space, quaternions perform calculations and analogy deductions in a hyper-complex vector space. The Hamilton product employed by quaternions provides a highly significant and meaningful computational method for enhancing session representation and reducing model parameters. A quaternion weight calculation fusion mechanism was designed, which uses quaternion calculations in key calculation steps, such as calculating relevant learnable weights for item representation and session representation. Experimental results on the benchmark database show that QEAC-SR outperforms some of the existing state-of-the-art methods, with Precision scores improving over the best-performing baseline method by 1.70 % to 1.80 % on Diginetica and 21.51 % to 23.49 % on Tmall; MRR scores improve over the best-performing baseline method by 2.31 % to 2.36 % on Diginetica and 19.56 % to 19.65 % on Tmall. Yuanbiao Guo, Yishan Liu, Canta Zheng, Wenming Cao 0001 |
Intell. Data Anal. | 4 |
| 2025 | RDNet: Region specific iterative deformation with multi-scale attention for medical image registration
Wenming Cao 0001, Naeem Hussain, Zhiyue Yan |
Neurocomputing | 1 |
| 2025 | Potential Knowledge Extraction Network for Class-Incremental Learning
Xidong Xi, Guitao Cao, Wenming Cao 0001, Yan Li 0063 |
Neurocomputing | 3 |
| 2025 | H-SGANet: Hybrid sparse graph attention network for deformable medical image registration
Wenming Cao 0001 |
Neurocomputing | 2 |
| 2025 | STHRA: selective transformer hierarchical reciprocal attention-based deformable medical image registration
Zhiyue Yan, Wenming Cao 0001, Naeem Hussain |
Multim. Syst. | 3 |
| 2025 | Progressively deeper attention networks for 3D human motion prediction
Jiangtao Huang, Wenming Cao 0001, Jianqi Zhong |
Multim. Syst. | 3 |
| 2025 | Optimizing human motion prediction through decoupled motion spatio-temporal trends
Huan Pan, Ruiya Ji, Wenming Cao 0001, Jianqi Zhong |
Multim. Syst. | 3 |
| 2025 | Uncertainty guided semi-supervised few-shot segmentation with prototype level fusion
Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
Neural Networks | 5 |
| 2025 | NLA-GNN: Non-local information aggregated graph neural network for heterogeneous graph embedding
Siheng Wang, Guitao Cao, Wenming Cao 0001, Yan Li 0063 |
Pattern Recognit. | 3 |
| 2025 | Asymmetric Context-Guided Adaptive Alignment Network for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, self-supervised pre-training paradigms have been extensively investigated. Particularly, masked autoencoders-like methods based on masked target reconstruction have pushed the performance of pre-training to a new height, which are committed to choose a better target for reconstruction. In this work, we propose an asymmetric context-guided adaptive alignment network (ACA2Net) for self-supervised skeleton-based action recognition by utilizing a transformer-based teacher encoder guiding the student encoder to learn richer action contextual information. To tackle the misalignment from the asymmetry, we devise an adaptive alignment module to better align the student representations to the teacher’s. Additionally, considering that the differential operation for temporal motion might cause the prior loss related to the changes of direction, we propose a motion compass-aware masking strategy with fusion prior supplemented by motion and direction intensity. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets demonstrate that our proposed ACA2Net outperforms previous MAE-like methods. Wenming Cao 0001, Liangxi Qian, Yicha Zhang, Xuelong Li 0001, Xinpeng Yin |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Progressive Feature Reconstruction Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to transfer the knowledge learned in the seen classes to the unseen classes through semantic knowledge. However, to ensure the model’s versatility on different datasets, existing methods divide the image into blocks of the same size, resulting in the loss of information between attributes. More importantly, existing methods ignore that not every image contains all attributes corresponding to that class. In this paper, we propose a progressive feature reconstruction network, called PFRN. PFRN consists of an attribute relation sub-net and an attention-based feature reconstruction sub-net. Specifically, the attribute relation sub-net first adopts the attribute-related region module to obtain the attribute-related regions in the visual features, which are input to the attribute relation discovery module to find the relationships between attributes. The attention-based feature reconstruction sub-net obtains the fine-grained features based on attributes by the attribute attention module and uses the feature reconstruction module to randomly lose some attributes to reconstruct the new visual features of the missing attributes. The new visual features are fed back into the network for training. Finally, the attribute information learned by the attribute relation sub-net is fused to the visual embedding learned by the attention-based features reconstruction sub-net, and the ideal visual semantic interaction is performed with the semantic vector classified by ZSL. Extensive experiments on three ZSL benchmark datasets demonstrate the significant generalization performance of our proposed method over the state-of-the-art methods. Linchun Hu, Wenming Cao 0001, Zhenqi Zhang, Yuchuang Liang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Adaptive Meta-Path Selection Based Heterogeneous Spatial Enhancement for circRNA-Disease Associations PredictionabstractCircular RNAs (circRNAs) play a crucial role in human biological processes as miRNA sponges, regulating gene expression and affecting disease manifestations. Establishing heterogeneous nodal feature relationships through meta-paths can effectively enhance the predictive capability of models. Previous research primarily constructed associations between meta-paths manually, and excessive noise made it difficult to capture highly correlated hidden features. Equally important is learning more about feature distributions, which is key to improving the generalization ability of algorithms. To address these challenges, we propose an adaptive meta-path selection method named AdaMH. The core scheme introduces an adaptive path selection method that automatically identifies highly relevant heterogeneous meta-paths during iterative training rounds. Considering the sparsity of data distribution, we introduce controlled random noise into the data through graph contrastive learning to ensure an even distribution of features. Subsequently, a multi-head attention mechanism is utilized to capture relationships in the high-dimensional heterogeneous feature space, enhancing feature representation capability. Comparing with state-of-the-art (SOTA) algorithms, AdaMH is the only one that surpasses a performance threshold of 0.95 across seven evaluation metrics. Guitao Cao, Wenming Cao 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Robust Tensor Ring Decomposition for Urban Traffic Data ImputationabstractIn urban transportation systems, missing data and noise contamination are almost inevitable. To address the challenge of imputing traffic data corrupted by noise and outliers in real-world scenarios, this paper proposes a novel algorithm based on spatiotemporal tensor completion. The proposed method transforms observed data into three-dimensional spatiotemporal tensors and utilizes tensor ring decomposition for data completion. Furthermore, spatial and temporal information is incorporated into the model by utilizing the graph Laplacian matrix. To handle outliers, they are treated as unknown parameters, and the ℓ0-norm is introduced to ensure their sparsity, thereby achieving the Spatio-Temporal Tensor Completion model with ℓ0-norm term (STTC-ℓ0). The solution to the model is derived using the alternating optimization framework with the alternating direction method of multipliers. Then, we discuss the convergence of the solution method. To further enhance the efficiency of our proposed method, we combine the unrolling algorithm with our iterative optimization model, creating a lightweight and efficient neural network tailored for tensor completion, called STTC-ℓ0-NN. Extensive experiments conducted on real datasets demonstrate the superiority of our proposed method over several state-of-the-art methods across various experimental scenarios. It is worth noting that STTC-ℓ0-NN reduces computational time by one to two orders of magnitude compared to existing methods while maintaining or even improving imputation accuracy. The code is available at https://github.com/TCCofWANG/STTC-L0-and-STTC-L0-NN. Linfang Yu, Chenyu Guan, Hao Wang 0075, Wenming Cao 0001, Andrew Chi-Sing Leung |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | SAT-Net: Structure-Aware Transformer-Based Attention Fusion Network for Low-Quality Retinal FunduImages EnhancementabstractIn ophthalmology diagnosis, high-fidelity fundus images are essential for disease diagnosis and intervention. However, many real-world clinical conditions may degrade the quality of the acquired images and thus affect clinical diagnostic accuracy. Traditional convolutional neural network-based retinal fundus image enhancement methods cannot always capture long-range dependencies, which reduces the overall visual quality of images, especially for real retinal fundus images. Furthermore, existing enhancement methods often fail to fully utilize low-resolution structural detail information, which potentially leads to inaccurate pivotal fundus vessel topology or capillary details. In this paper, we propose a novel Structure-Aware Transformer-based attention fusion Network (SAT-Net) for low-quality retinal fundus image enhancement. First, we introduce a Transformer-based attention fusion module which incorporates window-based self-attention and channel self-attention to capture global spatial dependencies and emphasize important feature channels simultaneously. This fusion significantly improves the overall perceptual quality of the image by enhancing both the local details and the uniformity of the non-vessel background regions. Second, we introduce a cross-quality knowledge distillation technique, which bridges the quality gap between high-quality and low-quality fundus images. By designing a high-performing teacher network to guide a lightweight student network, the student network enables to capture detailed features from low-quality fundus images, further preserving critical diagnostic information and fine topology structures. Moreover, we design a structure-aware multi-scale loss function by using a trainable subnetwork to obtain the edge structure from different scales to better constrain pivotal fundus vessel structure and capillary details. Comprehensive quantitative and qualitative experiments on both synthetic and real fundus image datasets robustly validate that our proposed SAT-Net outperforms other state-of-the-art methods for fundus image enhancement. In addition, extensive comparative experiments on both the vessel segmentation and Optic Disc/Cup detection tasks further validate the effectiveness and superiority of our proposed method. Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | A Lightweight Depthwise Separable ConvNet with Frequency-domain Enhancement for Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is crucial in the diagnosis and treatment of various cardiovascular and eye diseases. Although current vessel segmentation methods have achieved impressive performance, some challenging issues still need to be addressed. For example, existing methods always cannot segment complex capillaries well because they may be interfered with or covered by other components in the retina, and they need to further improve the continuity and consistency of vessel segmentation results. Moreover, the excellent vessel segmentation methods are usually built on bulky and cumbersome models which greatly limit their application range. In this article, we propose a novel efficient depthwise separable convolution network with frequency-domain enhancement (dubbed RetiNeXt) for retinal vessel segmentation. Firstly, we design a lightweight vessel enhancement module to extract global fine topological structure features from the frequency domain to enhance the complex capillary vessel details. Secondly, we propose a global feature extraction block to fully capture the large-scale spatial information and global characterizations, which enables the model to maintain vessel structural coherence from a global perspective. Thirdly, we construct a local feature mixing block based on SimAM attention mechanism to highlight the tiny capillary topological structure features and optimize the segmentation of low-contrast blood vessels, thereby improving the integrity and continuity of complex capillaries. Comprehensive comparison experiments on three well-benchmarked retinal vessel segmentation datasets fully verify the effectiveness and superiority of the proposed RetiNeXt. To further demonstrate the universality of RetiNeXt for medical image segmentation, we also conduct sufficient comparative experiments on two classical coronary angiography datasets. Extensive quantitative and qualitative experiments fully show that RetiNeXt outperforms other state-of-the-art methods with only 0.4M of trainable parameters. Shunzhe Shen, Wuzhen Shi, Wenming Cao 0001, Lei Bi 0001, Xiaokang Yang 0001, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | MSPFM: Multi-Scale Pyramid Fusion Mamba for Medical Image Classification
Wuzhen Shi, Daquan Feng, Wenming Cao 0001 |
Vis. Comput. | 5 |
| 2025 | Multi-weather Image Restoration via Histogram-Based Transformer Feature Enhancement
Anyu Lai, Wuzhen Shi, Wenming Cao 0001 |
Vis. Comput. | 5 |
| 2025 | Multi-prior guided depth map super-resolution based on a diffusion model
Wuzhen Shi, Jianhua Ji, Wenming Cao 0001, Zhiquan He |
Vis. Comput. | 5 |
| 2025 | Dual prior guided depth image super-resolution with multi-scale transformer fusion network
Jianhua Ji, Wuzhen Shi, Wenming Cao 0001 |
Vis. Comput. | 5 |
| 2024 | HyperEditor: Achieving Both Authenticity and Cross-Domain Capability in Image Editing via HypernetworksabstractEditing real images authentically while also achieving cross-domain editing remains a challenge. Recent studies have focused on converting real images into latent codes and accomplishing image editing by manipulating these codes. However, merely manipulating the latent codes would constrain the edited images to the generator's image domain, hindering the attainment of diverse editing goals. In response, we propose an innovative image editing method called HyperEditor, which utilizes weight factors generated by hypernetworks to reassign the weights of the pre-trained StyleGAN2's generator. Guided by CLIP's cross-modal image-text semantic alignment, this innovative approach enables us to simultaneously accomplish authentic attribute editing and cross-domain style transfer, a capability not realized in previous methods. Additionally, we ascertain that modifying only the weights of specific layers in the generator can yield an equivalent editing result. Therefore, we introduce an adaptive layer selector, enabling our hypernetworks to autonomously identify the layers requiring output weight factors, which can further improve our hypernetworks' efficiency. Extensive experiments on abundant challenging datasets demonstrate the effectiveness of our method. Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
AAAI | 5 |
| 2024 | A Generic Autoregressive Predictive Feedback Framework for Skeleton-Based Action Recognition
Xinpeng Yin, Wenming Cao 0001 |
ACCV (9) | 3 |
| 2024 | Fuzzy Attention-based Geometric Algebra Convolution in Recursive Fusion Network for Medical Image RegistrationabstractDeep learning-based registration methods have significantly improved accuracy; however, they still grapple with the inherent fuzziness in medical images. Challenges such as intensity inhomogeneity, partial volume effects, and noise can distort spatial dependencies, leading to discontinuities in registration results. In this study, we introduce FuzzyGA-Net, a novel model designed to address these issues. Our approach has three key features: (i) We propose type-1 fuzzy attention (T1FA), which re-weights the feature map to mitigate the impact of uncertainty and enhance the network’s focus on the target center, thereby improving registration accuracy. (ii) Geometric Algebra Convolution is utilized throughout the network to enhance the extraction of spatial dependency cues. (iii) Two fusion strategies are employed to provide flexibility in combining information at various potential locations within the network, optimizing overall performance. Experimental results demonstrate that FuzzyGA-Net achieves the highest registration accuracy compared to state-of-the-art methods, while maintaining the smoothness of the deformation field. Our method shows significant improvement in quantitative metrics and exhibits strong generalization capabilities, suggesting its broad applicability in medical image registration. Wenming Cao 0001, Yicha Zhang |
BIBM | 2 |
| 2024 | Multiple Weather Images Restoration Using the Task Transformer and Adaptive Mixup Strategy
Anyu Lai, Hao Wang 0075, Wuzhen Shi, Wenming Cao 0001 |
CGI (1) | 6 |
| 2024 | Tacking over-smoothing: Target-guide progressive dynamic graph learning for 3D skeleton-based human motion prediction
Jianqi Zhong, Jiang Huang, Wenming Cao 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Pure large kernel convolutional neural network transformer for medical image registrationabstractDeformable medical image registration is a fundamental and critical task in medical image analysis. Recently, deep learning-based methods have rapidly developed and have shown impressive results in deformable image registration. However, existing approaches still suffer from limitations in registration accuracy or generalization performance. To address these challenges, in this paper, we propose a pure convolutional neural network module (CVTF) to implement hierarchical transformers and enhance the registration performance of medical images. CVTF has a larger convolutional kernel, providing a larger global effective receptive field, which can improve the network’s ability to capture long-range dependencies. In addition, we introduce the spatial interaction attention (SIA) module to compute the interrelationship between the target feature pixel points and all other points in the feature map. This helps to improve the semantic understanding of the model by emphasizing important features and suppressing irrelevant ones. Based on the proposed CVTF and SIA, we construct a novel registration framework named PCTNet. We applied PCTNet to generate displacement fields and register medical images, and we conducted extensive experiments and validation on two public datasets, OASIS and LPBA40. The experimental results demonstrate the effectiveness and generality of our method, showing significant improvements in registration accuracy and generalization performance compared to existing methods. Our code has been available at https://github.com/fz852/PCTNet. Zhao Fang, Wenming Cao 0001 |
Intell. Data Anal. | 2 |
| 2024 | Multiple Distilling-based spatial-temporal attention networks for unsupervised human action recognitionabstractUnsupervised action recognition based on spatiotemporal fusion feature extraction has attracted much attention in recent years. However, existing methods still have several limitations: (1) The long-term dependence relationship is not effectively extracted at the time level. (2) The high-order motion relationship between non-adjacent nodes is not effectively captured at the spatial level. (3) The model complexity is too high when the cascade layer input sequence is long, or there are many key points. To solve these problems, a Multiple Distilling-based spatial-temporal attention (MD-STA) networks is proposed in this paper. This model can extract temporal and spatial features respectively and fuse them. Specifically, we first propose a Screening Self-attention (SSA) module; this module can find long-term dependencies in distant frames and high-order motion patterns between non-adjacent nodes in a single frame through a sparse metric on dot product pairs. Then, we propose the Frames and Keypoint-Distilling (FKD) module, which uses extraction operations to halve the input of the cascade layer to eliminate invalid key points and time frame features, thus reducing time and memory complexity. Finally, the Dim-reduction Fusion (DRF) module is proposed to reduce the dimension of existing features to further eliminate redundancy. Numerous experiments were conducted on three distinct datasets: NTU-60, NTU-120, and UWA3D, showing that MD-STA achieves state-of-the-art standards in skeleton-based unsupervised action recognition. Jianqi Zhong, Wenming Cao 0001, Jianhua Ji |
Intell. Data Anal. | 3 |
| 2024 | A Transformer-Assisted Cascade Learning Network for Choroidal Vessel Segmentation
Lei Bi 0001, Wuzhen Shi, Yupeng Xu, Wenming Cao 0001, David Dagan Feng |
J. Comput. Sci. Technol. | 8 |
| 2024 | Dual-model Collaborative Learning with Knowledge Clustering for Few-shot Image Classification
Wenming Cao 0001, Zhineng Zhao |
Multim. Tools Appl. | 2 |
| 2024 | Isolation and distillation network for generalized zero-shot learning
Yuchuan Liang, Wenming Cao 0001 |
Neural Comput. Appl. | 2 |
| 2024 | Gradient multi-foci networks for 3D skeleton-based human motion prediction
Junyu Shi, Jianqi Zhong, Zhiquan He, Wenming Cao 0001 |
Neural Comput. Appl. | 4 |
| 2024 | PANet: Pluralistic Attention Network for Few-Shot Image ClassificationabstractAbstract Traditional deep learning methods require a large amount of labeled data for model training, which is laborious and costly in real word. Few-shot learning (FSL) aims to recognize novel classes with only a small number of labeled samples to address these challenges. We focus on metric-based few-shot learning with improvements in both feature extraction and metric method. In our work, we propose the Pluralistic Attention Network (PANet), a novel attention-oriented framework, involving both a local encoded intra-attention(LEIA) module and a global encoded reciprocal attention(GERA) module. The LEIA is designed to capture comprehensive local feature dependencies within every single sample. The GERA concentrates on the correlation between two samples and learns the discriminability of representations obtained from the LEIA. The two modules are complementary to each other and ensure the feature information within and between images can be fully utilized. Furthermore, we also design a dual-centralization (DC) cosine similarity to eliminate the disparity of data distribution in different dimensions and enhance the metric accuracy between support and query samples. Our method is thoroughly evaluated with extensive experiments, and the results demonstrate that with the contribution of each component, our model can achieve high-performance on four widely used few-shot classification benchmarks of miniImageNet, tieredImageNet, CUB-200-2011 and CIFAR-FS. Wenming Cao 0001, Tianyuan Li, Qifan Liu, Zhiquan He |
Neural Process. Lett. | 1 |
| 2024 | Spatiotemporal Progressive Inward-Outward Aggregation Network for skeleton-based action recognition
Xinpeng Yin, Jianqi Zhong, Deliang Lian, Wenming Cao 0001 |
Pattern Recognit. | 4 |
| 2024 | Attribute disentanglement and re-entanglement for generalized zero-shot learning
Yucuan Liang, Zhenqi Zhang, Wenming Cao 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | Image Classification on Hypersphere LossabstractThe effectiveness of an image classification system depends on the following two key components: 1) the feature learning module and 2) the classification module. A well-designed loss function can not only enhance the classification ability of the latter but also improve the feature extraction capabilities of the former. This article devises a novel hypersphere loss function, which enhances the intraclass compactness and interclass separability of feature vectors given by the feature learning module. Furthermore, a new generalized class center is introduced into the loss function to handle the inevitable variability in samples (such as illumination, background, blurriness, low resolution, etc.) within the same class. Then, an alternative learning strategy is employed to optimize trainable parameters and class centers. Specifically, we first fix the trainable parameters of the deep learning model and calculate class centers using the exponentially weighted moving average method. Subsequently, we fix the generalized class centers and update the model's trainable parameters using mini-batch stochastic gradient descent. The proposed algorithm is evaluated on a range of typical tasks, including standard image classification, face verification, object detection, and retail product checkout. The results demonstrate that our proposed algorithm outperforms several state-of-the-art approaches. Hao Wang 0075, Jinpeng Cao, Zhanglei Shi, Andrew Chi-Sing Leung, Ruibin Feng, Wenming Cao 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2024 | Multi-Semantics Aggregation Network Based on the Dynamic-Attention Mechanism for 3D Human Motion PredictionabstractGraph convolutional network-based methods have recently shown promising performance in skeleton-based data processing. However, these methods have two critical issues in skeleton-based motion prediction tasks: First, graph modeling of motion poses is based on the fixed graph according to the physical connection of human joints and ignores the exploration of deep implicit information based on human dynamic kinetics. Second, existing methods usually use motion information in a single semantic space to model the whole motion sequences, underestimating diverse semantic patterns for improving the modeling ability. To address the first issue, we propose the Attention-based Dynamic Graph Convolution method, which tries to capture implicit semantic information dynamically. To address the second issue, we propose the Kinematic-based Semantics Aggregation Block (KSAB), which combines various semantic features from four semantic perspectives to rich motion representation. Integrating the above two designs, we propose a novel Multi-Semantics Aggregation Network (MANet), resulting in more comprehensive feature extraction in dynamic implicit semantics learning to enhance motion prediction. Extensive experiments are conducted to validate the effectiveness of MANet, which outperforms state-of-the-art methods by 10.9%, 6.6%, and 19.6% in terms of MPJPE for motion prediction on Human3.6M, CMU Mocap, and 3DPW datasets, respectively. Junyu Shi, Jianqi Zhong, Wenming Cao 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | Co-salient object detection with iterative purification and predictive optimizationabstractCo-salient object detection (Co-SOD) aims to identify and segment commonly salient objects in a set of related images. However, most current Co-SOD methods encounter issues with the inclusion of irrelevant information in the co-representation. These issues hamper their ability to locate co-salient objects and significantly restrict the accuracy of detection. To address this issue, this study introduces a novel Co-SOD method with iterative purification and predictive optimization (IPPO) comprising a common salient purification module (CSPM), predictive optimizing module (POM), and diminishing mixed enhancement block (DMEB). These components are designed to explore noise-free joint representations, assist the model in enhancing the quality of the final prediction results, and significantly improve the performance of the Co-SOD algorithm. Furthermore, through a comprehensive evaluation of IPPO and state-of-the-art algorithms focusing on the roles of CSPM, POM, and DMEB, our experiments confirmed that these components are pivotal in enhancing the performance of the model, substantiating the significant advancements of our method over existing benchmarks. Experiments on several challenging benchmark co-saliency datasets demonstrate that the proposed IPPO achieves state-of-the-art performance. Yuhuan Wang, Hao Wang 0075, Wuzhen Shi, Wenming Cao 0001 |
Virtual Real. Intell. Hardw. | 5 |
| 2023 | Class-Balanced Universal Perturbations for Adversarial TrainingabstractUniversal attack generates image-agnostic perturbation called universal adversarial perturbation (UAP), which can be added to all samples in the data distribution to fool the classifier. However, a universal perturbation will likely mislead the classifier to identify most adversarial examples as the same label, resulting in the imbalance of attack strength between classes. In this paper, we propose class-balanced UAPs that enlarge the dispersion of the predicted labels for adversarial examples. To ensure attack strength and balance simultaneously, we design a novel diversity objective containing probability calibration and penalty regularizer, which fully considers the predicted label distribution between samples and the predicted probability distribution within samples. Furthermore, we apply class-balanced attacks in adversarial training to defend against universal perturbations since the class-balanced UAP provides diverse perturbation directions. We correspondingly reformulate adversarial training from the min-max optimization problem into a new two-stage framework. Experiments on several benchmark datasets demonstrate that the class-balanced attack achieves better performance than the universal attack, while adversarial training with class-balanced UAP achieves state-of-the-art results in clean accuracy and robustness to universal perturbations. Kexue Ma, Guitao Cao, Mengqian Xu, Chunwei Wu, Wenming Cao 0001 |
IJCNN | 6 |
| 2023 | Chaos to Order: A Label Propagation Perspective on Source-Free Domain AdaptationabstractSource-free domain adaptation (SFDA), where only a pre-trained source model is used to adapt to the target distribution, is a more general approach to achieving domain adaptation in the real world. However, it can be challenging to capture the inherent structure of the target features accurately due to the lack of supervised information on the target domain. By analyzing the clustering performance of the target features, we show that they still contain core features related to discriminative attributes but lack the collation of semantic information. Inspired by this insight, we present Chaos to Order (CtO), a novel approach for SFDA that strives to constrain semantic credibility and propagate label information among target subpopulations. CtO divides the target data into inner and outlier samples based on the adaptive threshold of the learning state, customizing the learning strategy to fit the data properties best. Specifically, inner samples are utilized for learning intra-class structure thanks to their relatively well-clustered properties. The low-density outlier samples are regularized by input consistency to achieve high accuracy with respect to the ground truth labels. In CtO, by employing different learning strategies to propagate the labels from the inner local to outlier instances, it clusters the global samples from chaos to order. We further adaptively regulate the neighborhood affinity of the inner samples to constrain the local semantic credibility. In theoretical and empirical analyses, we demonstrate that our algorithm not only propagates from inner to outlier but also prevents local clustering from forming spurious clusters. Empirical evidence demonstrates that CtO outperforms the state of the arts on three public benchmarks: Office-31, Office-Home, and VisDA. Chunwei Wu, Guitao Cao, Yan Li 0063, Xidong Xi, Wenming Cao 0001 |
ACM Multimedia | 5 |
| 2023 | Explore Across-Dimensional Feature Correlations for Few-Shot Learning
Tianyuan Li, Wenming Cao 0001 |
PRCV (3) | 2 |
| 2023 | Deformable image registration with attention-guided fusion of multi-scale deformation fieldsabstractAbstract Deformable medical image registration plays a crucial role in theoretical research and clinical application. Traditional methods suffer from low registration accuracy and efficiency. Recent deep learning-based methods have made significant progresses, especially those weakly supervised by anatomical segmentations. However, the performance still needs further improvement, especially for images with large deformations. This work proposes a novel deformable image registration method based on an attention-guided fusion of multi-scale deformation fields. Specifically, we adopt a separately trained segmentation network to segment the regions of interest to remove the interference from the uninterested areas. Then, we construct a novel dense registration network to predict the deformation fields of multiple scales and combine them for final registration through an attention-weighted field fusion process. The proposed contour loss and image structural similarity index (SSIM) based loss further enhance the model training through regularization. Compared to the state-of-the-art methods on three benchmark datasets, our method has achieved significant performance improvement in terms of the average Dice similarity score (DSC), Hausdorff distance (HD), Average symmetric surface distance (ASSD), and Jacobian coefficient (JAC). For example, the improvements on the SHEN dataset are 0.014, 5.134, 0.559, and 359.936, respectively. Zhiquan He, Yupeng He, Wenming Cao 0001 |
Appl. Intell. | 3 |
| 2023 | Multi-layer noise reshaping and perceptual optimization for effective adversarial attack of imagesabstractAbstract Adversarial attack aims to fail the deep neural network by adding a small amount of perturbation to the input image, in which the attack success rate and resulting image quality are maximized under the lp norm perturbation constraint. However, the lp norm is not accurately correlated to human perception of image quality. Attack methods based on l0 norm constraint usually suffer from the high computational cost due to the iterative search for candidate pixels to modify. In this work, we explore how perceptual quality optimization can be incorporated into the adversarial attack design and propose a two-stage attack method to reshape the adversarial noise by an initial attack and optimize the visual quality of the attacked images without sacrificing the attack success rate. Specifically, we construct a visual attention network to generate a perceptual attention map to modulate the adversarial noise generated by a base attack method. The network is trained to maximize the visual quality in Structural Similarity Index Metric (SSIM) while achieving the same attack success rate. To improve the image perceptual quality further, we propose a fast search algorithm to perform an iterative block-wise pruning of the adversarial noise. We evaluate our method on the mini-ImageNet dataset against three different defense schemes. The results have demonstrated that our method can achieve better attack performance in image quality, attack success rate, and efficiency than the state-of-the-art attack methods. Zhiquan He, Xujia Lan, Jianhe Yuan, Wenming Cao 0001 |
Appl. Intell. | 4 |
| 2023 | A novel inference paradigm based on multi-view prototypes for one-shot semantic segmentation
Guitao Cao, Wenming Cao 0001 |
Appl. Intell. | 3 |
| 2023 | Geometric algebra-based multiscale encoder-decoder networks for 3D motion prediction
Jianqi Zhong, Wenming Cao 0001 |
Appl. Intell. | 2 |
| 2023 | Two analog neural models with the controllability on number of assets for sparse portfolio design
Hao Wang 0075, Andrew Chi-Sing Leung, Andy Hau-Ping Chan, Anthony G. Constantinides, Wenming Cao 0001 |
Neurocomputing | 5 |
| 2023 | Visual-semantic consistency matching network for generalized zero-shot learning
Zhenqi Zhang, Wenming Cao 0001 |
Neurocomputing | 2 |
| 2023 | A novel sample and feature dependent ensemble approach for Parkinson's disease detectionabstractAbstract Parkinson’s disease (PD) is a neurological disease that has been reported to have affected most people worldwide. Recent research pointed out that about 90% of PD patients possess voice disorders. Motivated by this fact, many researchers proposed methods based on multiple types of speech data for PD prediction. However, these methods either face the problem of low rate of accuracy or lack generalization. To develop an approach that will be free of these issues, in this paper we propose a novel ensemble approach. These paper contributions are two folds. First, investigating feature selection integration with deep neural network (DNN) and validating its effectiveness by comparing its performance with conventional DNN and other similar integrated systems. Second, development of a novel ensemble model namely EOFSC (Ensemble model with Optimal Features and Sample Dependant Base Classifiers) that exploits the findings of recently published studies. Recent research pointed out that for different types of voice data, different optimal models are obtained which are sensitive to different types of samples and subsets of features. In this paper, we further consolidate the findings by utilizing the proposed integrated system and propose the development of EOFSC. For multiple types of vowel phonations, multiple base classifiers are obtained which are sensitive to different subsets of features. These features and sample-dependent base classifiers are integrated, and the proposed EOFSC model is constructed. To evaluate the final prediction of the EOFSC model, the majority voting methodology is adopted. Experimental results point out that feature selection integration with neural networks improves the performance of conventional neural networks. Additionally, feature selection integration with DNN outperforms feature selection integration with conventional machine learning models. Finally, the newly developed ensemble model is observed to improve PD detection accuracy by 6.5%. Chinmay Chakraborty, Zhiquan He, Wenming Cao 0001, Yakubu Imrana, Joel J. P. C. Rodrigues |
Neural Comput. Appl. | 4 |
| 2023 | Correction to: Multi-level context-driven interaction modeling for human future trajectory prediction
Zhiquan He, Hao Sun 0024, Wenming Cao 0001, Henry Z. He |
Neural Comput. Appl. | 3 |
| 2023 | Dual-domain reciprocal learning design for few-shot image classification
Qifan Liu, Yaozong Chen, Wenming Cao 0001 |
Neural Comput. Appl. | 3 |
| 2023 | Diegraph: dual-branch information exchange graph convolutional network for deformable medical image registration
Longji Wang, Zhiyue Yan, Wenming Cao 0001, Jianhua Ji |
Neural Comput. Appl. | 3 |
| 2023 | Cycle optimization metric learning for few-shot classification
Qifan Liu, Wenming Cao 0001, Zhihai He |
Pattern Recognit. | 2 |
| 2023 | Geometric algebra-based multiview interaction networks for 3D human motion prediction
Jianqi Zhong, Wenming Cao 0001 |
Pattern Recognit. | 2 |
| 2023 | l0-norm based Short-term Sparse Portfolio Optimization Algorithm Based on Alternating Direction Method of MultipliersabstractThis paper proposes a novel short-term sparse portfolio optimization (SSPO) model based on ℓ 0 -norm. Compared with existing approaches, this model selects the portfolio based on the short-term increasing potential of assets, and an ℓ 0 -norm constraint is introduced to directly control the maximum number of non-zero assets in selected portfolios. Unlike the ℓ 1 -norm based methods, the no-short-selling constraints can be directly used in our proposed model. Besides, a sparse regularization term is introduced to eliminate trivial trades in the SSPO system. Moreover, to solve the contained non-convex optimization system, an algorithm based on the concept of the alternating direction method of multipliers (ADMM) is developed. The convergence of the proposed algorithm is also investigated. Finally, the effectiveness of the proposed approach is demonstrated by some numerical experiments on four real-world datasets. Hao Wang 0075, Wenming Cao 0001 |
Signal Process. | 4 |
| 2023 | PCFN: Progressive Cross-Modal Fusion Network for Human Pose TransferabstractThe goal of human pose transfer is to transfer the human in the image from the original pose to the desired one. Existing methods utilizing progressive manner have achieved great success. However, they fail to remove background distraction and preserve appearance details in synthesized images since the correlation between the image and pose is not fully exploit. To this end, we propose a novel progressive cross-modal fusion network (PCFN), which consists of multiple cascaded cross-modal fusion blocks (CMFBs). Each CMFB comprises a feature fusion module (FFM) and a cross-modal module (CMM) to take full advantage of appearance and shape information. From an overall perspective, FFM fully exploits the correlation between image features and pose features through the residual gated convolution. Benefitting from feature integration and dynamic selection, CMFB can extract useful information from the image-pose stream. From a local perspective, CMM utilizes the feature-conditioned gated convolution and the pose-guided heterogeneous attention mechanism to update all codes in a crossing manner and enhance the interaction between fusion information and structural information. Qualitative and quantitative experiments demonstrate the superiority of PCFN compared to state-of-the-art methods, which can transfer the correct human features and increase the authenticity of the generated images. At the same time, PCFN can also be applied to supplement the dataset for person re-identification (ReID). PCFN works well for human pose transfer, and our usage of the gated convolution and the attention mechanism also provides references for other conditional generation tasks. Rui Wang 0034, Wenming Cao 0001, Wei Xiang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Expensive Multiobjective Optimization Based on Information Transfer SurrogateabstractObjective value estimation based on computationally efficient surrogate models is widely used to reduce the computational cost in solving expensive multiobjective optimization problems (MOPs). However, due to the scarcity of training data and the lack of data sharing between training tasks in a surrogate-based system, the estimation effectiveness of the surrogate models might not be satisfactory. In this study, we present a novel surrogate methodology based on information transfer to deal with this problem. Particularly, in the proposed framework, the objectives of an MOP that may have little apparent similarity or correlation are linearly mapped to a number of related tasks. Afterward, the related tasks are used to train a multitask Gaussian process (MTGP). MTGP expands the training data leading to more confident learning of the parameters of the model. The predicted values of the objective functions can be obtained by a reverse mapping from the learned MTGP model. In this way, the computational burden of the expensive objective functions of an MOP can be substantially reduced while maintaining good estimation accuracy. MTGP facilitates mutual information transfer across tasks, avoids learning from scratch for new tasks, and captures the underlying structural information between tasks. The proposed surrogate approach is merged into MOEA/D to address MOPs. Experimental tests under various scenarios indicate that the resultant algorithm outperforms other state-of-the-art surrogate-based multiobjective optimization algorithms. Jianping Luo, YongFei Dong, Zexuan Zhu 0001, Wenming Cao 0001, Xia Li 0006 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Feature Pyramid Vision Transformer for MedMNIST Classification DecathlonabstractMedMNIST is a medical dataset proposed to block the need for medical knowledge, but there is currently no model that can generalize well on all its sub-datasets. Owing to the inadequacy of long-range relation modeling, models based on convolutional neural networks (CNNs) cannot fully learn the information of images. Besides, relying only on high-level features limits the generalization effect as well. All of these remain challenges for MedMNIST Classification Decathlon. In this paper, we proposed Feature Pyramid Vision Transformer (FPViT), a strong alternative for MedMNIST Classification Decathlon. Our FPViT exhibits enhanced feature learning and modeling capabilities, which merits both residual network (ResNet) and Vision Transformer (ViT). Transformers in our model take the features extracted by ResNet as sequences to capture global contexts which compensate for the lack of locality of convolution operations. Moreover, the feature pyramid designed in our model effectively utilizes the multi-scale feature maps from basic layers of ResNet. These multi-scale features from low-level to high level enable our model to have better adaptability. And, the final prediction is based on the multi-scale ViT and the original ResNet heads. Through experiments, our FPViT can achieve superior classification and generalization on MedMNIST than state-of-the-art methods. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
IJCNN | 5 |
| 2022 | Adversarial Discriminative Feature Separation for Generalization in Reinforcement LearningabstractImporving the generalization ability of an agent is an important and challenging task in deep reinforcement learning (RL). Procedually generated environment is an important benchmark for testing generalization in deep RL. In this benchmark, each game consists of multiple levels, each level is an algorithmically created environment instance with a unique configuration of its factors of variation. Existing methods (e.g., regularization, data augmentation) for improving the generalization of RL agent do not learn well the invariant representation among multiple levels. Besides, existing methods for learning invariant representations in RL using adversarial training can only learn invariant information across two levels. To solve this problem, we propose Adversarial Discriminative Feature Separate (ADFS). First, ADFS design a new discriminator for distinguishing whether two observations belong to the same level. Thus, the policy encoder is encouraged to learn invariant information between multiple levels. Second, it separates the representation of observation into level-invariant features and level-discriminative features, so that correction of the optimization direction of the discriminator. The discriminative features are learned by reducing the similarity of specific features intra-levels and increasing that of inter-levels, respectively. Experimental results demonstrate that our method is quite competitive with existing state-of-the-art methods on Procgen Benchmark. Chunwei Wu, Xidong Xi, Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
IJCNN | 6 |
| 2022 | PassAugment: Pass Nodes Importance in Graph Data Augmentation for Graph ClassificationabstractData augmentation has been widely introduced into graph-based tasks to improve the generalizability of models. Based on empirical hypothesis, we show that the node in a graph has different importance, the important one is critical for classification task while the unimportant one hurts the performance. However, there are few works in data augmentation addressing the information propagation of nodes with different importance. In this work, we propose a novel graph data augmentation algorithm for graph classification task, called PassAugment, aiming to pass these importance in graph data augmentation. After distinguishing the importance of all nodes in each graph using the saliency map, we design a data augmentation approach including two strategies: (i) randomly adding edges between the important nodes and the other nodes to globally improve the effective information passing, and (ii) randomly removing edges between the unimportant nodes and their neighbors to locally reduce the ineffective information passing. More importantly, our proposed approach as a standalone module can be combined with many GNNs architectures. Experimental results on graph classification task show that our approach consistently improves the accuracy and achieves or closely matches the state-of-the-art performance. Xiaohu Li, Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
SMC | 5 |
| 2022 | Geometric algebra graph neural network for cross-domain few-shot classification
Qifan Liu, Wenming Cao 0001 |
Appl. Intell. | 2 |
| 2022 | Geometric deep learning: progress, applications and challenges
Wenming Cao 0001, Canta Zheng, Zhiyue Yan, Weixin Xie |
Sci. China Inf. Sci. | 1 |
| 2022 | A dual attention model based on probabilistically mask for 3D human motion prediction
Wenming Cao 0001, Jianqi Zhong |
Neurocomputing | 1 |
| 2022 | Multi-strategy mutual learning network for deformable medical image registration
Wenming Cao 0001, Ye Duan, Guitao Cao, Deliang Lian |
Neurocomputing | 2 |
| 2022 | Implicit user relationships across sessions enhanced graph for session-based recommendationabstractSession-based recommendation aims to predict users’ next preference based on the sequence of their own history preferences in a short period. Most state-of-the-art methods model the session as a graph using graph neural networks (GNN) to capture the dynamic transitions between items within sessions. However, the complex and hidden correlations between different sessions are not adequately addressed, especially during the testing stage. We argue that session-based recommendation tasks can be improved by exploiting the correlations between different sessions in both training and testing. To this end, we propose a novel three-GNN-based recommendation framework to exploit the intra- and inter-session item correlations and the session-session correlations. The first one is a graph-based multi-layer perceptron to learn the inter-session item representations in a contrastive learning scheme guided by a contrastive loss function. The second one is a multi-relation graph attention network for intra-session item representations. The two item embeddings are combined through a position attention scheme to form the session representation, which is modulated and enhanced by the third extra session GNN by capturing the session-session correlations. The three levels of correlations are used in the training and testing stages in a joint prediction manner. To alleviate the data sparsity issue faced by the session GNN, we expand the session items by incorporating the neighboring items in the global item graph built from the entire training sessions. We have evaluated our method on multiple benchmark datasets. The results have shown that our joint recommendation method based on session correlation has significantly improved the recommendation accuracy over the state-of-the-art by more than 10%. Wenming Cao 0001, Yishan Liu, Guitao Cao, Zhiquan He |
Inf. Sci. | 1 |
| 2022 | Adaptive weight multi-channel center similar deep hashing
Xinghua Liu 0004, Guitao Cao, Qiubin Lin, Wenming Cao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Geometric machine learning: research and applications
Wenming Cao 0001, Canta Zheng, Zhiyue Yan, Zhihai He, Weixin Xie |
Multim. Tools Appl. | 1 |
| 2022 | Multi-level context-driven interaction modeling for human future trajectory predictionabstractAbstract Human trajectory prediction is a challenging task with important applications such as intelligent surveillance and autonomous driving. We recognize that pedestrians in close and distant neighborhoods have different impacts on the person’s decision of future movements. Local scene context and global scene layout also affect the movement decision differently. Existing methods have not adequately addressed these interactions between humans and the multi-level contexts occurring at different spatial and temporal scales. To this end, we propose a multi-level context-driven interaction modeling (MCDIM) method for human future trajectory learning and prediction. Specifically, we construct a multilayer graph attention network (GAT) to model the hierarchical human–human interactions. An extra set of long short-term memory networks is designed to capture the correlations of these human–human interactions at different temporal scales. To model the human–scene interactions, we explicitly extract and encode the global scene layout features and local context features in the neighborhood of the person at each time step and capture the spatial–temporal information of the interactions between human and the local scene contexts. The human–human and human–scene interactions are incorporated into the multi-level GAT-based network for accurate prediction of future trajectories. We have evaluated the method on benchmark datasets: the walking pedestrians dataset provided by ETH Zurich (ETH) and the crowd data provided by the University of Cyprus. The results demonstrate that our MCDIM method outperforms existing methods, being able to generate more accurate and plausible trajectories for pedestrians. The average performance gain is 2 and 3 percentage points in terms of the average displacement error and final displacement error, respectively. Zhiquan He, Hao Sun 0024, Wenming Cao 0001, Henry Z. He |
Neural Comput. Appl. | 3 |
| 2022 | Feature self-calibration network with global-local training strategy for multi-region deformable medical image registration
Wenming Cao 0001, Deliang Lian |
Neural Comput. Appl. | 2 |
| 2022 | QMEDNet: A quaternion-based multi-order differential encoder-decoder model for 3D human motion prediction
Wenming Cao 0001, Jianqi Zhong |
Neural Networks | 1 |
| 2022 | Self-guided information for few-shot classification
Zhineng Zhao, Qifan Liu, Wenming Cao 0001, Deliang Lian, Zhihai He |
Pattern Recognit. | 3 |
| 2022 | GA-CNN: Convolutional Neural Network Based on Geometric Algebra for Hyperspectral Image ClassificationabstractConvolutional neural networks (CNNs) have achieved state-of-the-art performance in hyperspectral images (HSIs) classification, which is widely used for the analysis of remotely sensed images. HSI includes spectral and spatial information from several hundreds of spectral data channels. Recent CNN models deal with various bands of HSIs as independent channels, which may lead to the loss of dependencies between different channels or the loss of associated information between each channel and the global. This article proposes a novel CNN model based on geometric algebra (GA), dubbed GA-CNN, to process the HSIs in a holistic way without losing the interrelationship among channels. Specifically, taking advantage of GA, different band images are represented as GA multivectors to capture the inherent structures and preserve the correlation of those channels. In particular, all the basic modules of our model, such as convolutional layers and the backpropagation algorithm, are extended to the GA domain. We evaluate the performance of the proposed GA-CNN model in classification tasks on four well-known HSI datasets. The experimental results indicate that our GA-CNN model outperforms traditional and state-of-the-art real-valued CNNs with higher classification accuracy and fewer model parameters. Rui Wang 0034, Yi Wang 0063, Xiangyang Wang 0003, Wenming Cao 0001, Wei Xiang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2021 | DCFG: Discovering Directional CounterFactual Generation for Chest X-raysabstractWhile Deep Neural Networks (DNNs) are achieving state-of-the-art performance on medical domains across a variety of tasks, the need for explainability of model predictions in these high-stakes tasks is still lacking. Current for the explainability in model predictions potentially relies on the supervised counterfactual generation that is time-consuming and direction uncontrollable. Yet, the counterfactual generation needs to be easy to implement and have a controllable direction. In light of this trend, we propose an approach for the unsupervised latent direction search of black-box models that are steerable to the user by enabling the user to effectively explore counterfactual generation in a directional way, without relying on domain- or data-specific assumptions. To identify these explainable directions, we use Principal Component Analysis (PCA), a general manifold learning framework to extract low-dimensional subspaces based on a local noise injection of the pre-trained generative model, so that a small perturbation in the subspaces would provide enough change in the resulting data. With experiments on three real-world CXR datasets involving 6 tasks, we find that our approach is capable of learning explainable predictions that discard unrelated confounding factors. Moreover, our method enables practitioners to edit directions to better understand which features are used for predictions. Yan Li 0063, Chunwei Wu, Xidong Xi, Guitao Cao, Wenming Cao 0001 |
BIBM | 6 |
| 2021 | Spatial Assembly Networks for Image Representation LearningabstractIt has been long recognized that deep neural networks are sensitive to changes in spatial configurations or scene structures. Image augmentations, such as random translation, cropping, and resizing, can be used to improve the robustness of deep neural networks under spatial transforms. However, changes in object part configurations, spatial layout of object, and scene structures of the images may still result in major changes in the their feature representations generated by the network, creating significant challenges for various visual learning tasks, including representation or metric learning, image classification and retrieval. In this work, we introduce a new learnable module, called spatial assembly network (SAN), to address this important issue. This SAN module examines the input image and performs a learned re-organization and assembly of feature points from different spatial locations conditioned by feature maps from previous network layers so as to maximize the discriminative power of the final feature representation. This differentiable module can be flexibly incorporated into existing network architectures, improving their capabilities in handling spatial variations and structural changes of the image scene. We demonstrate that the proposed SAN module is able to significantly improve the performance of various metric / representation learning, image retrieval and classification tasks, in both supervised and unsupervised learning scenarios. Yang Li 0091, Shichao Kan, Jianhe Yuan, Wenming Cao 0001, Zhihai He |
CVPR | 4 |
| 2021 | Sample Efficient Lung Segmentation Using Group Structured Conditional Variational Data ImputationabstractPatients infected with COVID-19 can lead to their Chest X-rays (CXRs) with opacifications rendered regions, which may produce incomplete lung segmentation in automated image analysis models. To tackle this issue, we propose a Group structured Conditional Variational data Imputation model to capture the missing data accurately with conditional distribution, where the high-dimensional probability distribution is narrow down to a small latent space to account for unobserved features. This work particularly arises in the fight against COVID-19 that effectively modeling a segmentation of plausible can be presented to a subsequent automated risk scoring and treatment. We train this model with limited CXRs data to demonstrate the abilities on the task of data imputation and proved to be effective though with relatively small datasets. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
ICME | 3 |
| 2021 | Structure-Oriented Progressive Low-Rank Image Restoration for Defending Adversarial AttacksabstractDeep neural networks recognize objects by analyzing local image details and summarizing their information along the inference layers to derive the final decision. Because of this, they are prone to adversarial attacks. On the other hand, human eyes recognize objects based on their global structures and semantic cues, instead of local image textures. In this work, we propose to develop a structure-oriented progressive low-rank image completion method to remove unneeded texture details from the input images and shift the bias of deep neural networks towards global object structures and semantic cues. We formulate the problem into a low-rank matrix completion problem with progressively smoothed rank functions to avoid local minimums. Our experimental results demonstrate the proposed method is able to successfully remove the insignificant local image details while preserving important global object structures. Zhiqun Zhao, Hengyou Wang, Hao Sun 0024, Wenming Cao 0001, Zhihai He |
ICME | 4 |
| 2021 | Debiased Prototype Network for Adversarial Domain AdaptationabstractDomain adaptation is an important and challenging task. Existing adversarial domain adaptation methods explore the relationship between the source and target domains, with the knowledge learned in the source domain supporting the target domain task. The quality of the knowledge will affect the task performance of the transfer, i.e., the higher the quality of the knowledge, the better the transfer task performance. To obtain better domain-invariant knowledge, we extract domain-invariant semantic information over the unit sphere via the prototype network. With the help of geometric constraints from the hypersphere, the features can be more tightly clustered with the estimated prototype (representatives of each class). Adaptation is achieved by adversarial learning to align the domain distribution, which enhances the transferability of the learned features and obtains the basic prototype. Since the basic prototypes dominantly computed from the source domain are biased against the expected domain-invariant prototype, a debiased method is further proposed to obtain the domain-invariant prototypes. Specifically, our method diminishes the intra- and inter- class bias to achieve the class-level alignment. Extensive experiments demonstrate that our model achieves state-of-the-art performance on several domain adaptation benchmark datasets. Our code is available at https://github.com/Chunweiwu-source/DPN. Chunwei Wu, Guitao Cao, Wenming Cao 0001 |
IJCNN | 3 |
| 2021 | RGA-CNNs: convolutional neural networks based on reduced geometric algebra
Rui Wang 0034, Miaomiao Shen, Xiangyang Wang 0003, Wenming Cao 0001 |
Sci. China Inf. Sci. | 4 |
| 2021 | IAE-ClusterGAN: A new Inverse autoencoder for Generative Adversarial Attention Clustering network
Chao Ling, Guitao Cao, Wenming Cao 0001 |
Neurocomputing | 3 |
| 2021 | Progressive anatomically constrained deep neural network for 3D deformable medical image registrationabstractThe 3D deformable image registration is one of the most challenging tasks in medical image analysis. Due to the large and complex deformation in 3D medical images, many deep neural network based methods have been proposed to improve the image similarity after registration, among which recursive cascading network structure is one of the state-of-the-art. However, most existing works rely on the pixel-level image similarities to achieve anatomical rationality and overlook the global-level resemblance between the two structures. Therefore, the resulting registration is not quite clinically valuable. To this end, in this work, we propose a Progressive Anatomically Constrained deep neural Network (PACN) to incorporate the anatomical priors into a progressive cascading registration network to improve the anatomical plausibility as well as the pixel-level similarity of the registration results. Specifically, an Anatomical Constraint Encoder (ACE) network is proposed to encode the global context of the anatomical segmentations and attached to the dense registration network to form a registration unit. Repeated such units forming a cascading framework progressively warps the moving image toward the fixed one, with the output warped image of one unit as the input of the next unit. In this design, the global anatomical priors along with the pixel-level local information are used to guide the model learning process to produce high quality deformation field. Based on this, we explore two frameworks to investigate their registration effectiveness, one attaches the anatomical constraint encoder (ACE) to every dense registration sub-network and the other one attaches ACE only to the last dense registration unit. We test the two frameworks on benchmarks of three liver image datasets SLIVER, LiTS and LSPIG, and one brain dataset LPBA. Our two frameworks have achieved significantly better results in terms of average Dice score than the state-of-the-art baseline method on three liver datasets and comparable on LPBA when both tested with up to three cascades. Wenming Cao 0001, Zhiquan He |
Neurocomputing | 2 |
| 2021 | A cascaded registration network RCINet with segmentation mask
Wenlan Zou, Wenming Cao 0001, Zhiquan He, Zhihai He |
Neural Comput. Appl. | 3 |
| 2021 | Learned Model Composition With Critical Sample Look-Ahead for Semi-Supervised Learning on Small Sets of Labeled SamplesabstractIn this work, we propose to push the performance limit of semi-supervised learning on very small sets of labeled samples by developing a new method called learned model composition with critical sample look-ahead (LMCS). Training efficient deep neural networks on much smaller sets of labeled samples is a challenging problem. With a small labeled set, the initial network suffers from low accuracy. Based on this error-prone network, the subsequent semi-supervised learning process will be fragile and unstable. To address this issue, we propose to introduce a look-ahead master model to identify the correct direction of model evolution to effectively guide the semi-supervised learning process of the student model. Specifically, our proposed LMCS method explores two major ideas. First, it introduces a new learned model composition structure so that we can compose a more efficient master network from student models of past iterations through a network learning process. Second, we develop a new method, called confined maximum entropy search, to discover new critical samples near the model decision boundary and provide the master model with look-ahead access to these samples to enhance its guidance capability. Our extensive experimental results demonstrate that the proposed LMCS method outperforms the state-of-the-art semi-supervised learning methods, especially on small sets of labeled samples. For example, on the CIFAR-10 dataset, with a very small set of 80 labeled samples, our method outperforms Google's MixMatch method, reducing the error rate by more than 10%. Yang Li 0091, Shichao Kan, Wenming Cao 0001, Zhihai He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Mask Cross-Modal Hashing NetworksabstractDue to the rapid development of deep learning, cross-modal retrieval has achieved significant progress in recent years. Moreover, cross-modal hashing has recently attracted considerable attention to multi-modal retrieval applications due to its advantages of low storage costs and fast retrieval speed. However, it is still a challenging problem due to an existing semantic heterogeneity gap between different modalities. In order to further narrow the gap and obtain more effective hash codes, we put forward a novel mask deep cross-modal hashing (MDCH) approach to explore the similarity between inter-modal instances. The main contributions of this paper are that: (1) we attempt to introduce semantic mask information into cross-modal hashing retrieval, (2) we alternately train intra-modal and inter-modal networks to fully mine the semantic relationship between different modalities. The semantic mask can improve the semantic information of the image feature. While inter-modal similarity, explored by inter-modal networks, focuses on enforcing images and their corresponding text tags to have similar hash codes, intra-modal similarity, explored by intra-modal networks, can retain local structural information embedded in each modality to achieve internal similarity. A large number of experiments conducted on three datasets demonstrate that our proposed MDCH approach is superior to several state-of-the-art cross-modal hashing approaches. Qiubin Lin, Wenming Cao 0001, Zhiquan He, Zhihai He |
IEEE Trans. Multim. | 2 |
| 2020 | A Dynamic Group Equivariant Convolutional Networks for Medical Image AnalysisabstractGroup equivariant Convolutional Neural Networks (G-CNNs) has led to big empirical success in the medical domain, one fundamental assumption is that equivariance provides a powerful inductive bias for medical images. By leveraging concepts from group representation theory, we can generalize vanilla Convolutional Neural Networks (CNNs) to G-CNN. Currently, although embedding an arbitrary equivariance to CNNs can learn powerful disentangled representations in a higher dimensional domain, they lack explicit means to learn meaningful relationships among the equivariant convolutional kernels. In this paper, we propose a generalization of the dynamic convolutional method, named as dynamic group equivariant convolution, to strengthen the relationships and increase model capability by aggregating multiple group convolutional kernels via attention. Meanwhile, we generalize attention to an equivariant one to preserve equivariant of dynamic group convolution. In our approach, this leads to a flexible framework that enables a dynamic convolutional in G-CNNs by means of a dynamic routing layer expansions. We demonstrate that breast tumor classification is substantial improvements when compared to a recent baseline architecture. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
BIBM | 3 |
| 2020 | Semantic deep cross-modal hashing
Qiubin Lin, Wenming Cao 0001, Zhihai He, Zhiquan He |
Neurocomputing | 2 |
| 2019 | Emotion Recognition from Children Speech Signals Using Attention Based Time Series Deep LearningabstractChildren's emotions expression concentrates in the acoustic aspects such as the tones and timbres of the voice instead of the semantics, and there are a lot of lengthy fragments in their speech. This paper proposes an emotion recognition model using the time series deep learning technology, named attention based Bi-directional Long Short-Term Memory (CNN-BiLSTM) to extract the emotional features. After preprocessing the speech signal, the forty-dimensional Mel Frequency Cepstral Coefficients (MFCC) related parameters are extracted, including the dynamic and static features. And these frequency domain features are enhanced by convolutional neural networks (CNNs) as the emotional features of children's speech recognition. BiLSTM is used to solve the problem of poor performance of long-term dependent learning features, and attention mechanism is used for only a few frames contain emotional features in the children speech signal. Compared with the related speech emotion recognition models such as LSTM-CNN and 2D-CNN-LSTM, our proposed speech emotion recognition model improves the accuracy up to 71.6% on the FAU-AIBO children's speech emotion database. Guitao Cao, Yunming Tang, Jiyu Sheng, Wenming Cao 0001 |
BIBM | 4 |
| 2019 | Stacking-based deep neural network for Facial Expression RecognitionabstractWe present a scalable stacking-based deep neural network(S-DNN) for facial expression recognition. The network is a congregate of basic learning models in series to synthesize a deep neural network with feedforward network architecture. Thur, choosing trainable learning modules is the core to effectively build S-DNN in an end-to-end manner. Inspired by the manifold learning archetype, we implement a Patch Discriminative Analysis(PDA) as a basic learning model, followed by hashing and block histogram on the top, which sample image in a low discriminative space, and finding an efficient representation of the training data. As those self-learnable models trained, a low dimensional discriminative feature is implicitly learned, which proves to be useful in facial expression recognition. Experimental results on the facial expression dataset(CK+) show that the proposed model is superior to its counterparts, capable of achieving state-of-the-art performance. Yan Li 0063, Guitao Cao, Wenming Cao 0001 |
BIBM | 3 |
| 2019 | Hybrid representation learning for cross-modal retrieval
Wenming Cao 0001, Qiubin Lin, Zhihai He, Zhiquan He |
Neurocomputing | 1 |
| 2019 | GA-SURF: A new Speeded-Up robust feature extraction algorithm for multispectral images based on geometric algebra
Rui Wang 0034, Yijie Shi, Wenming Cao 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | 3D Convolutional Neural Networks Fusion Model for Lung Nodule Detection onClinical CT Scans
Guitao Cao, Tiantian Huang, Wenming Cao 0001 |
BIBM | 4 |
| 2018 | Sparse Representation for Color Image Based on Geometric AlgebraabstractExisting sparse representation models represent RGB channels separately without thinking about the relationship color channels, which lose some color structures inevitably. In this paper, we introduce a novel sparse representation model for color image based on geometric algebra (GA) theory and its corresponding dictionary learning algorithm, namely K-GASVD is proposed. The model represents the color image as a multivector with the spatial and spectral information in GA space, providing a kind of vectorial representation for the inherent color structures rather than a scalar representation via current sparse image models. The proposed sparse model is validated in the applications of color image denoising and reconstruction. The experimental results demonstrate that our sparse image model avoids the hue bias phenomenon successfully and retained the color structures completely. It shows its potential as a general and powerful tool in various applications of color image analysis. Rui Wang 0034, Miaomiao Shen, Wenming Cao 0001 |
ICME | 4 |
| 2018 | Real-Time Moving Object Segmentation and Classification From HEVC Compressed Surveillance VideoabstractMoving object segmentation and classification from compressed video plays an important role in intelligent video surveillance. Compared with H.264/AVC, High Efficiency Video Coding (HEVC) introduces a host of new coding features that can be further exploited for moving object segmentation and classification. In this paper, we present a real-time approach to segment and classify moving objects using unique features directly extracted from the HEVC compressed domain for video surveillance. In the proposed method, first, motion vector (MV) interpolation for intra-coded prediction unit (PU) and MV outlier removal are employed for preprocessing. Second, blocks with nonzero MVs are clustered into the connected foreground regions using the four-connectivity component labeling algorithm. Third, object region tracking based on temporal consistency is applied to the connected foreground regions to remove the noise regions. The boundary of moving object region is further refined by the coding unit size and PU size. Finally, a person-vehicle classification model using bag of spatial-temporal HEVC syntax words is trained to classify the moving objects, either persons or vehicles. The experimental results demonstrate that the proposed method provides solid performance and can classify moving persons and vehicles accurately. Liang Zhao 0007, Zhihai He, Wenming Cao 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Action recognition based on depth image sequenceabstractHuman action recognition is the process of labeling image sequences with action labels. Robust solutions to this problem have applications in domains such as medical care, human-computer interaction and virtual training. The task is challenging for feature extraction due to variations in motion performance, recording settings and inter-personal differences. To meet these challenges, we propose two types of feature extraction methods based on the Kinect depth image sequences in this paper. One is assuming that there exists even distribute position lines in the three-dimensional space of frame difference, it will be active when the moving object touches them. The other is mapping the 16 successive frame sequences to a single image by Speed Time Mapping (STM) or Time Depth Mapping (STDM), obtaining 36-dimensiona spatial-temporal features in this image. These features are fed into Support Vector Machine (SVM) to identify the action categories. The experiments compare their performance and demonstrate the effectiveness of STDM. Liangcan Liao, Guitao Cao, Wenming Cao 0001 |
BIBM | 3 |
| 2017 | Sparse fast Clifford Fourier transformabstractThe Clifford Fourier transform (CFT) can be applied to both vector and scalar fields. However, due to problems with big data, CFT is not efficient, because the algorithm is calculated in each semaphore. The sparse fast Fourier transform (sFFT) theory deals with the big data problem by using input data selectively. This has inspired us to create a new algorithm called sparse fast CFT (SFCFT), which can greatly improve the computing performance in scalar and vector fields. The experiments are implemented using the scalar field and grayscale and color images, and the results are compared with those using FFT, CFT, and sFFT. The results demonstrate that SFCFT can effectively improve the performance of multivector signal processing. Rui Wang 0034, Yi-xuan Zhou, Yanliang Jin, Wenming Cao 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2016 | Facial expression recognition based on LLENetabstractFacial expression recognition plays an important role in lie detection, and computer-aided diagnosis. Many deep learning facial expression feature extraction methods have a great improvement in recognition accuracy and robutness than traditional feature extraction methods. However, most of current deep learning methods need special parameter tuning and ad hoc fine-tuning tricks. This paper proposes a novel feature extraction model called Locally Linear Embedding Network (LLENet) for facial expression recognition. The proposed LLENet first reconstructs image sets for the cropped images. Unlike previous deep convolutional neural networks that initialized convolutional kernels randomly, we learn multi-stage kernels from reconstructed image sets directly in a supervised way. Also, we create an improved LLE to select kernels, from which we can obtain the most representative feature maps. Furthermore, to better measure the contribution of these kernels, a new distance based on kernel Euclidean is proposed. After the procedure of multi-scale feature analysis, feature representations are finally sent into a linear classifier. Experimental results on facial expression datasets (CK+) show that the proposed model can capture most representative features and thus improves previous results. Dan Meng 0001, Guitao Cao, Zhihai He, Wenming Cao 0001 |
BIBM | 4 |
| 2016 | Automated human physical function measurement using constrained high dispersal network with SVM-linearabstractPhysical measurement have been becoming increasingly helpful in monitoring the humans health status. Manual measurement of physical status is time consuming and may result in misdiagnosing, so an automatic method for identification the status of physical is urgently needed. This paper presents a novel feature extraction method based on using constrained high dispersal network for depth images and coped with Support Vector Machines (SVM) to measure human physical function. The proposed method can catch the most representative features of depth images belonging to different actions and statuses. We analyze the representation efficiency of hand-crafted features (HOG features, and LBP features), deep learning features (CNN features, and PCANet features) and our proposed deep learning features separately in order to validate the efficiency and accuracy of our proposed method. The results show superior performance of 85.19% on 3840 samples (three actions, each with four different statuses, and every status contains sixteen sequences) when the proposed deep features combined with SVM. Dan Meng 0001, Guitao Cao, Weiting Chen, Wenming Cao 0001 |
BIBM | 5 |
| 2016 | Automatic fall detection of human in video using combination of featuresabstractThe problem of automatically fall detection of older people living alone is a popular research topic since falls are one of the major health hazards among the aging population aged 65 and above and the population of them in China is more than 100 million. In this paper, we present an automatic human fall detection framework based on video surveillance which can improve safety of elders in indoor environments. First, a vision component was used to detect and extract moving people in videos from static cameras. Then, we combine Histograms of Oriented Gradients(HOG),Local Binary Pattern(LBP)and feature extracted by the Deep Learning Framework Caffe to form a new augmented feature and the feature is named HLC. We use HLC to represent a person's motion state in a frame of a video sequence. Because the process of fall is a sequence of movements, we use HLC features which were extracted from continuous frames of a video sequence to implement the fall detection. With the help of the HLC feature, we achieve an average fall detection result of 93.7% sensitivity and 92.0% specificity on three different datasets. Guitao Cao, Dan Meng 0001, Weiting Chen, Wenming Cao 0001 |
BIBM | 5 |
| 2016 | Constellational contour parsing for deformable object detection
Tony X. Han, Zhihai He, Wenming Cao 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Task-Driven Progressive Part Localization for Fine-Grained Object RecognitionabstractThe problem of fine-grained object recognition is very challenging due to the subtle visual differences between different object categories. In this paper, we propose a task-driven progressive part localization (TPPL) approach for fine-grained object recognition. Most existing methods follow a two-step approach that first detects salient object parts to suppress the interference from background scenes and then classifies objects based on features extracted from these regions. The part detector and object classifier are often independently designed and trained. In this paper, our major finding is that the part detector should be jointly designed and progressively refined with the object classifier so that the detected regions can provide the most distinctive features for final object recognition. Specifically, we develop a part-based SPP-net (Part-SPP) as our baseline part detector. We then establish a TPPL framework, which takes the predicted boxes of Part-SPP as an initial guess, and then examines new regions in the neighborhood using a particle swarm optimization approach, searching for more discriminative image regions to maximize the objective function and the recognition performance. This procedure is performed in an iterative manner to progressively improve the joint part detection and object classification performance. Experimental results on the Caltech-UCSD-200-2011 dataset demonstrate that our method outperforms state-of-the-art fine-grained categorization methods both in part localization and classification, even without requiring a bounding box during testing. Zhihai He, Guitao Cao, Wenming Cao 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | Animal Detection From Highly Cluttered Natural Scenes Using Spatiotemporal Object Region Proposals and Patch VerificationabstractIn this paper, we consider the animal object detection and segmentation from wildlife monitoring videos captured by motion-triggered cameras, called camera-traps. For these types of videos, existing approaches often suffer from low detection rates due to low contrast between the foreground animals and the cluttered background, as well as high false positive rates due to the dynamic background. To address this issue, we first develop a new approach to generate animal object region proposals using multilevel graph cut in the spatiotemporal domain. We then develop a cross-frame temporal patch verification method to determine if these region proposals are true animals or background patches. We construct an efficient feature description for animal detection using joint deep learning and histogram of oriented gradient features encoded with Fisher vectors. Our extensive experimental results and performance comparisons over a diverse set of challenging camera-trap data demonstrate that the proposed spatiotemporal object proposal and patch verification framework outperforms the state-of-the-art methods, including the recent Faster-RCNN method, on animal object detection accuracy by up to 4.5%. Zhi Zhang 0005, Zhihai He, Guitao Cao, Wenming Cao 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | Content-based image retrieval using high-dimensional information geometry
Wenming Cao 0001, Qicong Kong |
Sci. China Inf. Sci. | 1 |
| 2012 | Multispectral image edge detection via Clifford gradient
Chen Xu 0004, Wenming Cao 0001, Jiqiang Feng |
Sci. China Inf. Sci. | 3 |
| 2009 | Geometry Algebra Neuron Based on Biomimetic Pattern Recognition
Wenming Cao 0001 |
ISNN (4) | 1 |
| 2008 | Coverage analysis for sensor networks based on Clifford algebra
Weixin Xie, Wenming Cao 0001, Shan Meng |
Sci. China Ser. F Inf. Sci. | 2 |
| 2007 | Analysis of Higher Order Voronoi Diagram for Fuzzy Information Coverage
Weixin Xie, Rui Wang 0034, Wenming Cao 0001 |
MSN | 3 |
| 2006 | A Novel Image Restoration Algorithm Based on High-Dimensional Space Geometry
Wenming Cao 0001, Mei-fen Xie, Shoujue Wang |
ICCSA (5) | 1 |
| 2006 | Research on Multi-Degree-of-Freedom Neurons with Weighted Graphs
Shoujue Wang, Singsing Liu, Wenming Cao 0001 |
ISNN (1) | 3 |
| 2005 | Digits Speech Recognition Based on Geometrical Learning
Wenming Cao 0001, Xiaoxia Pan, Shoujue Wang |
ADMA | 1 |
| 2005 | Application of Multi-weighted Neuron for Iris Recognition
Wenming Cao 0001, Jianhui Hu, Gang Xiao 0001, Shoujue Wang |
ISNN (2) | 1 |
| 2005 | Continuous Speech Research Based on Two-Weight Neural Network
Wenming Cao 0001, Xiaoxia Pan, Shoujue Wang |
ISNN (2) | 1 |
| 2004 | Determining of the Delay Time for a Heating Ventilating and Air-Conditioning Plant Using Two Weighted Neural Network Approach
Mengdi Hu, Wenming Cao 0001, Shoujue Wang |
ISNN (2) | 2 |
| 2004 | The application of DBF neural networks for object recognition
Wenming Cao 0001, Shoujue Wang |
Inf. Sci. | 1 |