VLDB 2026 Research / reviewers in the wild / expert
Dewen Hu
dblp:10/6414
· DBLP profile ↗
130ranked-venue papers
2as first author
58since 2021 · last 2026
0000-0001-7357-0053ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 79 · 2 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 11 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 since 2021Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ESTIM: Efficient and Scalable Tensorial Incomplete Multi-view Semi-supervised Classification
Tingjin Luo, XiangYao Li, Zhangqi Jiang, Shuanghui Zhang, Dewen Hu |
KDD (1) | 5 |
| 2026 | Transfer learning from 2D natural images to 4D fMRI brain images via geometric mapping
Kai Gao 0011, Liang Li 0006, Yu-Wei Wang, Xue-Ying Li, Hui-Xian Li, Yi-Fan Liao, Li-Ping Cao, Guan-Mao Chen, Jian-Shan Chen, Tao-Lin Chen, Yan-Rong Chen, Yu-Qi Cheng, Zhao-Song Chu, Shi-Xian Cui, Xi-Long Cui, Zhao-Yu Deng, Qing-Lin Gao, Qi-Yong Gong, Wen-Bin Guo, Can-Can He, Zheng-Jia-Yi Hu, Xin-Lei Ji, Feng-Nan Jia, Li Kuang, Bao-Juan Li, Tao Lian, Xiao-Yun Liu, Yan-Song Liu, Zhe-Ning Liu, Yi-Cheng Long, Jian-Ping Lu, Jiang Qiu, Xiao-Xiao Shan, Tian-Mei Si, Peng-Feng Sun, Chuan-Yue Wang, Han-Lin Wang, Ying Wang 0007, Chen-Nan Wu, Xiao-Ping Wu, Xin-Ran Wu, Yan-Kun Wu, Chun-Ming Xie, Guang-Rong Xie, Xiu-Feng Xu, Zhen-Peng Xue, Jian Yang 0003, Yong-Qiang Yu, Min-Lan Yuan, Yong-Gui Yuan, Ai-Xia Zhang, Ke-Rang Zhang, Wei Zhang 0090, Zi-Jing Zhang, Jing-Ping Zhao, Jia-Jia Zhu, Xi-Nian Zuo, Hua-Ning Wang, Chaogan Yan, Yufeng Zang, Dewen Hu |
Medical Image Anal. | 74 |
| 2026 | EEG-Guided Adversarial Alignment for EOG-Only Vigilance Estimation
Xiangyu Ju, Ming Li 0028, Dewen Hu |
IEEE Signal Process. Lett. | 4 |
| 2026 | Self-Supervised Disentangled Representation Learning via Compositional InvarianceabstractImages serve as a crucial information source for machine intelligence to understand the world, while how to represent images significantly impacting the generalizability and interpretability of intelligent systems. Disentangled representation learning offers a promising approach to improve both aspects. However, most of existing methods predominantly rely on statistical independence assumptions. This poses two key limitations: first, it fails to capture the reality that many concepts are both disentangled yet interrelated; second, it conflicts with human cognitive patterns where concepts naturally exhibit complex dependencies. These limitations further hinder collaboration between machine and human beings. To overcome these limitations, we propose Compositional Invariant Disentanglement (CID), a novel self-supervised learning method that enables models to learn composable representations aligned with human cognitive habits. Inspired by humans’ ability to flexibly recombine concepts, we reframe the definition of disentanglement through the lens of compositional invariance rather than statistical independence. This paradigm shift allows effective disentanglement even with correlated factors, achieving state-of-the-art disentanglement performance across multiple standard benchmarks (improved by 4.0% on Shapes3D, 4.5% on Dsprites, and 28.6% on MPI3D). Furthermore, by building upon and extending the successful self-supervised learning framework BYOL, CID demonstrates potential for large-scale disentanglement pre-training on unlabeled data. This work contributes to extracting more robust and interpretable representations from images for machine intelligence. Haoqiang Chen, Jianxiang Sun, Yadong Liu 0001, Dewen Hu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Decoding Driving Intentions via a Novel Brain-Computer Interface Paradigm With Low Cognitive Load and High Robustness
Jianxiang Sun, Zongtan Zhou, Yadong Liu 0001, Daxue Liu, Haoqiang Chen, Yingxin Liu, Dewen Hu |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2025 | HGSLoc: 3DGS-Based Heuristic Camera Pose RefinementabstractVisual localization refers to the process of determining camera poses and orientation within a known scene representation. This task is often complicated by factors such as changes in illumination and variations in viewing angles. In this paper, we propose HGSLoc, a novel lightweight plug-and-play pose optimization framework, which integrates 3D reconstruction with a heuristic refinement strategy to achieve higher pose estimation accuracy. Specifically, we introduce an explicit geometric map for 3D representation and high-fidelity rendering, allowing the generation of high-quality synthesized views to support accurate visual localization. Our method demonstrates higher localization accuracy compared to NeRFbased neural rendering localization approaches. We introduce a heuristic refinement strategy, its efficient optimization capability can quickly locate the target node, while we set the steplevel optimization step to enhance the pose accuracy in the scenarios with small errors. With carefully designed heuristic functions, it offers efficient optimization capabilities, enabling rapid error reduction in rough localization estimations. Our method mitigates the dependence on complex neural network models while demonstrating improved robustness against noise and higher localization accuracy in challenging environments, as compared to neural network joint optimization strategies. The optimization framework proposed in this paper introduces novel approaches to visual localization by integrating the advantages of 3D reconstruction and the heuristic refinement strategy, which demonstrates strong performance across multiple benchmark datasets, including 7Scenes and Deep Blending dataset. The implementation of our method has been released at https://github.com/anchang699/HGSLoc. Zhongyan Niu, Zhen Tan 0002, Jinpu Zhang, Xueliang Yang, Dewen Hu |
ICRA | 5 |
| 2025 | Tracking Any Point with Frame-Event Fusion Network at High Frame RateabstractTracking any point based on image frames is constrained by frame rates, leading to instability in high-speed scenarios and limited generalization in real-world applications. To overcome these limitations, we propose an image-event fusion point tracker, FE-TAP, which combines the contextual information from image frames with the high temporal resolution of events, achieving high frame rate and robust point tracking under various challenging conditions. Specifically, we designed an Evolution Fusion module (EvoFusion) to model the image generation process guided by events. This module can effectively integrate valuable information from both modalities operating at different frequencies. To achieve smoother point trajectories, we employed a transformer-based refinement strategy that updates the point’s trajectories and features iteratively. Extensive experiments demonstrate that our method outperforms state-of-the-art approaches, particularly improving expected feature age by 24% on EDS datasets. Finally, we qualitatively validated the robustness of our algorithm in real driving scenarios using our custom-designed image-event synchronization device. Jiaxiong Liu, Bo Wang 0144, Zhen Tan 0002, Jinpu Zhang, Hui Shen 0004, Dewen Hu |
IROS | 6 |
| 2025 | ETA: Learning Optical Flow with Efficient Temporal AttentionabstractConsidering the potential of using multi-frame information to solve the occlusion problem, we introduce a novel idea of multi-frame information integration, which uses the attention mechanism to fuse the temporal information from the previous frame. The idea can effectively improve the estimation accuracy in occluded regions and optimize the inference speed under multi-frame settings. Meanwhile, we suggest the concept of attention confidence to provide an explicit value criterion for the model to utilize useful attention information more efficiently. Furthermore, we propose an Efficient Temporal Attention network (ETA), which achieves promising results on Sintel and KITTI benchmarks, especially with a 9.4% error reduction compared to the baseline method GMA on Sintel (test) Clean. Bo Wang 0144, Zhenping Sun, Yang Yu 0014, Li Liu 0002, Jian Li 0003, Dewen Hu |
IROS | 6 |
| 2025 | Fully Spiking Neural Networks for Unified Frame-Event Object TrackingabstractThe integration of image and event streams offers a promising approach for achieving robust visual object tracking in complex environments. However, current fusion methods achieve high performance at the cost of significant computational overhead and struggle to efficiently extract the sparse, asynchronous information from event streams, failing to leverage the energy-efficient advantages of event-driven spiking paradigms. To address this challenge, we propose the first fully Spiking Frame-Event Tracking framework called SpikeFET. This network achieves synergistic integration of convolutional local feature extraction and Transformer-based global modeling within the spiking paradigm, effectively fusing frame and event data. To overcome the degradation of translation invariance caused by convolutional padding, we introduce a Random Patchwork Module (RPM) that eliminates positional bias through randomized spatial reorganization and learnable type encoding while preserving residual structures. Furthermore, we propose a Spatial-Temporal Regularization (STR) strategy that overcomes similarity metric degradation from asymmetric features by enforcing spatio-temporal consistency among temporal template features in latent space. Extensive experiments across multiple benchmarks demonstrate that the proposed framework achieves superior tracking accuracy over existing methods while significantly reducing power consumption, attaining an optimal balance between performance and efficiency. Jingjun Yang, Liangwei Fan, Jinpu Zhang, Xiangkai Lian, Hui Shen 0004, Dewen Hu |
NeurIPS | 6 |
| 2025 | Domain adversarial learning with multiple adversarial tasks for EEG emotion recognition
Xiangyu Ju, Sheng Dai, Ming Li 0028, Dewen Hu |
Expert Syst. Appl. | 5 |
| 2025 | Domain Adversarial Neural Network with Reliable Pseudo-labels Iteration for cross-subject EEG emotion recognitionabstractDomain adaptation (DA) for electroencephalography (EEG) plays an important role in cross-subject emotion recognition. However, traditional DA methods are often limited by target domain complexities, leading to inaccurate knowledge transfer. Recent advances in subdomain adaptation, which focuses on dividing data into subdomains using pseudo-labels, have shown promise, but still rely on the quality of the generated pseudo-labels. To address this issue, we propose a novel approach, a Domain Adversarial Neural Network with Reliable Pseudo-Label Iteration (DANN-RPLI), for cross-subject emotion recognition. This method assumes that high-quality samples are close to the center and stable under perturbations. Thus, we introduced a reliable pseudo-label generation strategy with an iterative process and increased the confidence in the selected labels using perturbations. A domain adversarial network was further used to confuse subdomains, enabling a more effective cross-domain emotion representation. Our method achieved state-of-the-art results on the SEED, SEED-IV, and DEAP datasets. The superior stability of the algorithm was proven through parameter comparison experiments. Furthermore, this study reduces the impact of unreliable pseudo-labels on EEG measurements and provides a new solution for emotion recognition in practical EEG-BCI scenarios. • We propose a domain adversarial neural network with reliable pseudo-label iteration. • A pseudo-label generation strategy was designed to achieve more accurate knowledge transfer. • This method improves the performance in cross-subject EEG emotion recognition. Xiangyu Ju, Jianpo Su, Sheng Dai, Ming Li 0028, Dewen Hu |
Knowl. Based Syst. | 6 |
| 2025 | Adaptive Learning for Dynamic Features and Noisy LabelsabstractApplying current machine learning algorithms in complex and open environments remains challenging, especially when different changing elements are coupled and the training data is scarce. For example, in the activity recognition task, the motion sensors may change position or fall off due to the intensity of the activity, leading to changes in feature space and finally resulting in label noise. Learning from such a problem where the dynamic features are coupled with noisy labels is crucial but rarely studied, particularly when the noisy samples in new feature space are limited. In this paper, we tackle the above problem by proposing a novel two-stage algorithm, called Adaptive Learning for Dynamic features and Noisy labels (ALDN). Specifically, optimal transport is first modified to map the previously learned heterogeneous model to the prior model of the current stage. Then, to fully reuse the mapped prior model, we add a simple yet efficient regularizer as the consistency constraint to assist both the estimation of the noise transition matrix and the model training in the current stage. Finally, two implementations with direct (ALDN-D) and indirect (ALDN-ID) constraints are illustrated for better investigation. More importantly, we provide theoretical guarantees for risk minimization of ALDN-D and ALDN-ID. Extensive experiments validate the effectiveness of the proposed algorithms. Shilin Gu, Chao Xu 0008, Dewen Hu, Chenping Hou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Sample Adaptive Localized Simple Multiple Kernel K-Means and its Application in Parcellation of Human Cerebral CortexabstractSimple multiple kernel k-means (SMKKM) introduces a new minimization-maximization learning paradigm for multi-view clustering and makes remarkable achievements in some applications. As one of its variants, localized SMKKM (LSMKKM) is recently proposed to capture the variation among samples, focusing on reliable pairwise samples, which should keep together and cut off unreliable, farther pairwise ones. Though demonstrating effectiveness, we observe that LSMKKM indiscriminately utilizes the variation of each sample, resulting in unsatisfying clustering performance. To overcome this limitation, we propose a sample adaptive localized SMKKM (SAL-SMKKM) algorithm where the weight of the local alignment for each sample can be adaptively adjusted, resulting in a more challenging tri-level minimization-minimization-maximization. To deal with it, we reformulate it into a minimization problem of an optimal function characterized by minimization-maximization dynamics, prove its differentiability, and develop a reduced gradient descent method to optimize it. We then theoretically analyze the clustering performance of the proposed SAL-SMKKM by deriving its generalization error bound. In addition, we empirically evaluate the clustering performance of the proposed SAL-SMKKM on several benchmark datasets. Experiment results clearly indicate that proposed algorithms consistently outperform state-of-the-art ones. Finally, we apply the proposed SAL-SMKKM to the multi-modal parcellation of the human cerebral cortex, which is essential and helpful to understanding brain organization and function. As seen, SAL-SMKKM achieves accurate parcellation in an automatic and objective manner without any manual intervention, which once again demonstrates its validity and effectiveness in practical applications. Xinwang Liu 0002, Yi Zhang 0104, Li Liu 0002, Chang Tang, Long Lan, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | SceneTracker: Long-Term Scene Flow Estimation NetworkabstractConsidering that scene flow estimation has the capability of the spatial domain to focus but lacks the coherence of the temporal domain, this study proposes long-term scene flow estimation (LSFE), a comprehensive task that can simultaneously capture the fine-grained and long-term 3D motion in an online manner. We introduce SceneTracker, the first LSFE network that adopts an iterative approach to approximate the optimal 3D trajectory. The network dynamically and simultaneously indexes and constructs appearance correlation and depth residual features. Transformers are then employed to explore and utilize long-range connections within and between trajectories. With detailed experiments, SceneTracker shows superior capabilities in addressing 3D spatial occlusion and depth noise interference, highly tailored to the needs of the LSFE task. We build a real-world evaluation dataset, LSFDriving, for the LSFE field and use it in experiments to further demonstrate the advantage of SceneTracker in generalization abilities. Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Li Liu 0002, Zhenping Sun, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Few-Shot Class-Incremental Learning for Classification and Object Detection: A SurveyabstractFew-shot Class-Incremental Learning (FSCIL) presents a unique challenge in Machine Learning (ML), as it necessitates the Incremental Learning (IL) of new classes from sparsely labeled training samples without forgetting previous knowledge. While this field has seen recent progress, it remains an active exploration area. This paper aims to provide a comprehensive and systematic review of FSCIL. In our in-depth examination, we delve into various facets of FSCIL, encompassing the problem definition, the discussion of the primary challenges of unreliable empirical risk minimization and the stability-plasticity dilemma, general schemes, and relevant problems of IL and Few-shot Learning (FSL). Besides, we offer an overview of benchmark datasets and evaluation metrics. Furthermore, we introduce the Few-shot Class-incremental Classification (FSCIC) methods from data-based, structure-based, and optimization-based approaches and the Few-shot Class-incremental Object Detection (FSCIOD) methods from anchor-free and anchor-based approaches. Beyond these, we present several promising research directions within FSCIL that merit further investigation. Li Liu 0002, Olli Silvén, Matti Pietikäinen, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | An Early Warning Approach for Pilots' Cognitive Tipping Points Based Multi-Modal SignalsabstractWhen executing complex missions in emergency scenarios, pilots’ cognitive state may deteriorate, posing significant challenges to flight safety and mission execution. This paper proposes a cross-subject and cross-session early warning approach based on small-sample and multi-modal signals for predicting cognitive collapse state. We designs an experimental paradigm that has been demonstrated to induce cognitive collapse in 87$\%$of trials by analyzing questionnaire scores, physiological signals, and task performance. The extracted multi-modal features are fused and selected to construct the optimal feature set. Further, a two-step early warning method is introduced to identify critical slowing down and to predict tipping points by detecting the high cognitive workload state and classifying the corresponding signals into early warning signal (EWS) and normal state. The early warning of pilot’s cognitive tipping point obtains the result that recall is 77.76$\%$, false alarm rate is 28.43$\%$, and the AUC of the warning method can reach 0.83. Our method shows better early warning performance compared with other classification models and has strong generalizability on other datasets.Note to Practitioners—The pilots’ cognitive state is crucial for mission completion and flight safety. Previous studies mainly focused on assessing current cognitive state, and there is little research on early warning of cognitive collapse that occurs in emergency scenarios (e.g., faults of system and the sudden increase in flight missions). In particular, identification is not effective for small-sample and instable data. Therefore, this paper presents an early warning method for cognitive tipping points, which can provide timely warning of pilots’ cognitive collapse and reduce human-caused risks in emergency scenarios. Yadong Liu 0001, Dewen Hu |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Cognitive Load Prediction From Multimodal Physiological Signals Using Multiview LearningabstractPredicting cognitive load is a crucial issue in the emerging field of human-computer interaction and holds significant practical value, particularly in flight scenarios. Although previous studies have realized efficient cognitive load classification, new research is still needed to adapt the current state-of-the-art multimodal fusion methods. Here, we proposed a feature selection framework based on multiview learning to address the challenges of information redundancy and reveal the common physiological mechanisms underlying cognitive load. Specifically, the multimodal signal features [electroencephalogram (EEG), electrodermal activity (EDA), electrocardiogram (ECG), electrooculogram (EOG), & eye movements] at three cognitive load levels were estimated during multiattribute task battery (MATB) tasks performed by 22 healthy participants and fed into a feature selection-multiview classification with cohesion and diversity (FS-MCCD) framework. The optimized feature set was extracted from the original feature set by integrating the weight of each view and the feature weights to formulate the ranking criteria. The cognitive load prediction model, evaluated using real-time classification results, achieved an average accuracy of 81.08% and an average F1-score of 80.94% for three-class classification among 22 participants. Furthermore, the weights of the physiological signal features revealed the physiological mechanisms related to cognitive load. Specifically, heightened cognitive load was linked to amplified $\delta$ and $\theta$ power in the frontal lobe, reduced $\alpha$ power in the parietal lobe, and an increase in pupil diameter. Thus, the proposed multimodal feature fusion framework emphasizes the effectiveness and efficiency of using these features to predict cognitive load. Yingxin Liu, Yang Yu 0014, Zeqi Ye, Hao Li 0086, Dewen Hu, Zongtan Zhou |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Few-Shot Class-Incremental Learning for Retinal Disease RecognitionabstractFew-Shot Class-Incremental Learning (FSCIL) techniques are essential for developing Deep Learning (DL) models that can continuously learn new classes with limited samples while retaining existing knowledge. This capability is particularly crucial for DL-based retinal disease diagnosis system, where acquiring large annotated datasets is challenging, and disease phenotypes evolve over time. This paper introduces Re-FSCIL, a novel framework for Few-Shot Class-Incremental Retinal Disease Recognition (FSCIRDR). Re-FSCIL integrates the RETFound model with a fine-grained module, employing a forward-compatible training strategy to improve adaptability, supervised contrastive learning to enhance feature discrimination, and feature fusion for robust representation quality. We convert existing datasets into the FSCIL format and reproduce numerous representative FSCIL methods to create two new benchmarks, RFMiD38 and JSIEC39, specifically for FSCIRDR. Our experimental results demonstrate that Re-FSCIL achieves State-of-the-art (SOTA) performance, significantly surpassing existing FSCIL methods on these benchmarks. Yongkun Zhao, Chen Li 0022, Dewen Hu |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Causal Confusion in Pedestrian Crossing Intention Prediction for Autonomous Vehicles: The Role of Ego-Vehicle SpeedabstractPedestrian crossing intention prediction is crucial for autonomous vehicles due to their inherent inertia, yet challenging. The prevailing practice is to leverage multi-modal data that correlate with pedestrian crossing intention as input to infer it, with ego-vehicle speed being a commonly used modality. However, from a causal perspective, we identify two critical issues overlooked by existing methods: 1) The causal relationship between ego-vehicle speed and pedestrian crossing intention is inconsistent across the training and testing phases, leading to a distribution shift in the data; 2) There exists a bidirectional causality between ego-vehicle speed and pedestrian crossing intention, comprising forward causality and anti-causality. The imbalanced distribution of these two causal directions in natural datasets results in causal confusion, further exacerbating the distribution shift. These issues lead to a counter-intuitive hypothesis: removing ego-vehicle speed as input can actually benefit prediction performance. Therefore, we propose LIM (Less Is More), a uni-modal model that utilizes only skeleton sequences as input. LIM features specially designed modules for efficient skeleton sequence processing, eliminating the reliance on unstable ego-vehicle speed. Moreover, LIM employs adversarial training to identify and remove any latent ego-vehicle speed information embedded within the skeleton sequences. LIM achieves competitive prediction performance compared to state-of-the-art multi-modal models while offering superior real-time capabilities and lower computational costs. By revealing critical issues overlooked by most existing work and providing a more robust crossing intention prediction model, our work contributes to the development of safer autonomous vehicles. Haoqiang Chen, Jianxiang Sun, Yadong Liu 0001, Zhiyong Peng 0002, Dewen Hu |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | M₂DC: A Meta-Learning Framework for Generalizable Diagnostic Classification of Major Depressive DisorderabstractPsychiatric diseases are bringing heavy burdens for both individual health and social stability. The accurate and timely diagnosis of the diseases is essential for effective treatment and intervention. Thanks to the rapid development of brain imaging technology and machine learning algorithms, diagnostic classification of psychiatric diseases can be achieved based on brain images. However, due to divergences in scanning machines or parameters, the generalization capability of diagnostic classification models has always been an issue. We propose Meta-learning with Meta batch normalization and Distance Constraint (M2DC) for training diagnostic classification models. The framework can simulate the train-test domain shift situation and promote intra-class cohesion, as well as inter-class separation, which can lead to clearer classification margins and more generalizable models. To better encode dynamic brain graphs, we propose a concatenated spatiotemporal attention graph isomorphism network (CSTAGIN) as the backbone. The network is trained for the diagnostic classification of major depressive disorder (MDD) based on multi-site brain graphs. Extensive experiments on brain images from over 3261 subjects show that models trained by M2DC achieve the best performance on cross-site diagnostic classification tasks compared to various contemporary domain generalization methods and SOTA studies. The proposed M2DC is by far the first framework for multi-source closed-set domain generalizable training of diagnostic classification models for MDD and the trained models can be applied to reliable auxiliary diagnosis on novel data. Jianpo Su, Bo Wang 0144, Hui Shen 0004, Dewen Hu |
IEEE Trans. Medical Imaging | 7 |
| 2025 | A Forward and Backward Compatible Framework for Few-Shot Class-Incremental Pill RecognitionabstractAutomatic pill recognition (APR) systems are crucial for enhancing hospital efficiency, assisting visually impaired individuals, and preventing cross-infection. However, most existing deep learning-based pill recognition systems can only perform classification on classes with sufficient training data. In practice, the high cost of data annotation and the continuous increase in new pill classes necessitate the development of a few-shot class-incremental pill recognition (FSCIPR) system. This article introduces the first FSCIPR framework, discriminative and bidirectional compatible few-shot class-incremental learning (DBC-FSCIL). It encompasses forward-compatible and backward-compatible learning components. In forward-compatible learning, we propose an innovative virtual class generation strategy and a center-triplet (CT) loss to enhance discriminative feature learning. These virtual classes serve as placeholders in the feature space for future class updates, providing diverse semantic knowledge for model training. For backward-compatible learning, we develop a strategy to synthesize reliable pseudo-features of old classes using uncertainty quantification, facilitating data replay (DR) and knowledge distillation (KD). This approach allows for the flexible synthesis of features and effectively reduces additional storage requirements for samples and models. Additionally, we construct a new pill image dataset for FSCIL and assess various mainstream FSCIL methods, establishing new benchmarks. Our experimental results demonstrate that our framework surpasses existing state-of-the-art (SOTA) methods. Li Liu 0002, Kai Gao 0011, Dewen Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | DMAE-EEG: A Pretraining Framework for EEG Spatiotemporal Representation LearningabstractElectroencephalography (EEG) plays a crucial role in neuroscience research and clinical practice, but it remains limited by nonuniform data, noise, and difficulty in labeling. To address these challenges, we develop a pretraining framework named DMAE-EEG, a denoising masked autoencoder for mining generalizable spatiotemporal representation from massive unlabeled EEG. First, we propose a novel brain region topological heterogeneity (BRTH) division method to partition the nonuniform data into fixed patches based on neuroscientific priors. Second, we design a denoised pseudo-label generator (DPLG), which utilizes a denoising reconstruction pretext task to enable the learning of generalizable representations from massive unlabeled EEG, suppressing the influence of noise and artifacts. Furthermore, we utilize an asymmetric autoencoder with self-attention as the backbone in the proposed DMAE-EEG, which captures long-range spatiotemporal dependencies and interactions from unlabeled EEG data across 14 public datasets. The proposed DMAE-EEG is validated on both generative (signal quality enhancement) and discriminative tasks (motion intention recognition). In the quality enhancement, DMAE-EEG outperforms existing statistical methods with normalized mean squared error (nMSE) reduction of 27.78%-50.00% under corruption levels of 25%, 50%, and 75%, respectively. In motion intention recognition, DMAE-EEG achieves a relative improvement of 2.71%-6.14% in intrasession classification balanced accuracy across 2-6 class motor execution and imagery tasks, outperforming state-of-the-art methods. Overall, the results suggest that the pretraining framework DMAE-EEG can capture generalizable spatiotemporal representations from massive unlabeled EEG and enhance the knowledge transferability across sessions, subjects, and tasks in various downstream scenarios, advancing EEG-aided diagnosis and brain-computer communication and control, and other clinical practice. Yang Yu 0014, Hao Li 0086, Anqi Wu, Xin Chen 0106, Jinfang Liu, Dewen Hu |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Grasp Like Humans: Learning Generalizable Multifingered Grasping From Human Proprioceptive Sensorimotor Integration
Ce Guo 0004, Xieyuanli Chen, Zirui Guo, Haoran Xiao, Dewen Hu, Huimin Lu 0002 |
IEEE Trans. Robotics | 7 |
| 2025 | Toward Scalable Multirobot Control: Fast Policy Learning in Distributed MPCabstractDistributed model predictive control (DMPC) is promising in achieving optimal cooperative control in multirobot systems (MRS). However, real-time DMPC implementation relies on numerical optimization tools to periodically calculate local control sequences online. This process is computationally demanding and lacks scalability for large-scale, nonlinear MRS. This article proposes a novel distributed learning-based predictive control framework for scalable multirobot control. Unlike conventional DMPC methods that calculate open-loop control sequences, our approach centers around a computationally fast and efficient distributed policy learning algorithm that generates explicit closed-loop DMPC policies for MRS without using numerical solvers. The policy learning is executed incrementally and forward in time in each prediction interval through an online distributed actor–critic implementation. The control policies are successively updated in a receding-horizon manner, enabling fast and efficient policy learning with the closed-loop stability guarantee. The learned control policies could be deployed online to MRS with varying robot scales, enhancing scalability and transferability for large-scale MRS. Furthermore, we extend our methodology to address the multirobot safe learning challenge through a force field-inspired policy learning approach. We validate our approach's effectiveness, scalability, and efficiency through extensive experiments on cooperative tasks of large-scale wheeled robots and multirotor drones. Our results demonstrate the rapid learning and deployment of DMPC policies for MRS with scales up to 10 000 units. Wei Pan 0004, Cong Li 0015, Xin Xu 0001, Xiangke Wang, Dewen Hu |
IEEE Trans. Robotics | 7 |
| 2024 | TD-NeRF: Novel Truncated Depth Prior for Joint Camera Pose and Neural Radiance Field OptimizationabstractThe reliance on accurate camera poses is a significant barrier to the widespread deployment of Neural Radiance Fields (NeRF) models for 3D reconstruction and SLAM tasks. The existing method introduces monocular depth priors to jointly optimize the camera poses and NeRF, which fails to fully exploit the depth priors and neglects the impact of their inherent noise. In this paper, we propose Truncated Depth NeRF (TD-NeRF), a novel approach that enables training NeRF from unknown camera poses - by jointly optimizing learnable parameters of the radiance field and camera poses. Our approach explicitly utilizes monocular depth priors through three key advancements: 1) we propose a novel depth-based ray sampling strategy based on the truncated normal distribution, which improves the convergence speed and accuracy of pose estimation; 2) to circumvent local minima and refine depth geometry, we introduce a coarse-to-fine training strategy that progressively improves the depth precision; 3) we propose a more robust inter-frame point constraint that enhances robustness against depth noise during training. The experimental results on three datasets demonstrate that TD-NeRF achieves superior performance in the joint optimization of camera pose and NeRF, surpassing prior works, and generates more accurate depth geometry. The implementation of our method has been released at https://github.com/nubot-nudt/TD-NeRF. Zhen Tan 0002, Zongtan Zhou, Yangbing Ge, Xieyuanli Chen, Dewen Hu |
IROS | 6 |
| 2024 | SplatFlow: Learning Multi-frame Optical Flow via Splatting
Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Zhenping Sun, Li Liu 0002, Dewen Hu |
Int. J. Comput. Vis. | 7 |
| 2024 | Few-Shot Domain-Adaptive Anomaly Detection for Cross-Site Brain ImagesabstractEarly screening is essential for effective intervention and treatment of individuals with mental disorders. Functional magnetic resonance imaging (fMRI) is a noninvasive tool for depicting neural activity and has demonstrated strong potential as a technique for identifying mental disorders. Due to the difficulty in data collection and diagnosis, imaging data from patients are rare at a single site, whereas abundant healthy control data are available from public datasets. However, joint use of these data from multiple sites for classification model training is hindered by cross-domain distribution discrepancy and diverse label spaces. Herein, we propose few-shot domain-adaptive anomaly detection (FAAD) to achieve cross-site anomaly detection of brain images based on only a few labeled samples. We introduce domain adaptation to mitigate cross-domain distribution discrepancy and jointly align the general and conditional feature distributions of imaging data across multiple sites. We utilize fMRI data of healthy subjects in the Human Connectome Project (HCP) as the source domain and fMRI images from six independent sites, including patients with mental disorders and demographically matched healthy controls, as target domains. Experiments showed the superiority of the proposed method compared with binary classification, traditional anomaly detection methods, and several recognized domain adaptation methods. Jianpo Su, Hui Shen 0004, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Gradient Matching Federated Domain Adaptation for Brain Image ClassificationabstractFederated learning has shown its unique advantages in many different tasks, including brain image analysis. It provides a new way to train deep learning models while protecting the privacy of medical image data from multiple sites. However, previous studies suggest that domain shift across different sites may influence the performance of federated models. As a solution, we propose a gradient matching federated domain adaptation (GM-FedDA) method for brain image classification, aiming to reduce domain discrepancy with the assistance of a public image dataset and train robust local federated models for target sites. It mainly includes two stages: 1) pretraining stage; we propose a one-common-source adversarial domain adaptation (OCS-ADA) strategy, i.e., adopting ADA with gradient matching loss to pretrain encoders for reducing domain shift at each target site (private data) with the assistance of a common source domain (public data) and 2) fine-tuning stage; we develop a gradient matching federated (GM-Fed) fine-tuning method for updating local federated models pretrained with the OCS-ADA strategy, i.e., pushing the optimization direction of a local federated model toward its specific local minimum by minimizing gradient matching loss between sites. Using fully connected networks as local models, we validate our method with the diagnostic classification tasks of schizophrenia and major depressive disorder based on multisite resting-state functional MRI (fMRI), respectively. Results show that the proposed GM-FedDA method outperforms other commonly used methods, suggesting the potential of our method in brain imaging analysis and other fields, which need to utilize multisite data while preserving data privacy. Jianpo Su, Min Gan, Hui Shen 0004, Dewen Hu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | PIPO-SLAM: Lightweight Visual-Inertial SLAM With Preintegration Merging Theory and Pose-Only Descriptions of Multiple View GeometryabstractOptimization-based VI-SLAM focuses on the establishment of the loss function using both inertial and visual constraints. Preintegration theory is commonly used to express inertial constraints, but it lacks the merging equation between keyframes, challenging VI-SLAM from culling and merging redundant keyframes. To address this, we establish an on-manifold preintegration merging theory, including the merging of preintegrated terms, noise covariance, and Jacobians for bias updating, which significantly improves the preintegration theory and provides theoretical support for the keyframe management function of VI-SLAM. Visual constraints are typically expressed using multiple view geometry with three-dimensional (3D) points optimized as scene structure parameters. However, the excessive dimensionality of the optimization parameters generated by 3D points can lead to computational bottlenecks. Through the recent pose-only imaging geometry representation, we construct a lightweight optimization algorithm for SLAM that avoids the dimensional explosion in bundle adjustment (BA). Based on the above, we propose a 3D points-free SLAM optimizer. The proposed algorithms are validated on simulation, public datasets, and real-world experiments, and compared against advanced open-source systems such as ORB-SLAM3 and VINS. Yangbing Ge, Lilian Zhang, Yuanxin Wu, Dewen Hu |
IEEE Trans. Robotics | 4 |
| 2023 | Neural Video: A Novel Framework for Interpreting the Spatiotemporal Activities of the Human Brain
Jingrui Xu, Jianpo Su, Kai Gao 0011, Ming Zhang 0027, Dewen Hu |
ICIG (5) | 6 |
| 2023 | Label Distribution Changing Learning with Sample Space ExpandingabstractWith the evolution of data collection ways, label ambiguity has arisen from various applications. How to reduce its uncertainty and leverage its effectiveness is still a challenging task. As two types of representative label ambiguities, Label Distribution Learning (LDL), which annotates each instance with a label distribution, and Emerging New Class (ENC), which focuses on model reusing with new classes, have attached extensive attentions. Nevertheless, in many applications, such as emotion distribution recognition and facial age estimation, we may face a more complicated label ambiguity scenario, i.e., label distribution changing with sample space expanding owing to the new class. To solve this crucial but rarely studied problem, we propose a new framework named as Label Distribution Changing Learning (LDCL) in this paper, together with its theoretical guarantee with generalization error bound. Our approach expands the sample space by re-scaling previous distribution and then estimates the emerging label value via scaling constraint factor. For demonstration, we present two special cases within the framework, together with their optimizations and convergence analyses. Besides evaluating LDCL on most of the existing 13 data sets, we also apply it in the application of emotion distribution recognition. Experimental results demonstrate the effectiveness of our approach in both tackling label ambiguity problem and estimating facial emotion Chao Xu 0008, Jing Zhang 0064, Dewen Hu, Chenping Hou |
J. Mach. Learn. Res. | 4 |
| 2023 | A Pose-Only Solution to Visual Reconstruction and NavigationabstractVisual navigation and three-dimensional (3D) scene reconstruction are essential for robotics to interact with the surrounding environment. Large-scale scenarios and computational robustness are great challenges facing the research community to achieve this goal. This paper raises a pose-only imaging geometry representation and algorithms that might help solve these challenges. The pose-only representation, equivalent to the classical multiple-view geometry, is discovered to be linearly related to camera global translations, which allows for efficient and robust camera motion estimation. As a result, the spatial feature coordinates can be analytically reconstructed and do not require nonlinear optimization. Comprehensive experiments demonstrate that the computational efficiency of recovering the scene and associated camera poses is significantly improved by 2-4 orders of magnitude. Lilian Zhang, Yuanxin Wu, Wenxian Yu, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Adaptive Feature Selection With Augmented AttributesabstractIn many dynamic environment applications, with the evolution of data collection ways, the data attributes are incremental and the samples are stored with accumulated feature spaces gradually. For instance, in the neuroimaging-based diagnosis of neuropsychiatric disorders, with emerging of diverse testing ways, we get more brain image features over time. The accumulation of different types of features will unavoidably bring difficulties in manipulating the high-dimensional data. It is challenging to design an algorithm to select valuable features in this feature incremental scenario. To address this important but rarely studied problem, we propose a novel Adaptive Feature Selection method (AFS). It enables the reusability of the feature selection model trained on previous features and adapts it to fit the feature selection requirements on all features automatically. Besides, an ideal$\ell _{0}$-norm sparse constraint for feature selection is imposed with a proposed effective solving strategy. We present the theoretical analyses about the generalization bound and convergence behavior. After tackling this problem in a one-shot case, we extend it to the multi-shot scenario. Plenty of experimental results demonstrate the effectiveness of reusing previous features and the superior of$\ell _{0}$-norm constraint in various aspects, together with its effectiveness in discriminating schizophrenic patients from healthy controls. Chenping Hou, Ruidong Fan, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | GeoTransformer: Fast and Robust Point Cloud Registration With Geometric TransformerabstractWe study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by 18 ∼ 31 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Slobodan Ilic, Dewen Hu, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | RayMVSNet++: Learning Ray-Based 1D Implicit Fields for Accurate Multi-View StereoabstractLearning-based multi-view stereo (MVS) has by far centered around 3D convolution on cost volumes. Due to the high computation and memory consumption of 3D CNN, the resolution of output depth is often considerably limited. Different from most existing works dedicated to adaptive refinement of cost volumes, we opt to directly optimize the depth value along each camera ray, mimicking the range (depth) finding of a laser scanner. This reduces the MVS problem to ray-based depth optimization which is much more light-weight than full cost volume optimization. In particular, we propose RayMVSNet which learns sequential prediction of a 1D implicit field along each camera ray with the zero-crossing point indicating scene depth. This sequential modeling, conducted based on transformer features, essentially learns the epipolar line search in traditional multi-view stereo. We devise a multi-task learning for better optimization convergence and depth accuracy. We found the monotonicity property of the SDFs along each ray greatly benefits the depth estimation. Our method ranks top on both the DTU and the Tanks & Temples datasets over all previous learning-based methods, achieving an overall reconstruction score of 0.33 mm on DTU and an F-score of 59.48% on Tanks & Temples. It is able to produce high-quality depth estimation and point cloud reconstruction in challenging scenarios such as objects/scenes with non-textured surface, severe occlusion, and highly varying depth range. Further, we propose RayMVSNet++ to enhance contextual feature aggregation for each ray through designing an attentional gating unit to select semantically relevant neighboring rays within the local frustum around that ray. This improves the performance on datasets with more challenging examples (e.g., low-quality images caused by poor lighting conditions or motion blur). RayMVSNet++ achieves state-of-the-art performance on the ScanNet dataset. In particular, it attains an AbsRel of 0.058m and produces accurate results on the two subsets of textureless regions and large depth variation. Junhua Xi, Dewen Hu, Zhiping Cai, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Learning to Detect 3D Symmetry From Single-View RGB-D Images With Weak Supervisionabstract3D symmetry detection is a fundamental problem in computer vision and graphics. Most prior works detect symmetry when the object model is fully known, few studies symmetry detection on objects with partial observation, such as single RGB-D images. Recent work addresses the problem of detecting symmetries from incomplete data with a deep neural network by leveraging the dense and accurate symmetry annotations. However, due to the tedious labeling process, full symmetry annotations are not always practically available. In this work, we present a 3D symmetry detection approach to detect symmetry from single-view RGB-D images without using symmetry supervision. The key idea is to train the network in a weakly-supervised learning manner to complete the shape based on the predicted symmetry such that the completed shape be similar to existing plausible shapes. To achieve this, we first propose a discriminative variational autoencoder to learn the shape prior in order to determine whether a 3D shape is plausible or not. Based on the learned shape prior, a symmetry detection network is present to predict symmetries that produce shapes with high shape plausibility when completed based on those symmetries. Moreover, to facilitate end-to-end network training and multiple symmetry detection, we introduce a new symmetry parametrization for the learning-based symmetry estimation of both reflectional and rotational symmetry. The proposed approach, coupled symmetry detection with shape completion, essentially learns the symmetry-aware shape prior, facilitating more accurate and robust symmetry detection. Experiments demonstrate that the proposed method is capable of detecting reflectional and rotational symmetries accurately, and shows good generality in challenging scenarios, such as objects with heavy occlusion and scanning noise. Moreover, it achieves state-of-the-art performance, improving the F1-score over the existing supervised learning method by 2%-11% on the ShapeNet and ScanNet datasets. Xin Xu 0001, Junhua Xi, Xiaochang Hu, Dewen Hu, Kai Xu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | SS-TBN: A Semi-Supervised Tri-Branch Network for COVID-19 Screening and Lesion SegmentationabstractInsufficient annotated data and minor lung lesions pose big challenges for computed tomography (CT)-aided automatic COVID-19 diagnosis at an early outbreak stage. To address this issue, we propose a Semi-Supervised Tri-Branch Network (SS-TBN). First, we develop a joint TBN model for dual-task application scenarios of image segmentation and classification such as CT-based COVID-19 diagnosis, in which pixel-level lesion segmentation and slice-level infection classification branches are simultaneously trained via lesion attention, and individual-level diagnosis branch aggregates slice-level outputs for COVID-19 screening. Second, we propose a novel hybrid semi-supervised learning method to make full use of unlabeled data, combining a new double-threshold pseudo labeling method specifically designed to the joint model and a new inter-slice consistency regularization method specifically tailored to CT images. Besides two publicly available external datasets, we collect internal and our own external datasets including 210,395 images (1,420 cases versus 498 controls) from ten hospitals. Experimental results show that the proposed method achieves state-of-the-art performance in COVID-19 classification with limited annotated data even if lesions are subtle, and that segmentation results promote interpretability for diagnosis, suggesting the potential of the SS-TBN in early screening in insufficient labeled data situations at the early stage of a pandemic outbreak like COVID-19. Kai Gao 0011, Dewen Hu, Zhichao Feng, Chenping Hou, Pengfei Rong, Wei Wang 0434 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Few-Shot Class-Incremental SAR Target Recognition via Cosine Prototype LearningabstractRecent years have witnessed a remarkable breakthrough in Synthetic Aperture Radar Automatic Target Recognition (SAR ATR) with the development of deep learning (DL). Nonetheless, once deployed, the DL-based methods’ ability to incrementally learn new knowledge from few-shot samples without forgetting the old is fragile, hindering them from discriminating unseen targets in real-world situations. In this paper, we propose a Cosine Prototype Learning (CPL) framework to first unlock few-shot class-incremental learning (FSCIL) in the SAR ATR field inspired by the intrinsic relationships between target azimuth-aware knowledge and semantic features under the cosine criterion. By condensing class-specific characteristics into individual prototypes, stable profiles of targets are depicted without losing generalization. For the model’s plasticity, a pairwise structure separation (PSS) loss is introduced to separate old and new classes and compact intra-class features. Meanwhile, the model’s transferability on new classes is guaranteed by a prototype consistency (PC) loss. For the model’s stability, we propose a prototype-exemplar distillation (PED) loss and a prototype re-calibration (PR) strategy to penalize semantic drifts of old-class feature spaces and alleviate the misalignment of the learned prototypes successively. At inference, a nearest-class-mean (NCM) classifier is adopted for evaluation by comparing cosine similarity scores between testing samples and class-specific prototypes. In experiments, the proposed components of our method are explored by ablation studies. Strong baselines are established, and extensive experiments conducted on the MSTAR dataset show that our method outperforms state-of-the-art methods under various FSCIL conditions, verifying its effectiveness for the FSCIL of SAR ATR. Yan Zhao 0026, Lingjun Zhao, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Multi-Brain Coding Expands the Instruction Set in SSVEP-Based Brain-Computer InterfacesabstractPrevious studies have made great efforts to expand the instruction set in steady-state visual evoked potential (SSVEP)-based brain-computer interfaces. However, most systems are limited to single persons and expand the instruction set by increasing the flicker stimulation frequency range or via multiple frequencies sequential coding or joint frequency/phase coding. In this article, we propose a multibrain coding SSVEP paradigm that encodes the SSVEP instructions generated byNindependent subjects withMflicker stimuli, thus increasing the instruction set toMNinstructions. A total of 40 subjects participated in online experiments in this article. The results show that there is no significant difference in accuracy (p> 0.05, paired t test) between the multibrain and single-brain coding systems, while the information transfer rate (ITR) increases significantly (pN= 3 andM= 5. In summary, the proposed multibrain coding paradigm exponentially increases the number of instructions without increasing the output instruction time, which is of great significance to the application and promotion of BCIs. Xingxing Chu, Yang Yu 0014, Zeqi Ye, Dewen Hu |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2023 | Fusion of Spatial, Temporal, and Spectral EEG Signatures Improves Multilevel Cognitive Load PredictionabstractCognitive load prediction is one of the most important issues in the nascent field of neuroergonomics, and it has significant value in real-world applications. Most of the previous studies of cognitive load prediction only utilized electroencephalography (EEG)-based spectral signatures or interchannel connectivity, ignoring abundant temporal microstate features, which may represent the transient topologies of EEG signals. Furthermore, previous studies have mostly focused on the binary-level classification of cognitive load for single-type cognitive tasks. To date, there are few studies on the multilevel prediction of cognitive load during mixed cognitive tasks. Here, we first designed a new paradigm termed the “finding fault game,” mixing multiple tasks of memory, counting, and visual search, and then developed a multidimensional analysis framework to improve cognitive load prediction using a fusion of spatial, temporal, and spectral EEG features. Specifically, EEG-based functional connectivity, microstates and power spectral densities (PSD) were calculated for three cognitive load levels. Twelve adult subjects participated in the study. The experimental results show that increased cognitive load was associated with elevated theta and degraded alpha power and significant changes in interchannel connectivity and microstates, and that fusing the three types of EEG features improved the performance of three-level cognitive load prediction, achieving the accuracies of greater than 80% in the cross-validation, real-time, and over-time prediction. The findings suggest that all three types of EEG features can serve as signatures of cognitive load and that their fusion can improve multilevel prediction. Yingxin Liu, Yang Yu 0014, Zeqi Ye, Ming Li 0028, Zongtan Zhou, Dewen Hu |
IEEE Trans. Hum. Mach. Syst. | 7 |
| 2023 | Incomplete Multi-View Learning Under Label ShiftabstractIn image processing, images are usually composed of partial views due to the uncertainty of collection and how to efficiently process these images, which is called incomplete multi-view learning, has attracted widespread attention. The incompleteness and diversity of multi-view data enlarges the difficulty of annotation, resulting in the divergence of label distribution between the training and testing data, named as label shift. However, existing incomplete multi-view methods generally assume that the label distribution is consistent and rarely consider the label shift scenario. To address this new but important challenge, we propose a novel framework termed as Incomplete Multi-view Learning under Label Shift (IMLLS). In this framework, we first give the formal definitions of IMLLS and the bidirectional complete representation which describes the intrinsic and common structure. Then, a multilayer perceptron which combines the reconstruction and classification loss is employed to learn the latent representation, whose existence, consistency and universality are proved with the theoretical satisfaction of label shift assumption. After that, to align the label distribution, the learned representation and trained source classifier are used to estimate the importance weight by designing a new estimation scheme which balances the error generated by finite samples in theory. Finally, the trained classifier reweighted by the estimated weight is fine-tuned to reduce the gap between the source and target representations. Extensive experimental results validate the effectiveness of our algorithm over existing state-of-the-arts methods in various aspects, together with its effectiveness in discriminating schizophrenic patients from healthy controls. Ruidong Fan, Xiao Ouyang, Tingjin Luo, Dewen Hu, Chenping Hou |
IEEE Trans. Image Process. | 4 |
| 2023 | Semi-Supervised Learning With Label ProportionabstractThe scarcity of labels is common and great challenge in traditional supervised learning. Semi-supervised learning (SSL) leverages unlabeled samples to alleviate the absence of label information. Similar with annotation, label proportion is another type of prior information and plays a significant role in classification tasks. Compared with the acquisition of labels, label proportion can be obtained more easily. For example, only a small number of patients have been diagnosed with or not with cancers in hospital database, while the proportion with cancer can be generally estimated by historical records. How to incorporate such prior information of label proportion is crucial but rarely studied in literature. Traditional SSL methods often ignore this prior information and will lead to performance degradation inevitably. To solve this problem, we propose a novel SSL with Label Proportion (SSLLP). Our approach encourages to preserve label consistency and label proportion by imposing the cardinality bound constraints. Our formulated problem equals to a mixed-integer constrained submodular minimization and it is difficult to be solved directly. Therefore, we transformed the original problem into a convex one by Lov$\acute{\text{a}}$sz extension and designed an efficient solving algorithm. Extensive experimental results present the improved performance of our method over several state-of-the-art methods. Ningzhao Sun, Tingjin Luo, Wenzhang Zhuge, Chenping Hou, Dewen Hu |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Facial Kinship Verification: A Comprehensive Review and OutlookabstractThe goal of Facial Kinship Verification (FKV) is to automatically determine whether two individuals have a kin relationship or not from their given facial images or videos. It is an emerging and challenging problem that has attracted increasing attention due to its practical applications. Over the past decade, significant progress has been achieved in this new field. Handcrafted features and deep learning techniques have been widely studied in FKV. The goal of this paper is to conduct a comprehensive review of the problem of FKV. We cover different aspects of the research, including problem definition, challenges, applications, benchmark datasets, a taxonomy of existing methods, and state-of-the-art performance. In retrospect of what has been achieved so far, we identify gaps in current research and discuss potential future research directions. Xiaoting Wu, Xiaoyi Feng, Xiaochun Cao, Xin Xu 0001, Dewen Hu, Miguel Bordallo López, Li Liu 0002 |
Int. J. Comput. Vis. | 5 |
| 2022 | Coarse-to-fine pseudo supervision guided meta-task optimization for few-shot object classification
Yawen Cui, Qing Liao 0001, Dewen Hu, Wei An 0003, Li Liu 0002 |
Pattern Recognit. | 3 |
| 2022 | Deep Ladder-Suppression Network for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. Most existing approaches learn domain-invariant features by adapting the entire information of the images. However, forcing adaptation of domain-specific variations undermines the effectiveness of the learned features. To address this problem, we propose a novel, yet elegant module, called the deep ladder-suppression network (DLSN), which is designed to better learn the cross-domain shared content by suppressing domain-specific variations. Our proposed DLSN is an autoencoder with lateral connections from the encoder to the decoder. By this design, the domain-specific details, which are only necessary for reconstructing the unlabeled target data, are directly fed to the decoder to complete the reconstruction task, relieving the pressure of learning domain-specific variations at the later layers of the shared encoder. As a result, DLSN allows the shared encoder to focus on learning cross-domain shared content and ignores the domain-specific variations. Notably, the proposed DLSN can be used as a standard module to be integrated with various existing UDA frameworks to further boost performance. Without whistles and bells, extensive experimental results on four gold-standard domain adaptation datasets, for example: 1) Digits; 2) Office31; 3) Office-Home; and 4) VisDA-C, demonstrate that the proposed DLSN can consistently and significantly improve the performance of various popular UDA frameworks. Wanxia Deng, Lingjun Zhao, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Cybern. | 4 |
| 2022 | Self-Paced Dynamic Infinite Mixture Model for Fatigue Evaluation of Pilots' BrainsabstractCurrent brain cognitive models are insufficient in handling outliers and dynamics of electroencephalogram (EEG) signals. This article presents a novel self-paced dynamic infinite mixture model to infer the dynamics of EEG fatigue signals. The instantaneous spectrum features provided by ensemble wavelet transform and Hilbert transform are extracted to form four fatigue indicators. The covariance of log likelihood of the complete data is proposed to accurately identify similar components and dynamics of the developed mixture model. Compared with its seven peers, the proposed model shows better performance in automatically identifying a pilot's brain workload. Qi Wu 0003, MengChu Zhou, Dewen Hu, Longjun Zhu, Xu-Yi Qiu, Ping-Yu Deng, Limin Zhu 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | Attentional Feature Refinement and Alignment Network for Aircraft Detection in SAR ImageryabstractAircraft detection in synthetic aperture radar (SAR) imagery is a challenging task in SAR automatic target recognition (SAR ATR) areas due to aircraft’s extremely discrete appearance, obvious intraclass variation, small size, and serious background’s interference. In this article, a single shot detector (SSD), namely, attentional feature refinement and alignment network (AFRAN), is proposed for detecting aircraft in SAR images with competitive accuracy and speed. Specifically, three significant components, including attention feature fusion module (AFFM), deformable lateral connection module (DLCM), and anchor-guided detection module (ADM), are carefully designed in our method for refining and aligning informative characteristics of aircraft. To represent the characteristics of aircraft with less interference, low-level textural and high-level semantic features of aircraft are fused and refined in AFFM thoroughly. The alignment between aircraft’s discrete backscatting points and convolutional sampling spots is promoted in DLCM. Eventually, the locations of aircraft are predicted precisely in ADM based on aligned features revised by refined anchors. To evaluate the performance of our method, a self-built SAR aircraft sliced dataset and a large scene SAR image are collected. Extensive quantitative and qualitative experiments with detailed analysis illustrate the effectiveness of the three proposed components. Furthermore, the topmost detection accuracy and competitive speed are achieved by our method compared with other domain-specific methods, e.g., dense attention pyramid network (DAPN) and pyramid attention dilated network (PADN), and general convolutional neural network (CNN)-based methods, e.g., Feature Pyramid Network (FPN), Cascade R-CNN, SSD, RefineDet, and RepPoints Detector (RPDet). Yan Zhao 0026, Lingjun Zhao, Zhong Liu 0002, Dewen Hu, Gangyao Kuang, Li Liu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Informative Feature Disentanglement for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims at learning a classifier for an unlabeled target domain by transferring knowledge from a labeled source domain with a related but different distribution. The strategy of aligning the two domains in latent feature space via metric discrepancy or adversarial learning has achieved considerable progress. However, these existing approaches mainly focus on adapting the entire image and ignore the bottleneck that occurs when forced adaptation of uninformative domain-specific variations undermines the effectiveness of learned features. To address this problem, we propose a novel component called Informative Feature Disentanglement (IFD), which is equipped with the adversarial network or the metric discrepancy model, respectively. Accordingly, the new network architectures, named IFDAN and IFDMN, enable informative feature refinement before the adaptation. The proposed IFD is designed to disentangle informative features from the uninformative domain-specific variations, which are produced by a Variational Autoencoder (VAE) with lateral connections from the encoder to the decoder. We cooperatively apply the IFD to conduct supervised disentanglement for the source domain and unsupervised disentanglement for the target domain. In this way, informative features are disentangled from the domain-specific details before the adaptation. Extensive experimental results on three gold-standard domain adaptation datasets, e.g., Office31, Office-Home and VisDA-C, demonstrate the effectiveness of the proposed IFDAN and IFDMN models for UDA. Wanxia Deng, Lingjun Zhao, Qing Liao 0001, Deke Guo, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
IEEE Trans. Multim. | 6 |
| 2021 | Pixel Difference Networks for Efficient Edge DetectionabstractRecently, deep Convolutional Neural Networks (CNNs) can achieve human-level performance in edge detection with the rich and abstract edge representation capacities. However, the high performance of CNN based edge detection is achieved with a large pretrained CNN backbone, which is memory and energy consuming. In addition, it is surprising that the previous wisdom from the traditional edge detectors, such as Canny, Sobel, and LBP are rarely investigated in the rapid-developing deep learning era. To address these issues, we propose a simple, lightweight yet effective architecture named Pixel Difference Network (PiDiNet) for efficient edge detection. PiDiNet adopts novel pixel difference convolutions that integrate the traditional edge detection operators into the popular convolutional operations in modern CNNs for enhanced performance on the task, which enjoys the best of both worlds. Extensive experiments on BSDS500, NYUD, and Multicue are provided to demonstrate its effectiveness, and its high training and inference efficiency. Surprisingly, when training from scratch with only the BSDS500 and VOC datasets, PiDiNet can surpass the recorded result of human perception (0.807 vs. 0.803 in ODS F-measure) on the BSDS500 dataset with 100 FPS and less than 1M parameters. A faster version of PiDiNet with less than 0.1M parameters can still achieve comparable performance among state of the arts with 200 FPS. Results on the NYUD and Multicue datasets show similar observations. The codes are available at https://github.com/zhuoinoulu/pidinet. Zhuo Su 0002, Zitong Yu, Dewen Hu, Qing Liao 0001, Qi Tian 0001, Matti Pietikäinen, Li Liu 0002 |
ICCV | 4 |
| 2021 | Transferable Discriminative Feature Mining For Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to seek an effective model for unlabeled target domain by leveraging knowledge from a labeled source domain with a related but different distribution. Many existing approaches ignore the underlying discriminative features of the target data and the discrepancy of conditional distributions. To address these two issues simultaneously, the paper presents a Transferable Discriminative Feature Mining (TDFM) approach for UDA, which can naturally unify the mining of domain-invariant discriminative features and the alignment of class-wise features into one single framework. To be specific, to achieve the domain-invariant discriminative features, TDFM jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data. It then conducts the class-wise alignment by decreasing intra-class variations and increasing inter-class differences across domains, encouraging the emergence of transferable discriminative features. When combined, these two procedures are mutually beneficial. Comprehensive experiments verify that TDFM can obtain remarkable margins over state-of-the-art domain adaptation methods. Lingjun Zhao, Wanxia Deng, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
ICIP | 4 |
| 2021 | A Federated Deep Learning Framework for 3D Brain MRI ImagesabstractDeep learning has shown unique advantages in multiple applications. However, medical image data such as 3D brain MRI scans contain private information of subjects and cannot be shared and used without special permits. This makes it difficult for aggregated analysis of medical images from multiple sites. Federated learning provides an approach to aggregate learned features from different sites without transferring raw data, ensuring the security of subject information. Traditional federated learning algorithms such as FEDAVG has been proved to be effective, but the performance is always poor due to high data heterogeneity across sites. Moreover, federated learning for 3D medical images has rarely been explored. Here we proposed a Guide-Weighted Federated Deep Learning framework for multisite 3D brain MRI images. We used weighted aggregation on gradients and then used the guide gradients to update federated models to deal with cross-site heterogeneity. The framework was validated on the publicly accessible multi-site brain MRI database of Autism Spectrum Disorder (ASD) - ABIDE I&II. Results showed that federated models outperform local models with improvements in accuracy of 0.92-4.02%. As the first federated learning framework for 3D MRI data, the proposed framework is expected to facilitate imaging data sharing across sites and accelerate the Medical Imaging Alliance's formation. In this way, a robust federated model with better performance can then be generated for future diagnostic classification tasks and assisted in the diagnosing and screening of neuropsychiatric disorders. Jianpo Su, Kai Gao 0011, Dewen Hu |
IJCNN | 4 |
| 2021 | Informative Class-Conditioned Feature Alignment for Unsupervised Domain AdaptationabstractThe goal of unsupervised domain adaptation is to learn a task classifier that performs well for the unlabeled target domain by borrowing rich knowledge from a well-labeled source domain. Although remarkable breakthroughs have been achieved in learning transferable representation across domains, two bottlenecks remain to be further explored. First, many existing approaches focus primarily on the adaptation of the entire image, ignoring the limitation that not all features are transferable and informative for the object classification task. Second, the features of the two domains are typically aligned without considering the class labels; this can lead the resulting representations to be domain-invariant but non-discriminative to the category. To overcome the two issues, we present a novel Informative Class-Conditioned Feature Alignment (IC2FA) approach for UDA, which utilizes a twofold method: informative feature disentanglement and class-conditioned feature alignment, designed to address the above two challenges, respectively. More specifically, to surmount the first drawback, we cooperatively disentangle the two domains to obtain informative transferable features; here, Variational Information Bottleneck (VIB) is employed to encourage the learning of task-related semantic representations and suppress task-unrelated information. With regard to the second bottleneck, we optimize a new metric, termed Conditional Sliced Wasserstein Distance (CSWD), which explicitly estimates the intra-class discrepancy and the inter-class margin. The intra-class and inter-class CSWDs are minimized and maximized, respectively, to yield the domain-invariant discriminative features. IC2FA equips class-conditioned feature alignment with informative feature disentanglement and causes the two procedures to work cooperatively, which facilitates informative discriminative features adaptation. Extensive experimental results on three domain adaptation datasets confirm the superiority of IC2FA. Wanxia Deng, Yawen Cui, Zhen Liu 0004, Gangyao Kuang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ACM Multimedia | 5 |
| 2021 | Dual-branch combination network (DCN): Towards accurate diagnosis and lesion segmentation of COVID-19 using CT images
Kai Gao 0011, Jianpo Su, Zhongbiao Jiang, Zhichao Feng, Hui Shen 0004, Pengfei Rong, Xin Xu 0001, Yuexiang Yang, Wei Wang 0434, Dewen Hu |
Medical Image Anal. | 12 |
| 2021 | Joint Embedding Learning and Low-Rank Approximation: A Framework for Incomplete Multiview LearningabstractIn real-world applications, not all instances in the multiview data are fully represented. To deal with incomplete data, incomplete multiview learning (IML) rises. In this article, we propose the joint embedding learning and low-rank approximation (JELLA) framework for IML. The JELLA framework approximates the incomplete data by a set of low-rank matrices and learns a full and common embedding by linear transformation. Several existing IML methods can be unified as special cases of the framework. More interestingly, some linear transformation-based complete multiview methods can be adapted to IML directly with the guidance of the framework. Thus, the JELLA framework improves the efficiency of processing incomplete multiview data, and bridges the gap between complete multiview learning and IML. Moreover, the JELLA framework can provide guidance for developing new algorithms. For illustration, within the framework, we propose the IML with the block-diagonal representation (IML-BDR) method. Assuming that the sampled examples have an approximate linear subspace structure, IML-BDR uses the block-diagonal structure prior to learning the full embedding, which would lead to more correct clustering. A convergent alternating iterative algorithm with the successive over-relaxation optimization technique is devised for optimization. The experimental results on various datasets demonstrate the effectiveness of IML-BDR. Chenping Hou, Dongyun Yi, Jubo Zhu, Dewen Hu |
IEEE Trans. Cybern. | 5 |
| 2021 | Nonparametric Bayesian Prior Inducing Deep Network for Automatic Detection of Cognitive StatusabstractPilots' brain fatigue status recognition faces two important issues. They are how to extract brain cognitive features and how to identify these fatigue characteristics. In this article, a gamma deep belief network is proposed to extract multilayer deep representations of high-dimensional cognitive data. The Dirichlet distributed connection weight vector is upsampled layer by layer in each iteration, and then the hidden units of the gamma distribution are downsampled. An effective upper and lower Gibbs sampler is formed to realize the automatic reasoning of the network structure. In order to extract the 3-D instantaneous time-frequency distribution spectrum of electroencephalogram (EEG) signals and avoid signal modal aliasing, this article also proposes a smoothed pseudo affine Wigner-Ville distribution method. Finally, experimental results show that our model achieves satisfactory results in terms of both recognition accuracy and stability. Qi Wu 0003, Dewen Hu, Ping-Yu Deng, Yulian Cao, Wen-Ming Zhang, Limin Zhu 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Joint Clustering and Discriminative Feature Alignment for Unsupervised Domain AdaptationabstractUnsupervised Domain Adaptation (UDA) aims to learn a classifier for the unlabeled target domain by leveraging knowledge from a labeled source domain with a different but related distribution. Many existing approaches typically learn a domain-invariant representation space by directly matching the marginal distributions of the two domains. However, they ignore exploring the underlying discriminative features of the target data and align the cross-domain discriminative features, which may lead to suboptimal performance. To tackle these two issues simultaneously, this paper presents a Joint Clustering and Discriminative Feature Alignment (JCDFA) approach for UDA, which is capable of naturally unifying the mining of discriminative features and the alignment of class-discriminative features into one single framework. Specifically, in order to mine the intrinsic discriminative information of the unlabeled target data, JCDFA jointly learns a shared encoding representation for two tasks: supervised classification of labeled source data, and discriminative clustering of unlabeled target data, where the classification of the source domain can guide the clustering learning of the target domain to locate the object category. We then conduct the cross-domain discriminative feature alignment by separately optimizing two new metrics: 1) an extended supervised contrastive learning, i.e., semi-supervised contrastive learning 2) an extended Maximum Mean Discrepancy (MMD), i.e., conditional MMD, explicitly minimizing the intra-class dispersion and maximizing the inter-class compactness. When these two procedures, i.e., discriminative features mining and alignment are integrated into one framework, they tend to benefit from each other to enhance the final performance from a cooperative learning perspective. Experiments are conducted on four real-world benchmarks (e.g., Office-31, ImageCLEF-DA, Office-Home and VisDA-C). All the results demonstrate that our JCDFA can obtain remarkable margins over state-of-the-art domain adaptation methods. Comprehensive ablation studies also verify the importance of each key component of our proposed algorithm and the effectiveness of combining two learning strategies into a framework. Wanxia Deng, Qing Liao 0001, Lingjun Zhao, Deke Guo, Gangyao Kuang, Dewen Hu, Li Liu 0002 |
IEEE Trans. Image Process. | 6 |
| 2021 | Enhancing ISAR Image Efficiently via Convolutional Reweighted l1 MinimizationabstractInverse synthetic aperture radar (ISAR) imaging for the sparse aperture data is affected by considerable artifacts, because under-sampling of data produces high-level grating and side lobes. Noting the ISAR image generally exhibits strong sparsity, it is often obtained by sparse signal recovery (SSR) in case of sparse aperture. The image obtained by SSR, however, is often dominated by strong isolated scatterers, resulting in difficulty to recognize the structure of target. This paper proposes a novel approach to enhance the ISAR image obtained from the sparse aperture data. Although the scatterers of target are isolated in the ISAR image, they should be associated with the neighborhood to reflect some intrinsic structural information of the target. A convolutional reweighted l1minimization model, therefore, is proposed to model the structural sparsity of ISAR image. Specifically, the ISAR image is reconstructed by solving a sequence of reweighted l1problems, where the weight of each pixel used for the next iteration is calculated from the convolution of its neighbor values in the current solution. The problem is solved by the alternating direction of multipliers (ADMM) and linearized approximation, respectively, to improve the computational efficiency. Experimental results based on both simulated and measured data validate that the proposed algorithm is effective to enhance the ISAR image, robust to noise, and more impressively, very efficient to implement. Shuanghui Zhang, Yongxiang Liu, Xiang Li 0014, Dewen Hu |
IEEE Trans. Image Process. | 4 |
| 2021 | Removal of Micro-Doppler Effect of ISAR Image Based on Laplacian Regularized Nonconvex Low-Rank RepresentationabstractThe micro-Doppler (m-D) effect caused by micro-motion degrades the readability of the inverse synthetic aperture radar (ISAR) image. To achieve well-focused ISAR image of the target with the micro-motion part, this paper proposes a novel approach for the removal of m-D effect of ISAR image. Note that the range profiles of the rigid body are similar to each other, making the respective data matrix low-rank. Those of the micro-motion part, in contrary, generally fluctuate in different range cells, whose data matrix is sparse. Therefore, the removal of m-D effect can be naturally solved by the robust principal component analysis (RPCA)-a convenient convex program to decompose an auxiliary matrix into a low-rank matrix and a sparse one. In RPCA, the rank of a matrix is described by the nuclear norm, which is convex but leads to a suboptimal solution. To address it, we utilize a nonconvex surrogate, i.e., the summation of logistic function of the singular values of a matrix, to approximate the rank. Moreover, the range profiles of the rigid body are generally locally similar. To capture this geometric structured information, we further introduce a Laplacian regularization into the model. Then, the Laplacian regularized nonconvex low-rank (LRNL) model is solved efficiently by the linearized alternating direction method (ADM). Extensive experimental results based on both simulated and measured data demonstrate the effectiveness of the proposed approach on the removal of m-D effect of ISAR image. Shuanghui Zhang, Yongxiang Liu, Xiang Li 0014, Dewen Hu |
IEEE Trans. Image Process. | 4 |
| 2020 | Dynamic Group Convolution for Accelerating Convolutional Neural Networks
Zhuo Su 0002, Linpu Fang, Wenxiong Kang, Dewen Hu, Matti Pietikäinen, Li Liu 0002 |
ECCV (6) | 4 |
| 2020 | GAMA: Graph Attention Multi-agent reinforcement learning algorithm for cooperation
Haoqiang Chen, Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Ming Zhang 0027 |
Appl. Intell. | 4 |
| 2020 | A Dynamic User Interface Based BCI Environmental Control SystemabstractIn this study, a dynamic user interface (UI) is proposed in visual P300 Brain-Computer Interface (BCI) based environmental control system. A head-mounted Augmented Reality (AR) glass is used as the interactive media, which is used to assists the BCI system to build the dynamic UI with the scene in subject’s field of view. In the dynamic UI, based on the objects detected by the AR glass, options are dynamically generated. The subject can assign tasks by selecting different options in the dynamic UI. Five subjects successfully completed the task of controlling household appliances and navigating wheelchairs to designated destinations. Compared to static UI, the proposed dynamic UI has a 17.4% improvement in time delay. On average, only 1.9% of the commands resulted in incorrect operations. The dynamic UI makes progress in reducing time delay and incorrect operations. The proposed system provides a brand-new interactive method in BCI based applications. Saisai Zhong, Yadong Liu 0001, Yang Yu 0014, Jingsheng Tang, Zongtan Zhou, Dewen Hu |
Int. J. Hum. Comput. Interact. | 6 |
| 2020 | 3D human pose estimation by depth map
Jianzhai Wu, Dewen Hu, Fengtao Xiang, Xingsheng Yuan, Jiongming Su |
Vis. Comput. | 2 |
| 2019 | Dense-CAM: Visualize the Gender of Brains with MRI ImagesabstractStudying the gender differences of brains is important to understand the brain cognitive mechanism. With the increasing amounts of brain imaging data, researchers start to use deep learning for brain image processing. However, the interpretability of deep models is the big gap between them. At present, deep model visualization is the main method to solve the interpretability problem of deep networks. Class Activation Map (CAM) is a commonly-used deep model visualization method. But the resolution of CAMs are restricted since it only uses the last layer for visualization. In this paper, we proposed a novel convolutional network named Dense-CAM by combining DenseNet and CAM and realized the visualization of the whole network so that it can generate more accurate and more robust deep model visualization. The network was tested with the gender classification problem using more than 6000 samples and achieved an accuracy of 92.93%. Brain regions with significant differences between men and women are found with the proposed method, which can be used for future brain imaging studies. Kai Gao 0011, Hui Shen 0004, Yadong Liu 0001, Dewen Hu |
IJCNN | 5 |
| 2019 | An Tactile ERP-Based Brain-Computer Interface for CommunicationabstractA classical visual event relative potential (ERP) brain–computer interface (BCI) system relies on visual stimuli to choose commands. Users obtain most information about their surroundings visually as well. This large amount of information can aggravate visual burden and fatigue. In our study, we proposed a novel approach to evoke ERP with a tactile stimulus. To achieve this approach, we first designed a wireless stimulus module with vibrators to provide a tactile stimulus for the system. The vibrators were located on the subject’s arm to imitate the joint motion of a robotic arm. Then, the ERP feature and the parameters of classifiers were obtained through offline experimental data analysis. Based on the analysis, the suitable electrode channels, stimulus onset asynchrony (SOA), and filter upper limit were different for different subjects. According to those outcomes, a unique classifier was designed for each subject. Finally, 10 healthy BCI-naive subjects participated in online experiments to evaluate the performance of our tactile BCI system; they achieved an accuracy range from 78.67% to 100% with an average of 89.1% and an instantaneous transmission rate (ITR) range from 7.77 to 28.70 bits/min with an average of 14.77 bits/min. The accuracy of different subjects and SOAs remained relatively stable, the ITR fluctuated mainly due to the different SOAs, and we achieved balance between ITR and accuracy. Yadong Liu 0001, Jingjun Wang, Erwei Yin, Yang Yu 0014, Zongtan Zhou, Dewen Hu |
Int. J. Hum. Comput. Interact. | 6 |
| 2019 | Toward Brain-Actuated Mobile PlatformabstractThis study presents a brain–computer interface (BCI) system aimed at providing disabled patients with mobile solutions for practical use. The proposed system employs an omnidirectional chassis and a bionic robot arm to construct a multi-functional mobile platform. In addition, the system is equipped with a Kinect and 12 ultrasonic sensors to capture environment information. Based on artificial intelligence technology, the mobile system can understand the environment and smartly completes certain tasks. A hybrid BCI combined with movement imagery paradigm and asynchronous P300 paradigm is designed to translate human intent to computer commands. The users interact with the system in a flexible way: on the one hand, the user issues commands to drive the system directly; on the other hand, the system searches for predefined operable targets and reports the results to the user. Once the user confirms the target, the system will automatically complete the associated operation. To evaluate the system’s performance, a testing environment with a small room, aisle, and an elevator was built to simulate the mobile tasks in the daily scene. Participants were instructed to operate the mobile system in the room, aisle, and using the elevator to go outdoors. In this study, four subjects participated in the test, and all of them completed the task. Jingsheng Tang, Yadong Liu 0001, Jun Jiang 0001, Yang Yu 0014, Dewen Hu, Zongtan Zhou |
Int. J. Hum. Comput. Interact. | 5 |
| 2019 | Correlation-based channel selection and regularized feature optimization for MI-based BCI
Jing Jin 0001, Yangyang Miao, Ian Daly, Cili Zuo, Dewen Hu, Andrzej Cichocki |
Neural Networks | 5 |
| 2019 | Safe Classification with Augmented FeaturesabstractWith the evolution of data collection methods, it is possible to produce abundant data described by multiple feature sets. Previous studies show that including more features does not necessarily bring positive effects. How to prevent the augmented features from worsening classification performance is crucial but rarely studied. In this paper, we study this challenging problem by proposing a safe classification approach, whose accuracy is never degenerated when exploiting augmented features. We propose two ways to achieve the safeness of our method named as SAfe Classification (SAC). First, to leverage augmented features, we learn various types of classifiers and adapt them by employing a specially designed robust loss. It provides various candidate classifiers to meet the assumption of safeness operation. Second, we search for a safe prediction by integrating all candidate classifiers. Under a mild assumption, the integrated classifier has theoretical safeness guarantee. Several new optimization methods have been developed to accommodate the problems with proved convergence. Besides evaluating SAC on 16 data sets, we also apply SAC in the application of diagnostic classification of schizophrenia since it has vast application potentiality. Experimental results demonstrate the effectiveness of SAC in both tackling safeness problem and discriminating schizophrenic patients from healthy controls. Chenping Hou, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Towards correlation-based time window selection method for motor imagery BCIs
Jiankui Feng, Erwei Yin, Jing Jin 0001, Rami Saab, Ian Daly, Xingyu Wang 0004, Dewen Hu, Andrzej Cichocki |
Neural Networks | 7 |
| 2018 | Gender Identification of Human Brain Image with A Novel 3D DescriptorabstractDetermining gender by examining the human brain is not a simple task because the spatial structure of the human brain is complex, and no obvious differences can be seen by the naked eyes. In this paper, we propose a novel three-dimensional feature descriptor, the three-dimensional weighted histogram of gradient orientation (3D WHGO) to describe this complex spatial structure. The descriptor combines local information for signal intensity and global three-dimensional spatial information for the whole brain. We also improve a framework to address the classification of three-dimensional images based on MRI. This framework, three-dimensional spatial pyramid, uses additional information regarding the spatial relationship between features. The proposed method can be used to distinguish gender at the individual level. We examine our method by using the gender identification of individual magnetic resonance imaging (MRI) scans of a large sample of healthy adults across four research sites, resulting in up to individual-level accuracies under the optimized parameters for distinguishing between females and males. Compared with previous methods, the proposed method obtains higher accuracy, which suggests that this technology has higher discriminative power. With its improved performance in gender identification, the proposed method may have the potential to inform clinical practice and aid in research on neurological and psychiatric disorders. Fanglin Chen 0001, Dewen Hu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2017 | Detect visual field using eye tracking and steady-state visual evoked potentialabstractThis paper makes the subjects' sight locked in a certain area using an eye tracker, getting Steady-state visual evoked potential (SSVEP) from flickering stimuli with a fixed frequency but at random positions, in order to observe the impact of stimulus at different positions and their distances on the electroencephalogram (EEG). The result suggests that if human have to select the positions of stimuli of SSVEP-BCI, it is an agreeable strategy to separate them at least 4 ° for avoiding the possible mistakes. We hope that it could help in setting distances between stimuli or updating pattern selection algorithms in the future BCI system and other paradigms. Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Erwei Yin |
SMC | 4 |
| 2017 | Toward a Hybrid BCI: Self-Paced Operation of a P300-based Speller by Merging a Motor Imagery-Based "Brain Switch" into a P300 Spelling ApproachabstractThis study presents the self-paced operation of a brain–computer interface (BCI) speller, which can be voluntarily turned on/off by merging a motor imagery (MI)-based brain switch into a P300-based BCI speller. From an off state (idle state), the users can generate a “control signal” by consciously changing the cognitive state differential from the idle state to turn on a P300-based spelling system when he or she wants to spell words. With the system turned on, the user can spell words, and then, the spelling system can be voluntarily turned off and switched to the initial state using a command. In this paradigm, the participants tried to perform the two different cognitive tasks sequentially, rather than simultaneously, and multiple EEG components were processed sequentially. The practicability and effectiveness of the proposed approach were validated by eleven participants, and all of them achieved a satisfactory performance. For the P300 speller, they achieved an average PITR of 42.61 bits/min. The preliminary results indicated that the proposed hybrid BCI system with different mental strategies operating sequentially is feasible and has potential applications for practical self-paced control. Yang Yu 0014, Zongtan Zhou, Jun Jiang 0001, Erwei Yin, Kunjia Liu, Jingjun Wang, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 8 |
| 2017 | Blind source separation of functional MRI scans of the human brain based on canonical correlation analysis
Xingjie Wu, Hui Shen 0004, Ming Li 0028, Yun-an Hu, Dewen Hu |
Neurocomputing | 6 |
| 2017 | Salient Region Detection Using Diffusion Process on a Two-Layer Sparse GraphabstractDiffusion-based salient region detection has recently received intense research attention. In this paper, we present some effective improvements concerning two important aspects of diffusion-based methods: the construction of the diffusion matrix and the seed vector. First, we construct a two-layer sparse graph, which is generated by connecting each node to its neighboring nodes and the most similar node that shares common boundaries with its neighboring nodes. Compared with the most frequently used two-layer neighborhood graph, our graph not only effectively uses local spatial relationships, but also removes dissimilar redundant nodes. Second, we use the spatial variance of superpixel clusters to obtain the seed vector and, compared with the previously most-used boundary prior, our approach can better distinguish saliency seeds from the background seeds, especially when salient objects appear near the image boundaries. Finally, we calculate two preliminary saliency maps using the saliency and background seed vectors, and more accurate results are obtained using the manifold ranking diffusion method. Integrating these two diffusion-based saliency maps, we obtain the final saliency map. Extensive experiments in which we compare our method with 20 existing state-of-the-art methods on five benchmark data sets: ASD, DUT-OMRON, ECSSD, MSRA5K, and MSRA10K, show that the proposed method performs better in terms of various evaluation metrics. Li Zhou 0003, Zhaohui Yang 0008, Zongtan Zhou, Dewen Hu |
IEEE Trans. Image Process. | 4 |
| 2016 | Evaluation of LBP and Deep Texture Descriptors with a New Robustness Benchmark
Li Liu 0002, Paul W. Fieguth, Xiaogang Wang 0001, Matti Pietikäinen, Dewen Hu |
ECCV (3) | 5 |
| 2016 | A P300-Based Brain-Computer Interface for Chinese Character InputabstractThe majority of previously developed assistive communication brain–computer interface systems have primarily focused on languages that are written in alphabetic scripts. However, languages that are written in logographic scripts, such as those in Chinese hanzi (or sinograms), pose a challenge for the implementation of visual spelling systems because it is impossible to simultaneously display thousands of items in a stimulus matrix of a reasonable size. In this study, a P300 visual spelling system that uses a novel method to input Chinese sinograms developed with a Hanyu Pinyin-based method is presented. This method transcribes a Chinese Pinyin into initial consonant and vowel components according to its Mandarin pronunciation. In this paradigm, each sinogram is input by selecting the initial consonant and then the vowel components and subsequently selecting the sinogram itself. Ten healthy subjects participated in the study and achieved an average offline accuracy of 92.6% with a mean information transfer rate of 39.2 bits/min and an average online input speed of one sinogram per 43.9 s. The preliminary results presented here indicated that the online input of Chinese text using a Pinyin-based visual speller is feasible. Yang Yu 0014, Zongtan Zhou, Erwei Yin, Jun Jiang 0001, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 6 |
| 2016 | An Auditory-Tactile Visual Saccade-Independent P300 Brain-Computer InterfaceabstractMost P300 event-related potential (ERP)-based brain-computer interface (BCI) studies focus on gaze shift-dependent BCIs, which cannot be used by people who have lost voluntary eye movement. However, the performance of visual saccade-independent P300 BCIs is generally poor. To improve saccade-independent BCI performance, we propose a bimodal P300 BCI approach that simultaneously employs auditory and tactile stimuli. The proposed P300 BCI is a vision-independent system because no visual interaction is required of the user. Specifically, we designed a direction-congruent bimodal paradigm by randomly and simultaneously presenting auditory and tactile stimuli from the same direction. Furthermore, the channels and number of trials were tailored to each user to improve online performance. With 12 participants, the average online information transfer rate (ITR) of the bimodal approach improved by 45.43% and 51.05% over that attained, respectively, with the auditory and tactile approaches individually. Importantly, the average online ITR of the bimodal approach, including the break time between selections, reached 10.77 bits/min. These findings suggest that the proposed bimodal system holds promise as a practical visual saccade-independent P300 BCI. Erwei Yin, Timothy J. Zeyl, Rami Saab, Dewen Hu, Zongtan Zhou, Tom Chau |
Int. J. Neural Syst. | 4 |
| 2016 | Extended local binary patterns for face recognition
Li Liu 0002, Paul W. Fieguth, Guoying Zhao 0001, Matti Pietikäinen, Dewen Hu |
Inf. Sci. | 5 |
| 2016 | Introduction to the Special Issue on Unmanned Intelligent Vehicles in ChinaabstractThe papers in this special section present some recent advances in ongoing research of unmanned intelligent vehicles in China. Lingxi Li 0001, Dewen Hu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2015 | Improve scene classification by using feature and kernel combination
Fanglin Chen 0001, Li Zhou 0003, Dewen Hu |
Neurocomputing | 4 |
| 2015 | Fusing Sorted Random Projections for Robust Texture and Material ClassificationabstractThis paper presents a conceptually simple, and robust, yet highly effective, approach to both texture classification and material categorization. The proposed system is composed of three components: 1) local, highly discriminative, and robust features based on sorted random projections (RPs), built on the universal and information-preserving properties of RPs; 2) an effective bag-of-words global model; and 3) a novel approach for combining multiple features in a support vector machine classifier. The proposed approach encompasses the simplicity, broad applicability, and efficiency of the three methods. We have tested the proposed approach on eight popular texture databases, including Flickr Materials Database, a highly challenging materials database. We compare our method with 13 recent state-of-the-art methods, and the experimental results show that our texture classification system yields the best classification rates of which we are aware of 99.37% for Columbia-Utrecht, 97.16% for Brodatz, 99.30% for University of Maryland Database, and 99.29% for Kungliga Tekniska högskolan-textures under varying illumination, pose, and scale. Moreover, the proposed approach significantly outperforms the current state-of-the-art approach in materials categorization, with an improvement to classification accuracy of 67%. Li Liu 0002, Paul W. Fieguth, Dewen Hu, Yingmei Wei, Gangyao Kuang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Salient Region Detection via Integrating Diffusion-Based Compactness and Local ContrastabstractSalient region detection is a challenging problem and an important topic in computer vision. It has a wide range of applications, such as object recognition and segmentation. Many approaches have been proposed to detect salient regions using different visual cues, such as compactness, uniqueness, and objectness. However, each visual cue-based method has its own limitations. After analyzing the advantages and limitations of different visual cues, we found that compactness and local contrast are complementary to each other. In addition, local contrast can very effectively recover incorrectly suppressed salient regions using compactness cues. Motivated by this, we propose a bottom-up salient region detection method that integrates compactness and local contrast cues. Furthermore, to produce a pixel-accurate saliency map that more uniformly covers the salient objects, we propagate the saliency information using a diffusion process. Our experimental results on four benchmark data sets demonstrate the effectiveness of the proposed method. Our method produces more accurate saliency maps with better precision-recall curve and higher F-Measure than other 19 state-of-the-arts approaches on ASD, CSSD, and ECSSD data sets. Li Zhou 0003, Zhaohui Yang 0008, Zongtan Zhou, Dewen Hu |
IEEE Trans. Image Process. | 5 |
| 2015 | Including Signal Intensity Increases the Performance of Blind Source Separation on Brain Imaging DataabstractWhen analyzing brain imaging data, blind source separation (BSS) techniques critically depend on the level of dimensional reduction. If the reduction level is too slight, the BSS model would be overfitted and become unavailable. Thus, the reduction level must be set relatively heavy. This approach risks discarding useful information and crucially limits the performance of BSS techniques. In this study, a new BSS method that can work well even at a slight reduction level is presented. We proposed the concept of "signal intensity" which measures the significance of the source. Only picking the sources with significant intensity, the new method can avoid the overfitted solutions which are nonexistent artifacts. This approach enables the reduction level to be set slight and retains more useful dimensions in the preliminary reduction. Comparisons between the new and conventional algorithms were performed on both simulated and real data. Ming Li 0028, Yadong Liu 0001, Fanglin Chen 0001, Dewen Hu |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Adaptive NN Control of a Class of Nonlinear Systems With Asymmetric Saturation ActuatorsabstractIn this note, adaptive neural network (NN) control is investigated for a class of uncertain nonlinear systems with asymmetric saturation actuators and external disturbances. To handle the effect of nonsmooth asymmetric saturation nonlinearity, a Gaussian error function-based continuous differentiable asymmetric saturation model is employed such that the backstepping technique can be used in the control design. The explosion of complexity in traditional backstepping design is avoided using dynamic surface control. Using radial basis function NN, adaptive control is developed to guarantee that all the signals in the closed-loop system are semiglobally uniformly ultimately bounded, and the tracking error converges to a small neighborhood of origin by appropriately choosing design constants. The effectiveness of the proposed control is demonstrated in the simulation study. Shuzhi Sam Ge, Zhiqiang Zheng 0002, Dewen Hu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2015 | Multiobjective Reinforcement Learning: A Comprehensive OverviewabstractReinforcement learning (RL) is a powerful paradigm for sequential decision-making under uncertainties, and most RL algorithms aim to maximize some numerical value which represents only one long-term objective. However, multiple long-term objectives are exhibited in many real-world decision and control systems, so recently there has been growing interest in solving multiobjective reinforcement learning (MORL) problems where there are multiple conflicting objectives. The aim of this paper is to present a comprehensive overview of MORL. The basic architecture, research topics, and naïve solutions of MORL are introduced at first. Then, several representative MORL approaches and some important directions of recent research are comprehensively reviewed. The relationships between MORL and other related research are also discussed, which include multiobjective optimization, hierarchical RL, and multiagent RL. Moreover, research challenges and open problems of MORL techniques are suggested. Chunming Liu, Xin Xu 0001, Dewen Hu |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2014 | Learning Effective Event Models to Recognize a Large Number of Human ActionsabstractHuman action recognition in videos is an important problem in computer vision, but it is very challenging, especially when recognizing a large number of human actions. First, it is difficult to capture the crucial motion patterns that discriminate among these actions. Second, the method should be scalable for large datasets because more training examples are often collected for more action classes. In this paper, we employ latent models to capture the crucial motion patterns, and we propose an effective learning algorithm that can efficiently address large datasets. To capture the crucial motion patterns, we define an “event” for each category, and we add a latent variable that indicates the start of the event. The event has a length of several frames that can differ across the categories. To train effective latent models for a large number of action classes, we employ a multi-class formulation with latent variables, and we address this problem by solving a dual quadratic programming (QP) problem with linear inequality constraints. To make the algorithm scalable for large datasets, we propose an improved QP solver that converges quickly for large QP problems that have a very large number of linear inequality constraints in real-world applications. We examine the proposed approach on the HMDB51 and UCF50 datasets. Comparison results have been reported to demonstrate the effectiveness of the proposed technique. Our approach outperforms state-of-the-art results for both datasets. Jianzhai Wu, Dewen Hu |
IEEE Trans. Multim. | 2 |
| 2014 | Action recognition by hidden temporal models
Jianzhai Wu, Dewen Hu, Fanglin Chen 0001 |
Vis. Comput. | 2 |
| 2013 | Graph-based image segmentation using directional nearest neighbor graph
Dewen Hu, Hui Shen 0004, Guiyu Feng |
Sci. China Inf. Sci. | 2 |
| 2013 | Scene recognition combining structural and textural features
Li Zhou 0003, Dewen Hu, Zongtan Zhou |
Sci. China Inf. Sci. | 2 |
| 2013 | Scene classification using multi-resolution low-level feature combination
Li Zhou 0003, Zongtan Zhou, Dewen Hu |
Neurocomputing | 3 |
| 2013 | Scene classification using a multi-resolution bag-of-features model
Li Zhou 0003, Zongtan Zhou, Dewen Hu |
Pattern Recognit. | 3 |
| 2012 | Robust Mean Shift Tracking with Background Information
Guiyu Feng, Dewen Hu |
ISNN (2) | 3 |
| 2012 | Tracking objects using shape context matching
Hui Shen 0004, Guiyu Feng, Dewen Hu |
Neurocomputing | 4 |
| 2011 | Continuous-action reinforcement learning with fast policy search and adaptive basis function selection
Xin Xu 0001, Chunming Liu, Dewen Hu |
Soft Comput. | 3 |
| 2011 | Cerebral Artery-Vein Separation Using 0.1-Hz Oscillation in Dual-Wavelength Optical ImagingabstractWe present a novel artery-vein separation method using 0.1-Hz oscillation at two wavelengths with optical imaging of intrinsic signals (OIS). The 0.1-Hz oscillation at a green light wavelength of 546 nm exhibits greater amplitude in arteries than in veins and is primarily caused by vasomotion, whereas the 0.1-Hz oscillation at a red light wavelength of 630 nm exhibits greater amplitude in veins than in arteries and is primarily caused by changes of deoxyhemoglobin concentration. This spectral feature enables cortical arteries and veins to be segmented independently. The arteries can be segmented on the 0.1-Hz amplitude image at 546 nm using matched filters of a modified dual Gaussian model combining with a single Gaussian model. The veins are a combination of vessels segmented on both amplitude images at the two wavelengths using multiscale matched filters of single Gaussian model. Our method can separate most of the thin arteries and veins from each other, especially the thin arteries with low contrast in raw gray images. In vivo OIS experiments demonstrate the separation ability of the 0.1-Hz based segmentation method in cerebral cortex of eight rats. Two validation studies were undertaken to evaluate the performance of the method by quantifying the arterial and venous length based on a reference standard. The results indicate that our 0.1-Hz method is very effective in separating both large and thin arteries and veins regardless of vessel crossover or overlapping to great extent in comparison with previous methods. Dewen Hu, Yadong Liu 0001, Ming Li 0028 |
IEEE Trans. Medical Imaging | 2 |
| 2011 | Hierarchical Approximate Policy Iteration With Binary-Tree State Space DecompositionabstractIn recent years, approximate policy iteration (API) has attracted increasing attention in reinforcement learning (RL), e.g., least-squares policy iteration (LSPI) and its kernelized version, the kernel-based LSPI algorithm. However, it remains difficult for API algorithms to obtain near-optimal policies for Markov decision processes (MDPs) with large or continuous state spaces. To address this problem, this paper presents a hierarchical API (HAPI) method with binary-tree state space decomposition for RL in a class of absorbing MDPs, which can be formulated as time-optimal learning control tasks. In the proposed method, after collecting samples adaptively in the state space of the original MDP, a learning-based decomposition strategy of sample sets was designed to implement the binary-tree state space decomposition process. Then, API algorithms were used on the sample subsets to approximate local optimal policies of sub-MDPs. The original MDP was decomposed into a binary-tree structure of absorbing sub-MDPs, constructed during the learning process, thus, local near-optimal policies were approximated by API algorithms with reduced complexity and higher precision. Furthermore, because of the improved quality of local policies, the combined global policy performed better than the near-optimal policy obtained by a single API algorithm in the original MDP. Three learning control problems, including path-tracking control of a real mobile robot, were studied to evaluate the performance of the HAPI method. With the same setting for basis function selection and sample collection, the proposed HAPI obtained better near-optimal policies than previous API methods such as LSPI and KLSPI. Xin Xu 0001, Chunming Liu, Simon X. Yang, Dewen Hu |
IEEE Trans. Neural Networks | 4 |
| 2009 | Handprint Recognition: A Novel Biometric Technology
Guiyu Feng, MiYi Duan, Dewen Hu, Yabin Hu |
ISNN (3) | 4 |
| 2009 | Local region structured noise reduction for cortical optical imaging
Yadong Liu 0001, Dewen Hu, Zongtan Zhou, Fayi Liu |
Neurocomputing | 2 |
| 2009 | Incremental Laplacian eigenmaps by preserving adjacent information between data points
Junsong Yin, Xinsheng Huang, Dewen Hu |
Pattern Recognit. Lett. | 4 |
| 2008 | Design of an artificial bionic neural network to control fish-robot's locomotion
Daibing Zhang, Dewen Hu, Lincheng Shen, Haibin Xie |
Neurocomputing | 2 |
| 2008 | A Direct Locality Preserving Projections (DLPP) Algorithm for Image Recognition
Guiyu Feng, Dewen Hu, Zongtan Zhou |
Neural Process. Lett. | 2 |
| 2008 | Globally Consistent Reconstruction of Ripped-Up DocumentsabstractOne of the most crucial steps for automatically reconstructing ripped-up documents is to find a globally consistent solution from the ambiguous candidate matches. However, little work has been done so far to solve this problem in a general computational framework without using application-specific features. In this paper, we propose a global approach for reconstructing ripped-up documents by first finding candidate matches from document fragments using curve matching and then disambiguating these candidates through a relaxation process to reconstruct the original document. The candidate disambiguation problem is formulated in a relaxation scheme, in which the definition of compatibility between neighboring matches is proposed and global consistency is defined as the global criterion. Initially, global match confidences are assigned to each of the candidate matches. After that, the overall local relationships among neighboring matches are evaluated by computing their global consistency. Then these confidences are iteratively updated using the gradient projection method to maximize the criterion. This leads to a globally consistent solution and thus provides a sound document reconstruction. The overall performance of our approach in several practical experiments is illustrated. The results indicate that the reconstruction of ripped-up documents up to fifty pieces is possibly accomplished automatically. Liangjia Zhu, Zongtan Zhou, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2008 | Comment on "two-dimensional locality preserving projections (2DLPP) with its application to palmprint recognition"
Dewen Hu, Guiyu Feng, Zongtan Zhou |
Pattern Recognit. | 1 |
| 2008 | Noisy manifold learning using neighborhood smoothing embedding
Junsong Yin, Dewen Hu, Zongtan Zhou |
Pattern Recognit. Lett. | 2 |
| 2008 | Convergence of Nonautonomous Cohen-Grossberg-Type Neural Networks With Variable DelaysabstractThis paper is concerned with the global convergence of the solutions of a nonautonomous system with variable delays, arising from the description of the states of neurons in delayed Cohen-Grossberg type in a time-varying situation. By exploring intrinsic features between nonautonomous system and its asymptotic equation, several novel sufficient conditions are established to ensure that all solutions of the networks converge to a periodic function or a constant vector for delayed Cohen-Grossberg-type neural network (NN) models in time-varying situation. The results can be applied directly to group of NNs models including Hopfield NNs, bidirectional association memory NNs, and cellular NNs. Our results are not only presented in terms of system parameters and can be easily verified but also are less restrictive than previously known criteria. Numerical simulations have also been presented to demonstrate the theoretical analysis. Zhaohui Yuan, Dewen Hu, Bingwen Liu |
IEEE Trans. Neural Networks | 3 |
| 2007 | Manifold Learning using Growing Locally Linear EmbeddingabstractLocally linear embedding (LLE) is an effective nonlinear dimensionality reduction method for exploring the intrinsic characteristics of high dimensional data. This paper mainly proposes a hierarchical framework manifold learning method, based on LLE and growing neural gas (GNG), named growing locally linear embedding (GLLE). First, we address the major limitations of the original LLE: intrinsic dimensionality estimation, neighborhood number selection and computational complexity. Then by embedding the topology learning mechanism in GNG, the proposed GLLE algorithm is able to preserve the global topological structures and hold the geometric characteristics of the input patterns, which make the projections more stable and robust. Theoretical analysis and experimental simulations show that GLLE with global topology preservation tackles the three limitations, gives faster learning procedure and lower reconstruction error, and stimulates the wide applications of manifold learning Junsong Yin, Dewen Hu, Zongtan Zhou |
CIDM | 2 |
| 2007 | Dynamic Analysis of a Novel Artificial Neural Oscillator
Daibing Zhang, Dewen Hu, Lincheng Shen, Haibin Xie |
ISNN (1) | 2 |
| 2007 | Advanced fuzzy cellular neural network: Application to CT liver images
Shitong Wang 0001, Duan Fu, Dewen Hu |
Artif. Intell. Medicine | 4 |
| 2007 | Two-dimensional locality preserving projections (2DLPP) with its application to palmprint recognition
Dewen Hu, Guiyu Feng, Zongtan Zhou |
Pattern Recognit. | 1 |
| 2007 | Possibility Theoretic Clustering and its Preliminary Application to Large Image Segmentation
Korris Fu-Lai Chung, Shitong Wang 0001, Dewen Hu |
Soft Comput. | 4 |
| 2007 | Kernel-Based Least Squares Policy Iteration for Reinforcement LearningabstractIn this paper, we present a kernel-based least squares policy iteration (KLSPI) algorithm for reinforcement learning (RL) in large or continuous state spaces, which can be used to realize adaptive feedback control of uncertain dynamic systems. By using KLSPI, near-optimal control policies can be obtained without much a priori knowledge on dynamic models of control plants. In KLSPI, Mercer kernels are used in the policy evaluation of a policy iteration process, where a new kernel-based least squares temporal-difference algorithm called KLSTD-Q is proposed for efficient policy evaluation. To keep the sparsity and improve the generalization ability of KLSTD-Q solutions, a kernel sparsification procedure based on approximate linear dependency (ALD) is performed. Compared to the previous works on approximate RL methods, KLSPI makes two progresses to eliminate the main difficulties of existing results. One is the better convergence and (near) optimality guarantee by using the KLSTD-Q algorithm for policy evaluation with high precision. The other is the automatic feature selection using the ALD-based kernel sparsification. Therefore, the KLSPI algorithm provides a general RL method with generalization performance and convergence guarantee for large-scale Markov decision problems (MDPs). Experimental results on a typical RL task for a stochastic chain problem demonstrate that KLSPI can consistently achieve better learning efficiency and policy quality than the previous least squares policy iteration (LSPI) algorithm. Furthermore, the KLSPI method was also evaluated on two nonlinear feedback control problems, including a ship heading control problem and the swing up control of a double-link underactuated pendulum called acrobot. Simulation results illustrate that the proposed method can optimize controller performance using little a priori information of uncertain dynamic systems. It is also demonstrated that KLSPI can be applied to online learning control by incorporating an initial controller to ensure online performance. Xin Xu 0001, Dewen Hu, Xicheng Lu |
IEEE Trans. Neural Networks | 2 |
| 2006 | Local Stability Analysis of Maximum Nongaussianity Estimation in Independent Component Analysis
Xin Xu 0001, Dewen Hu |
ISNN (1) | 3 |
| 2006 | Classification of Movement-Related Potentials for Brain-Computer Interface: A Reinforcement Training Approach
Zongtan Zhou, Dewen Hu |
ISNN (2) | 3 |
| 2006 | An alternative formulation of kernel LPP with application to image recognition
Guiyu Feng, Dewen Hu, David Zhang 0001, Zongtan Zhou |
Neurocomputing | 2 |
| 2006 | Clustering Analysis of Gene Expression Data based on Semi-supervised Visual Clustering Algorithm
Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng, Chen Shu, Dewen Hu |
Soft Comput. | 5 |
| 2006 | Robust maximum entropy clustering algorithm with its labeling for outliers
Shitong Wang 0001, Korris Fu-Lai Chung, Zhaohong Deng, Dewen Hu, Xisheng Wu |
Soft Comput. | 4 |
| 2006 | Experimental study on parameter choices in norm-r support vector regression machines with noisy input
Shitong Wang 0001, Jiagang Zhu, Korris Fu-Lai Chung, Dewen Hu |
Soft Comput. | 4 |
| 2006 | CATSMLP: Toward a Robust and Interpretable Multilayer Perceptron With Sigmoid Activation FunctionsabstractEnhancing the robustness and interpretability of a multilayer perceptron (MLP) with a sigmoid activation function is a challenging topic. As a particular MLP, additive TS-type MLP (ATSMLP) can be interpreted based on single-stage fuzzy IF-THEN rules, but its robustness will be degraded with the increase in the number of intermediate layers. This paper presents a new MLP model called cascaded ATSMLP (CATSMLP), where the ATSMLPs are organized in a cascaded way. The proposed CATSMLP is a universal approximator and is also proven to be functionally equivalent to a fuzzy inference system based on syllogistic fuzzy reasoning. Therefore, the CATSMLP may be interpreted based on syllogistic fuzzy reasoning in a theoretical sense. Meanwhile, due to the fact that syllogistic fuzzy reasoning has distinctive advantage over single-stage IF-THEN fuzzy reasoning in robustness, this paper proves in an indirect way that the CATSMLP is more robust than the ATSMLP in an upper-bound sense. Several experiments were conducted to confirm such a claim. Korris Fu-Lai Chung, Shitong Wang 0001, Zhaohong Deng, Dewen Hu |
IEEE Trans. Syst. Man Cybern. Part B | 4 |
| 2005 | Possibility Theoretic Clustering
Shitong Wang 0001, Korris Fu-Lai Chung, Dewen Hu, Lin Qing |
ICIC (1) | 4 |
| 2005 | Self-adaptive FastICA Based on Generalized Gaussian Model
Xin Xu 0001, Dewen Hu |
ISNN (1) | 3 |
| 2005 | A novel method for spatio-temporal pattern analysis of brain fMRI dataabstractA novel data processing procedure for fMRI was suggested in this paper, by which spatial and temporal characteristics of stimuli-induced signal dynamic responses can be investigated simultaneously. First the multitaper spectral estimation was utilized to estimate the spectrum of each voxel; the significance of the line frequency components at the interested frequency was tested to detect the task-related cortex areas; the temporal independent component analysis (tICA) was then applied to the activated voxels to obtain stimuli-induced signal dynamic responses. The advantages of this procedure are: few assumptions are needed for the cerebral hemodynamics and spatial distribution of task-related areas, problems which often appear in tICA analysis of fMRI data, such as the lack of stability, reliability and robustness, are overcome by the suggested method. Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Lirong Yan, Changlian Tan, Daxing Wu, Shuqiao Yao |
Sci. China Ser. F Inf. Sci. | 3 |
| 2005 | DSOM: a novel self-organizing model based on NO dynamic diffusing mechanism
Junsong Yin, Dewen Hu, Zongtan Zhou |
Sci. China Ser. F Inf. Sci. | 2 |
| 2005 | A new gaussian noise filter based on interval type-2 fuzzy logic systems
Shitong Wang 0001, Korris Fu-Lai Chung, Y. Y. Li, Dewen Hu, Xisheng Wu |
Soft Comput. | 4 |
| 2005 | Theoretically Optimal Parameter Choices for Support Vector Regression Machines with Noisy Input
Shitong Wang 0001, Jiagang Zhu, Korris Fu-Lai Chung, Lin Qing, Dewen Hu |
Soft Comput. | 5 |
| 2005 | Unscented Kalman filtering for additive noise case: augmented versus nonaugmentedabstractThis paper concerns the unscented Kalman filtering (UKF) for the nonlinear dynamic systems with additive process and measurement noises. It is widely accepted for such a case that the system state needs not to be augmented with noise vectors and the resultant nonaugmented UKF yields similar, if not the same, results to the augmented UKF. In this letter, we find that under the condition of n+/spl kappa/=const, the basic difference between them is that the augmented UKF draws a sigma set only once within a filtering recursion, while the nonaugmented UKF has to redraw a new set of sigma points to incorporate the effect of additive process noise. This difference generally favors the augmented UKF in that the odd-order moment information is partly captured by the nonlinearly transformed sigma points and propagated throughout the recursion. The simulation results agree well with the analyses. Yuanxin Wu, Dewen Hu, Meiping Wu, Xiaoping Hu 0002 |
IEEE Signal Process. Lett. | 2 |
| 2004 | Diffusion and Growing Self-Organizing Map: A Nitric Oxide Based Neural Model
Zongtan Zhou, Dewen Hu |
ISNN (1) | 3 |
| 2004 | A New Computational Model of Biological Vision for Stereopsis
Baoquan Song, Zongtan Zhou, Dewen Hu, Zhengzhi Wang |
ISNN (2) | 3 |
| 2004 | The Existence of Spurious Equilibrium in FastICA
Dewen Hu |
ISNN (1) | 2 |
| 2004 | Mobile Robot Path-Tracking Using an Adaptive Critic Learning PD Controller
Xin Xu 0001, Dewen Hu |
ISNN (2) | 3 |
| 2002 | Evolutionary adaptive-critic methods for reinforcement learningabstractIn this paper, a novel hybrid learning method is proposed for reinforcement learning problems with continuous state and action spaces. The reinforcement learning problems are modeled as Markov decision processes (MDPs) and the hybrid learning method combines evolutionary algorithms with gradient-based adaptive heuristic critic (AHC) algorithms to approximate the optimal policy of MDPs. The suggested method takes the advantages of evolutionary learning and gradient-based reinforcement learning to solve reinforcement learning problems. Simulation results on the learning control of an acrobot illustrate the efficiency of the presented method. Xin Xu 0001, Hangen He, Dewen Hu |
IEEE Congress on Evolutionary Computation | 3 |
| 2002 | Efficient Reinforcement Learning Using Recursive Least-Squares MethodsabstractThe recursive least-squares (RLS) algorithm is one of the most well-known algorithms used in adaptive filtering, system identification and adaptive control. Its popularity is mainly due to its fast convergence speed, which is considered to be optimal in practice. In this paper, RLS methods are used to solve reinforcement learning problems, where two new reinforcement learning algorithms using linear value function approximators are proposed and analyzed. The two algorithms are called RLS-TD(lambda) and Fast-AHC (Fast Adaptive Heuristic Critic), respectively. RLS-TD(lambda) can be viewed as the extension of RLS-TD(0) from lambda=0 to general lambda within interval [0,1], so it is a multi-step temporal-difference (TD) learning algorithm using RLS methods. The convergence with probability one and the limit of convergence of RLS-TD(lambda) are proved for ergodic Markov chains. Compared to the existing LS-TD(lambda) algorithm, RLS-TD(lambda) has advantages in computation and is more suitable for online learning. The effectiveness of RLS-TD(lambda) is analyzed and verified by learning prediction experiments of Markov chains with a wide range of parameter settings. The Fast-AHC algorithm is derived by applying the proposed RLS-TD(lambda) algorithm in the critic network of the adaptive heuristic critic method. Unlike conventional AHC algorithm, Fast-AHC makes use of RLS methods to improve the learning-prediction efficiency in the critic. Learning control experiments of the cart-pole balancing and the acrobot swing-up problems are conducted to compare the data efficiency of Fast-AHC with conventional AHC. From the experimental results, it is shown that the data efficiency of learning control can also be improved by using RLS methods in the learning-prediction process of the critic. The performance of Fast-AHC is also compared with that of the AHC method using LS-TD(lambda). Furthermore, it is demonstrated in the experiments that different initial values of the variance matrix in RLS-TD(lambda) are required to get better performance not only in learning prediction but also in learning control. The experimental results are analyzed based on the existing theoretical work on the transient phase of forgetting factor RLS methods. Xin Xu 0001, Hangen He, Dewen Hu |
J. Artif. Intell. Res. | 3 |