EDBT 2026 Demo / reviewers in the wild / expert
Zongtan Zhou
dblp:71/2970
· DBLP profile ↗
53ranked-venue papers
1as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 1 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decoding Driving Intentions via a Novel Brain-Computer Interface Paradigm With Low Cognitive Load and High Robustness
Jianxiang Sun, Zongtan Zhou, Yadong Liu 0001, Daxue Liu, Haoqiang Chen, Yingxin Liu, Dewen Hu |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2025 | Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling
Wenting Luan, Siqi Lu, Yongbin Zheng, Wanying Xu, Lang Nie, Zongtan Zhou, Kang Liao |
ICCV | 6 |
| 2025 | InsCMPR: Efficient Cross-Modal Place Recognition via Instance-Aware Hybrid Mamba-TransformerabstractPlace recognition is an important technique for autonomous mobile robotic applications. While single-modal sensor-based approaches have shown satisfactory performance, cross-modal place recognition remains underexplored due to the challenge of bridging the cross-modal heterogeneity gap. In this work, we introduce an instance-aware cross-modal place recognition approach, named InsCMPR. We design a novel instance-aware modality alignment module, which aligns multi-modal data at both pixel-level and instance-level by leveraging a pre-trained vision foundation model SAM. Then a novel dual-branch hybrid Mamba-Transformer network is proposed to efficiently enhance the distinctiveness of the produced descriptors by integrating global features with local instance features. Experimental results on the KITTI, NCLT, and HAOMO datasets show that our proposed methods achieve state-of-the-art performance while operating in real time. We will open source the implementation of our method at: https://github.com/nubot-nudt/InsCMPR. Shuaifeng Jiao, Zhuoqun Su, Lun Luo, Hongshan Yu, Zongtan Zhou, Huimin Lu 0002, Xieyuanli Chen |
ICRA | 5 |
| 2025 | LuSeg: Efficient Negative and Positive Obstacles Segmentation via Contrast-Driven Multi-Modal Feature Fusion on the LunarabstractAs lunar exploration missions grow increasingly complex, ensuring safe and autonomous rover-based surface exploration has become one of the key challenges in lunar exploration tasks. In this work, we have developed a lunar surface simulation system called the Lunar Exploration Simulator System (LESS) and the LunarSeg dataset, which provides RGB-D data for lunar obstacle segmentation that includes both positive and negative obstacles. Additionally, we propose a novel two-stage segmentation network called LuSeg. Through contrastive learning, it enforces semantic consistency between the RGB encoder from Stage I and the depth encoder from Stage II. Experimental results on our proposed LunarSeg dataset and additional public real-world NPO road obstacle dataset demonstrate that LuSeg achieves state-of-the-art segmentation performance for both positive and negative obstacles while maintaining a high inference speed of approximately 57 Hz. We have released the implementation of our LESS system, LunarSeg dataset, and the code of LuSeg at: https://github.com/nubot-nudt/LuSeg. Shuaifeng Jiao, Zhuoqun Su, Xieyuanli Chen, Zongtan Zhou, Huimin Lu 0002 |
IROS | 5 |
| 2025 | ResLPR: A LiDAR Data Restoration Network and Benchmark for Robust Place Recognition Against Weather CorruptionsabstractLiDAR-based place recognition (LPR) is a key component for autonomous driving, and its resilience to environmental corruption is critical for safety in high-stakes applications. While state-of-the-art (SOTA) LPR methods perform well in clean weather, they still struggle with weather-induced corruption commonly encountered in driving scenarios. To tackle this, we propose ResLPRNet, a novel LiDAR data restoration network that largely enhances LPR performance under adverse weather by restoring corrupted LiDAR scans using a wavelet transform-based network. ResLPRNet is efficient, lightweight and can be integrated plug-and-play with pretrained LPR models without substantial additional computational cost. Given the lack of LPR datasets under adverse weather, we introduce ResLPR, a novel benchmark that examines SOTA LPR methods under a wide range of LiDAR distortions induced by severe snow, fog, and rain conditions. Experiments on our proposed WeatherKITTI and WeatherNCLT datasets demonstrate the resilience and notable gains achieved by using our restoration method with multiple LPR approaches in challenging weather scenarios. Our code and benchmark are publicly available here: https://github.com/nubot-nudt/ResLPR. Wenqing Kuang, Xiongwei Zhao, Yehui Shen, Congcong Wen, Huimin Lu 0002, Zongtan Zhou, Xieyuanli Chen |
IROS | 6 |
| 2025 | Pessimistic policy iteration with bounded uncertaintyabstractOffline Reinforcement Learning (RL) aims to learn policies by using static datasets. The extrapolation error in out-of-distribution (OOD) samples can cause off-policy RL algorithms to perform poorly on offline datasets. Hence, it is critical to avoid visiting OOD states and taking OOD actions in offline RL. Several recent methods have used uncertainty estimation to distinguish OOD samples. However, errors in the uncertainty estimation make the purely uncertainty-based method unstable and require additional components to ensure sufficient pessimism . In this study, we propose a Bounded Uncertainty based Pessimistic policy iteration algorithm (BUP). The BUP pessimistically estimates the value function via bounded uncertainty, and the uncertainty bound is achieved by constraining the actor from taking highly uncertain actions. The suboptimality bound of BUP is theoretically guaranteed in linear Markov Decision Processes (MDPs), and experiments on D4RL datasets show that BUP matches the state-of-the-art performance. Moreover, BUP is simple to implement with low computational cost and does not require any additional components. Zhiyong Peng 0002, Changlin Han, Yadong Liu 0001, Jingsheng Tang, Zongtan Zhou |
Expert Syst. Appl. | 5 |
| 2025 | PhyTransformer: A unified framework for learning spatial-temporal representation from physiological signals
Yuke Qu, Jingsheng Tang, Zongtan Zhou |
Neural Networks | 7 |
| 2025 | sEMG-Based Gesture-Free Hand Intention Recognition: System, Dataset, Toolbox, and Benchmark ResultsabstractIn sensitive scenarios, such as meetings, negotiations, and team sports, messages must be conveyed without detection by noncollaborators. Previous methods, such as encrypting messages, eye contact, and micro-gestures, had problems with either inaccurate information transmission or leakage of interaction intentions. To this end, a novel gesture-free hand intention recognition scheme was proposed, that adopted surface electromyography (sEMG) and isometric contraction theory to recognize hand intentions without any gesture. Specifically, this work includes four aspects: first, the experimental system, consisting of the self-conducted myoelectric wristband, the matched host computer software, and the sports platform, is built to get sEMG signals and simulate multiple usage scenarios; second, the paradigm is designed to standard prompt and collect the gesture-free sEMG datasets. Eight-channel signals of ten subjects were recorded twice per subject at about 5–10 days intervals; third, the toolbox integrates preprocessing methods (data segmentation, filter, normalization, etc.), widely used sEMG classification methods, and various plotting functions, to facilitate future research based this dataset; fourth, the benchmark results of widely used methods are provided. The results involve single-day, cross-day, and cross-subject experiments of six-class and 12-class gesture-free hand intention when subjects have different time windows. Jingsheng Tang, Xuechao Xu, Wei Dai 0014, Junhao Xiao 0001, Huimin Lu 0002, Zongtan Zhou |
IEEE Trans. Ind. Informatics | 8 |
| 2025 | Cognitive Load Prediction From Multimodal Physiological Signals Using Multiview LearningabstractPredicting cognitive load is a crucial issue in the emerging field of human-computer interaction and holds significant practical value, particularly in flight scenarios. Although previous studies have realized efficient cognitive load classification, new research is still needed to adapt the current state-of-the-art multimodal fusion methods. Here, we proposed a feature selection framework based on multiview learning to address the challenges of information redundancy and reveal the common physiological mechanisms underlying cognitive load. Specifically, the multimodal signal features [electroencephalogram (EEG), electrodermal activity (EDA), electrocardiogram (ECG), electrooculogram (EOG), & eye movements] at three cognitive load levels were estimated during multiattribute task battery (MATB) tasks performed by 22 healthy participants and fed into a feature selection-multiview classification with cohesion and diversity (FS-MCCD) framework. The optimized feature set was extracted from the original feature set by integrating the weight of each view and the feature weights to formulate the ranking criteria. The cognitive load prediction model, evaluated using real-time classification results, achieved an average accuracy of 81.08% and an average F1-score of 80.94% for three-class classification among 22 participants. Furthermore, the weights of the physiological signal features revealed the physiological mechanisms related to cognitive load. Specifically, heightened cognitive load was linked to amplified $\delta$ and $\theta$ power in the frontal lobe, reduced $\alpha$ power in the parietal lobe, and an increase in pupil diameter. Thus, the proposed multimodal feature fusion framework emphasizes the effectiveness and efficiency of using these features to predict cognitive load. Yingxin Liu, Yang Yu 0014, Zeqi Ye, Hao Li 0086, Dewen Hu, Zongtan Zhou |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | Efficient-PIP: Large-scale Pixel-level Aligned Image Pair Generation for Cross-time Infrared-RGB TranslationabstractGenerative models are gaining momentum in both academic and industrial applications driven by the availability of large-scale datasets, especially in tasks involving Image-to-Image Translation. Meanwhile, poor human perception of nighttime environment has led to a demand for translation from night-vision infrared to day-vision RGB images. However, collecting such cross-modal training data at the same time is impossible due to the thermal imaging properties of infrared cameras, the challenge lies in constructing image pairs during the day and at night respectively, where the requirement for data alignment poses significant difficulties. In this paper, we propose a Pixel-level aligned Image Pair generation framework PIP to explore efficient colorization of high-resolution infrared images. Specifically, we first construct a 3D high-precision point cloud map for the purpose of establishing the correlation between day and night scenes. Corresponding point clouds of modal images are collected simultaneously during data acquisition to obtain image sensor poses by Global Matching with the map, which allows us to calculate the transformation relationship from infrared to RGB image coordinate systems based on the sensor parameters and depth information of the map. Leveraging the relationship, the pixel values of RGB image is projected onto the infrared image followed by optimization as the colored image. Accordingly, we present a dataset NUDT-PIP, the first of its kind containing large-scale pixel-level aligned cross-time infrared-RGB image pairs of complicated real road scenes. Experimental results demonstrate the reliability and strong applicability of our dataset in Image-to-Image Translation. Our code will be released at https://github.com/wjjjjyourFA/NUDT-PIP. Jian Li 0003, Kexin Fei, Bokai Liu, Zongtan Zhou, Yongbin Zheng, Zhenping Sun |
IROS | 6 |
| 2024 | TD-NeRF: Novel Truncated Depth Prior for Joint Camera Pose and Neural Radiance Field OptimizationabstractThe reliance on accurate camera poses is a significant barrier to the widespread deployment of Neural Radiance Fields (NeRF) models for 3D reconstruction and SLAM tasks. The existing method introduces monocular depth priors to jointly optimize the camera poses and NeRF, which fails to fully exploit the depth priors and neglects the impact of their inherent noise. In this paper, we propose Truncated Depth NeRF (TD-NeRF), a novel approach that enables training NeRF from unknown camera poses - by jointly optimizing learnable parameters of the radiance field and camera poses. Our approach explicitly utilizes monocular depth priors through three key advancements: 1) we propose a novel depth-based ray sampling strategy based on the truncated normal distribution, which improves the convergence speed and accuracy of pose estimation; 2) to circumvent local minima and refine depth geometry, we introduce a coarse-to-fine training strategy that progressively improves the depth precision; 3) we propose a more robust inter-frame point constraint that enhances robustness against depth noise during training. The experimental results on three datasets demonstrate that TD-NeRF achieves superior performance in the joint optimization of camera pose and NeRF, surpassing prior works, and generates more accurate depth geometry. The implementation of our method has been released at https://github.com/nubot-nudt/TD-NeRF. Zhen Tan 0002, Zongtan Zhou, Yangbing Ge, Xieyuanli Chen, Dewen Hu |
IROS | 2 |
| 2024 | CollOR: Distributed collaborative offloading and routing for tasks with QoS demands in multi-robot system
Huimin Lu 0002, Songtao Guo, Zongtan Zhou |
Ad Hoc Networks | 5 |
| 2024 | SyRoC: Symbiotic robotics for QoS-aware heterogeneous applications in IoT-edge-cloud computing paradigm
Huimin Lu 0002, Songtao Guo, Mingfang Ma, Zongtan Zhou |
Future Gener. Comput. Syst. | 6 |
| 2024 | Sigmoid distance metric-based spline adaptive filters for nonlinear adaptive noise cancellation
Wenqi Li 0003, Zongtan Zhou, Jingsheng Tang |
Inf. Sci. | 2 |
| 2024 | Deadly triad matters for offline reinforcement learning
Zhiyong Peng 0002, Yadong Liu 0001, Zongtan Zhou |
Knowl. Based Syst. | 3 |
| 2023 | Weighted Policy Constraints for Offline Reinforcement LearningabstractOffline reinforcement learning (RL) aims to learn policy from the passively collected offline dataset. Applying existing RL methods on the static dataset straightforwardly will raise distribution shift, causing these unconstrained RL methods to fail. To cope with the distribution shift problem, a common practice in offline RL is to constrain the policy explicitly or implicitly close to behavioral policy. However, the available dataset usually contains sub-optimal or inferior actions, constraining the policy near all these actions will make the policy inevitably learn inferior behaviors, limiting the performance of the algorithm. Based on this observation, we propose a weighted policy constraints (wPC) method that only constrains the learned policy to desirable behaviors, making room for policy improvement on other parts. Our algorithm outperforms existing state-of-the-art offline RL algorithms on the D4RL offline gym datasets. Moreover, the proposed algorithm is simple to implement with few hyper-parameters, making the proposed wPC algorithm a robust offline RL method with low computational complexity. Zhiyong Peng 0002, Changlin Han, Yadong Liu 0001, Zongtan Zhou |
AAAI | 4 |
| 2023 | Overfitting-avoiding goal-guided exploration for hard-exploration multi-goal reinforcement learning
Changlin Han, Zhiyong Peng 0002, Yadong Liu 0001, Jingsheng Tang, Yang Yu 0014, Zongtan Zhou |
Neurocomputing | 6 |
| 2023 | Conservative network for offline reinforcement learning
Zhiyong Peng 0002, Yadong Liu 0001, Haoqiang Chen, Zongtan Zhou |
Knowl. Based Syst. | 4 |
| 2023 | Fusion of Spatial, Temporal, and Spectral EEG Signatures Improves Multilevel Cognitive Load PredictionabstractCognitive load prediction is one of the most important issues in the nascent field of neuroergonomics, and it has significant value in real-world applications. Most of the previous studies of cognitive load prediction only utilized electroencephalography (EEG)-based spectral signatures or interchannel connectivity, ignoring abundant temporal microstate features, which may represent the transient topologies of EEG signals. Furthermore, previous studies have mostly focused on the binary-level classification of cognitive load for single-type cognitive tasks. To date, there are few studies on the multilevel prediction of cognitive load during mixed cognitive tasks. Here, we first designed a new paradigm termed the “finding fault game,” mixing multiple tasks of memory, counting, and visual search, and then developed a multidimensional analysis framework to improve cognitive load prediction using a fusion of spatial, temporal, and spectral EEG features. Specifically, EEG-based functional connectivity, microstates and power spectral densities (PSD) were calculated for three cognitive load levels. Twelve adult subjects participated in the study. The experimental results show that increased cognitive load was associated with elevated theta and degraded alpha power and significant changes in interchannel connectivity and microstates, and that fusing the three types of EEG features improved the performance of three-level cognitive load prediction, achieving the accuracies of greater than 80% in the cross-validation, real-time, and over-time prediction. The findings suggest that all three types of EEG features can serve as signatures of cognitive load and that their fusion can improve multilevel prediction. Yingxin Liu, Yang Yu 0014, Zeqi Ye, Ming Li 0028, Zongtan Zhou, Dewen Hu |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2022 | Robot navigation in a crowd by integrating deep reinforcement learning and online planning
Zhiqian Zhou, Pengming Zhu, Junhao Xiao 0001, Huimin Lu 0002, Zongtan Zhou |
Appl. Intell. | 6 |
| 2022 | HQA-Trans: An end-to-end high-quality-awareness image translation framework for unsupervised cross-domain pedestrian detectionabstractAbstract Unsupervised cross‐domain pedestrian detection has attracted attention in recent years. Although some works adopted unsupervised image translation frameworks to generate an intermediate domain to narrow the gap between source and target domains, the images in the intermediate domain tend to be distorted due to the instability of the generation network. In this work, we propose a new framework to improve the image quality of the generated intermediate domain via an end‐to‐end translation framework. First, an image quality assessment index is adopted and adjusted appropriately. The part that controls the image quality is kept, and the part that adversely affects the domain style translation is discarded. Secondly, the adjusted image quality assessment index is integrated into the unsupervised image translation framework, where a new loss with the index's weight is proposed. An end‐to‐end high‐quality‐awareness image translation framework is constructed to generate a high‐quality intermediate domain directly through this process. Finally, the intermediate domain with high‐quality images is applied for cross‐domain pedestrian detection. Experimental results on benchmark datasets show that the proposed framework can effectively improve unsupervised cross‐domain pedestrian detection performance. Compared with some state‐of‐the‐art works, the proposed framework can also achieve superior performance under miss rate metrics. Gelin Shen, Haoqiang Chen, Zongtan Zhou |
IET Comput. Vis. | 5 |
| 2022 | Navigating Robots in Dynamic Environment With Deep Reinforcement LearningabstractIn the fight against COVID-19, many robots replace human employees in various tasks that involve a risk of infection. Among these tasks, the fundamental problem of navigating robots among crowds, named robot crowd navigation, remains open and challenging. Therefore, we propose HGAT-DRL, a heterogeneous GAT-based deep reinforcement learning algorithm. This algorithm encodes the constrained human-robot-coexisting environment in a heterogeneous graph consisting of four types of nodes. It also constructs an interactive agent-level representation for objects surrounding the robot, and incorporates the kinodynamic constraints from the non-holonomic motion model into the deep reinforcement learning (DRL) framework. Simulation results show that our proposed algorithm achieves a success rate of 92%, at least 6% higher than four baseline algorithms. Furthermore, the hardware experiment on a Fetch robot demonstrates our algorithm’s successful and convenient migration to real robots. Zhiqian Zhou, Lin Lang 0001, Weijia Yao, Huimin Lu 0002, Zhiqiang Zheng 0002, Zongtan Zhou |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | Giant magneto-impedance sensor with working point selfadaptation for unshielded human bio-magnetic detectionabstractCompared with traditional biomagnetic field detection devices, such as superconducting quantum interference devices (SQUIDs) and atomic magnetometers, only giant magnetoimpedance (GMI) sensors can be applied for unshielded human brain biomagnetic detection, and they have the potential for application in next-generation wearable equipment for brain-computer interfaces (BCIs). Achieving a better GMI sensor without magnetic shielding requires the stimulation of the GMI effect to be maximized and environmental noise interference to be minimized. Moreover, the GMI effect stimulated in an amorphous filament is closely related to its working point, which is sensitive to both the external magnetic field and the drive current of the filament. In this paper, we propose a new noisereducing GMI gradiometer with a dual-loop self-adapting structure. Noise reduction is realized by a direction-flexible differential probe, and the dual-loop structure optimizes and stabilizes the working point by automatically controlling the external magnetic field and drive current. This dual-loop structure is fully program controlled by a micro control unit (MCU), which not only simplifies the traditional constantparameter sensor circuit, saving the time required to adjust the circuit component parameters, but also improves the sensor performance and environmental adaptation. In the performance test, within 2 min of self-adaptation, our sensor showed a better sensitivity and signal-to-noise ratio (SNR) than those of the traditional designs and achieved a background noise of 12 pT/√Hz at 10 Hz and 7pT/√Hz at 200 Hz. To the best of our knowledge, our sensor is the first to realize self-adaptation of both the external magnetic field and the drive current. Changlin Han, Ming Xu 0022, Jingsheng Tang, Yadong Liu 0001, Zongtan Zhou |
Virtual Real. Intell. Hardw. | 5 |
| 2021 | A Virtual Mouse Based on Parallel Cooperation of Eye Tracker and Motor Imagery
Zeqi Ye, Yingxin Liu, Yang Yu 0014, Zongtan Zhou, Fengyu Xie |
ICIG (3) | 5 |
| 2021 | Brain-computer interface for human-multirobot strategic consensus with a differential world model
Wei Dai 0014, Huimin Lu 0002, Yadong Liu 0001, Zongtan Zhou |
Appl. Intell. | 5 |
| 2020 | GAMA: Graph Attention Multi-agent reinforcement learning algorithm for cooperation
Haoqiang Chen, Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Ming Zhang 0027 |
Appl. Intell. | 3 |
| 2020 | A Dynamic User Interface Based BCI Environmental Control SystemabstractIn this study, a dynamic user interface (UI) is proposed in visual P300 Brain-Computer Interface (BCI) based environmental control system. A head-mounted Augmented Reality (AR) glass is used as the interactive media, which is used to assists the BCI system to build the dynamic UI with the scene in subject’s field of view. In the dynamic UI, based on the objects detected by the AR glass, options are dynamically generated. The subject can assign tasks by selecting different options in the dynamic UI. Five subjects successfully completed the task of controlling household appliances and navigating wheelchairs to designated destinations. Compared to static UI, the proposed dynamic UI has a 17.4% improvement in time delay. On average, only 1.9% of the commands resulted in incorrect operations. The dynamic UI makes progress in reducing time delay and incorrect operations. The proposed system provides a brand-new interactive method in BCI based applications. Saisai Zhong, Yadong Liu 0001, Yang Yu 0014, Jingsheng Tang, Zongtan Zhou, Dewen Hu |
Int. J. Hum. Comput. Interact. | 5 |
| 2020 | R4 Det: Refined single-stage detector with feature recursion and refinement for rotating object detection in aerial images
Yongbin Zheng, Zongtan Zhou, Wanying Xu |
Image Vis. Comput. | 3 |
| 2019 | An Tactile ERP-Based Brain-Computer Interface for CommunicationabstractA classical visual event relative potential (ERP) brain–computer interface (BCI) system relies on visual stimuli to choose commands. Users obtain most information about their surroundings visually as well. This large amount of information can aggravate visual burden and fatigue. In our study, we proposed a novel approach to evoke ERP with a tactile stimulus. To achieve this approach, we first designed a wireless stimulus module with vibrators to provide a tactile stimulus for the system. The vibrators were located on the subject’s arm to imitate the joint motion of a robotic arm. Then, the ERP feature and the parameters of classifiers were obtained through offline experimental data analysis. Based on the analysis, the suitable electrode channels, stimulus onset asynchrony (SOA), and filter upper limit were different for different subjects. According to those outcomes, a unique classifier was designed for each subject. Finally, 10 healthy BCI-naive subjects participated in online experiments to evaluate the performance of our tactile BCI system; they achieved an accuracy range from 78.67% to 100% with an average of 89.1% and an instantaneous transmission rate (ITR) range from 7.77 to 28.70 bits/min with an average of 14.77 bits/min. The accuracy of different subjects and SOAs remained relatively stable, the ITR fluctuated mainly due to the different SOAs, and we achieved balance between ITR and accuracy. Yadong Liu 0001, Jingjun Wang, Erwei Yin, Yang Yu 0014, Zongtan Zhou, Dewen Hu |
Int. J. Hum. Comput. Interact. | 5 |
| 2019 | Toward Brain-Actuated Mobile PlatformabstractThis study presents a brain–computer interface (BCI) system aimed at providing disabled patients with mobile solutions for practical use. The proposed system employs an omnidirectional chassis and a bionic robot arm to construct a multi-functional mobile platform. In addition, the system is equipped with a Kinect and 12 ultrasonic sensors to capture environment information. Based on artificial intelligence technology, the mobile system can understand the environment and smartly completes certain tasks. A hybrid BCI combined with movement imagery paradigm and asynchronous P300 paradigm is designed to translate human intent to computer commands. The users interact with the system in a flexible way: on the one hand, the user issues commands to drive the system directly; on the other hand, the system searches for predefined operable targets and reports the results to the user. Once the user confirms the target, the system will automatically complete the associated operation. To evaluate the system’s performance, a testing environment with a small room, aisle, and an elevator was built to simulate the mobile tasks in the daily scene. Participants were instructed to operate the mobile system in the room, aisle, and using the elevator to go outdoors. In this study, four subjects participated in the test, and all of them completed the task. Jingsheng Tang, Yadong Liu 0001, Jun Jiang 0001, Yang Yu 0014, Dewen Hu, Zongtan Zhou |
Int. J. Hum. Comput. Interact. | 6 |
| 2019 | Towards a Hybrid BCI Gaming Paradigm Based on Motor Imagery and SSVEPabstractBrain-computer interfaces (BCIs) not only can allow individuals to voluntarily control external devices, helping to restore lost motor functions of the disabled, but can also be used by healthy users for entertainment and gaming applications. In this study, we proposed a hybrid BCI paradigm to explore a feasible and natural way to play games by using electroencephalogram (EEG) signals in a practical environment. In this paradigm, we combined motor imagery (MI) and steady-state visually evoked potentials (SSVEPs) to generate multiple commands. A classic game, Tetris, was chosen as the control object. The novelty of this study includes the effective usage of a “dwell time” approach and fusion rules to design BCI games. To demonstrate the feasibility of the proposed hybrid paradigm, ten subjects were chosen to participate in online control experiments. The experimental results showed that all subjects successfully completed the predefined tasks with high accuracy. This proposed hybrid BCI paradigm could potentially provide those who suffer disability or paralysis with additional entertainment options, such as brain-actuated games, that could improve their happiness and quality of life.Abbreviations: BCI: brain-computer interface; EEG: electroencephalogram; MI: motor imagery; SSVEP: steady-state visually evoked potential; ERP: event-related potential; SMR: sensorimotor rhythm; VEP: visual evoked potential; TCP/IP: transmission control protocol/internet protocol; GUI: graphical user interface; ERD/ERS: event-related desynchronization/synchronization; CIC: control intention classifier; LRC: left/right classifier; CSP: common spatial pattern; LDA: linear discriminant analysis; ROC: receiver operating characteristic; TPR: true positive rate; FPR: false positive rate; CCA: canonical correlation analysis. Zhihua Wang 0002, Yang Yu 0014, Ming Xu 0022, Yadong Liu 0001, Erwei Yin, Zongtan Zhou |
Int. J. Hum. Comput. Interact. | 6 |
| 2017 | Detect visual field using eye tracking and steady-state visual evoked potentialabstractThis paper makes the subjects' sight locked in a certain area using an eye tracker, getting Steady-state visual evoked potential (SSVEP) from flickering stimuli with a fixed frequency but at random positions, in order to observe the impact of stimulus at different positions and their distances on the electroencephalogram (EEG). The result suggests that if human have to select the positions of stimuli of SSVEP-BCI, it is an agreeable strategy to separate them at least 4 ° for avoiding the possible mistakes. We hope that it could help in setting distances between stimuli or updating pattern selection algorithms in the future BCI system and other paradigms. Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Erwei Yin |
SMC | 3 |
| 2017 | Toward a Hybrid BCI: Self-Paced Operation of a P300-based Speller by Merging a Motor Imagery-Based "Brain Switch" into a P300 Spelling ApproachabstractThis study presents the self-paced operation of a brain–computer interface (BCI) speller, which can be voluntarily turned on/off by merging a motor imagery (MI)-based brain switch into a P300-based BCI speller. From an off state (idle state), the users can generate a “control signal” by consciously changing the cognitive state differential from the idle state to turn on a P300-based spelling system when he or she wants to spell words. With the system turned on, the user can spell words, and then, the spelling system can be voluntarily turned off and switched to the initial state using a command. In this paradigm, the participants tried to perform the two different cognitive tasks sequentially, rather than simultaneously, and multiple EEG components were processed sequentially. The practicability and effectiveness of the proposed approach were validated by eleven participants, and all of them achieved a satisfactory performance. For the P300 speller, they achieved an average PITR of 42.61 bits/min. The preliminary results indicated that the proposed hybrid BCI system with different mental strategies operating sequentially is feasible and has potential applications for practical self-paced control. Yang Yu 0014, Zongtan Zhou, Jun Jiang 0001, Erwei Yin, Kunjia Liu, Jingjun Wang, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 2 |
| 2017 | Salient Region Detection Using Diffusion Process on a Two-Layer Sparse GraphabstractDiffusion-based salient region detection has recently received intense research attention. In this paper, we present some effective improvements concerning two important aspects of diffusion-based methods: the construction of the diffusion matrix and the seed vector. First, we construct a two-layer sparse graph, which is generated by connecting each node to its neighboring nodes and the most similar node that shares common boundaries with its neighboring nodes. Compared with the most frequently used two-layer neighborhood graph, our graph not only effectively uses local spatial relationships, but also removes dissimilar redundant nodes. Second, we use the spatial variance of superpixel clusters to obtain the seed vector and, compared with the previously most-used boundary prior, our approach can better distinguish saliency seeds from the background seeds, especially when salient objects appear near the image boundaries. Finally, we calculate two preliminary saliency maps using the saliency and background seed vectors, and more accurate results are obtained using the manifold ranking diffusion method. Integrating these two diffusion-based saliency maps, we obtain the final saliency map. Extensive experiments in which we compare our method with 20 existing state-of-the-art methods on five benchmark data sets: ASD, DUT-OMRON, ECSSD, MSRA5K, and MSRA10K, show that the proposed method performs better in terms of various evaluation metrics. Li Zhou 0003, Zhaohui Yang 0008, Zongtan Zhou, Dewen Hu |
IEEE Trans. Image Process. | 3 |
| 2016 | A P300-Based Brain-Computer Interface for Chinese Character InputabstractThe majority of previously developed assistive communication brain–computer interface systems have primarily focused on languages that are written in alphabetic scripts. However, languages that are written in logographic scripts, such as those in Chinese hanzi (or sinograms), pose a challenge for the implementation of visual spelling systems because it is impossible to simultaneously display thousands of items in a stimulus matrix of a reasonable size. In this study, a P300 visual spelling system that uses a novel method to input Chinese sinograms developed with a Hanyu Pinyin-based method is presented. This method transcribes a Chinese Pinyin into initial consonant and vowel components according to its Mandarin pronunciation. In this paradigm, each sinogram is input by selecting the initial consonant and then the vowel components and subsequently selecting the sinogram itself. Ten healthy subjects participated in the study and achieved an average offline accuracy of 92.6% with a mean information transfer rate of 39.2 bits/min and an average online input speed of one sinogram per 43.9 s. The preliminary results presented here indicated that the online input of Chinese text using a Pinyin-based visual speller is feasible. Yang Yu 0014, Zongtan Zhou, Erwei Yin, Jun Jiang 0001, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 2 |
| 2016 | An Auditory-Tactile Visual Saccade-Independent P300 Brain-Computer InterfaceabstractMost P300 event-related potential (ERP)-based brain-computer interface (BCI) studies focus on gaze shift-dependent BCIs, which cannot be used by people who have lost voluntary eye movement. However, the performance of visual saccade-independent P300 BCIs is generally poor. To improve saccade-independent BCI performance, we propose a bimodal P300 BCI approach that simultaneously employs auditory and tactile stimuli. The proposed P300 BCI is a vision-independent system because no visual interaction is required of the user. Specifically, we designed a direction-congruent bimodal paradigm by randomly and simultaneously presenting auditory and tactile stimuli from the same direction. Furthermore, the channels and number of trials were tailored to each user to improve online performance. With 12 participants, the average online information transfer rate (ITR) of the bimodal approach improved by 45.43% and 51.05% over that attained, respectively, with the auditory and tactile approaches individually. Importantly, the average online ITR of the bimodal approach, including the break time between selections, reached 10.77 bits/min. These findings suggest that the proposed bimodal system holds promise as a practical visual saccade-independent P300 BCI. Erwei Yin, Timothy J. Zeyl, Rami Saab, Dewen Hu, Zongtan Zhou, Tom Chau |
Int. J. Neural Syst. | 5 |
| 2015 | Salient Region Detection via Integrating Diffusion-Based Compactness and Local ContrastabstractSalient region detection is a challenging problem and an important topic in computer vision. It has a wide range of applications, such as object recognition and segmentation. Many approaches have been proposed to detect salient regions using different visual cues, such as compactness, uniqueness, and objectness. However, each visual cue-based method has its own limitations. After analyzing the advantages and limitations of different visual cues, we found that compactness and local contrast are complementary to each other. In addition, local contrast can very effectively recover incorrectly suppressed salient regions using compactness cues. Motivated by this, we propose a bottom-up salient region detection method that integrates compactness and local contrast cues. Furthermore, to produce a pixel-accurate saliency map that more uniformly covers the salient objects, we propagate the saliency information using a diffusion process. Our experimental results on four benchmark data sets demonstrate the effectiveness of the proposed method. Our method produces more accurate saliency maps with better precision-recall curve and higher F-Measure than other 19 state-of-the-arts approaches on ASD, CSSD, and ECSSD data sets. Li Zhou 0003, Zhaohui Yang 0008, Zongtan Zhou, Dewen Hu |
IEEE Trans. Image Process. | 4 |
| 2013 | Scene recognition combining structural and textural features
Li Zhou 0003, Dewen Hu, Zongtan Zhou |
Sci. China Inf. Sci. | 3 |
| 2013 | Scene classification using multi-resolution low-level feature combination
Li Zhou 0003, Zongtan Zhou, Dewen Hu |
Neurocomputing | 2 |
| 2013 | Scene classification using a multi-resolution bag-of-features model
Li Zhou 0003, Zongtan Zhou, Dewen Hu |
Pattern Recognit. | 2 |
| 2009 | Local region structured noise reduction for cortical optical imaging
Yadong Liu 0001, Dewen Hu, Zongtan Zhou, Fayi Liu |
Neurocomputing | 3 |
| 2008 | A Direct Locality Preserving Projections (DLPP) Algorithm for Image Recognition
Guiyu Feng, Dewen Hu, Zongtan Zhou |
Neural Process. Lett. | 3 |
| 2008 | Globally Consistent Reconstruction of Ripped-Up DocumentsabstractOne of the most crucial steps for automatically reconstructing ripped-up documents is to find a globally consistent solution from the ambiguous candidate matches. However, little work has been done so far to solve this problem in a general computational framework without using application-specific features. In this paper, we propose a global approach for reconstructing ripped-up documents by first finding candidate matches from document fragments using curve matching and then disambiguating these candidates through a relaxation process to reconstruct the original document. The candidate disambiguation problem is formulated in a relaxation scheme, in which the definition of compatibility between neighboring matches is proposed and global consistency is defined as the global criterion. Initially, global match confidences are assigned to each of the candidate matches. After that, the overall local relationships among neighboring matches are evaluated by computing their global consistency. Then these confidences are iteratively updated using the gradient projection method to maximize the criterion. This leads to a globally consistent solution and thus provides a sound document reconstruction. The overall performance of our approach in several practical experiments is illustrated. The results indicate that the reconstruction of ripped-up documents up to fifty pieces is possibly accomplished automatically. Liangjia Zhu, Zongtan Zhou, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Comment on "two-dimensional locality preserving projections (2DLPP) with its application to palmprint recognition"
Dewen Hu, Guiyu Feng, Zongtan Zhou |
Pattern Recognit. | 3 |
| 2008 | Noisy manifold learning using neighborhood smoothing embedding
Junsong Yin, Dewen Hu, Zongtan Zhou |
Pattern Recognit. Lett. | 3 |
| 2007 | Manifold Learning using Growing Locally Linear EmbeddingabstractLocally linear embedding (LLE) is an effective nonlinear dimensionality reduction method for exploring the intrinsic characteristics of high dimensional data. This paper mainly proposes a hierarchical framework manifold learning method, based on LLE and growing neural gas (GNG), named growing locally linear embedding (GLLE). First, we address the major limitations of the original LLE: intrinsic dimensionality estimation, neighborhood number selection and computational complexity. Then by embedding the topology learning mechanism in GNG, the proposed GLLE algorithm is able to preserve the global topological structures and hold the geometric characteristics of the input patterns, which make the projections more stable and robust. Theoretical analysis and experimental simulations show that GLLE with global topology preservation tackles the three limitations, gives faster learning procedure and lower reconstruction error, and stimulates the wide applications of manifold learning Junsong Yin, Dewen Hu, Zongtan Zhou |
CIDM | 3 |
| 2007 | Two-dimensional locality preserving projections (2DLPP) with its application to palmprint recognition
Dewen Hu, Guiyu Feng, Zongtan Zhou |
Pattern Recognit. | 3 |
| 2006 | Classification of Movement-Related Potentials for Brain-Computer Interface: A Reinforcement Training Approach
Zongtan Zhou, Dewen Hu |
ISNN (2) | 1 |
| 2006 | An alternative formulation of kernel LPP with application to image recognition
Guiyu Feng, Dewen Hu, David Zhang 0001, Zongtan Zhou |
Neurocomputing | 4 |
| 2005 | A novel method for spatio-temporal pattern analysis of brain fMRI dataabstractA novel data processing procedure for fMRI was suggested in this paper, by which spatial and temporal characteristics of stimuli-induced signal dynamic responses can be investigated simultaneously. First the multitaper spectral estimation was utilized to estimate the spectrum of each voxel; the significance of the line frequency components at the interested frequency was tested to detect the task-related cortex areas; the temporal independent component analysis (tICA) was then applied to the activated voxels to obtain stimuli-induced signal dynamic responses. The advantages of this procedure are: few assumptions are needed for the cerebral hemodynamics and spatial distribution of task-related areas, problems which often appear in tICA analysis of fMRI data, such as the lack of stability, reliability and robustness, are overcome by the suggested method. Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Lirong Yan, Changlian Tan, Daxing Wu, Shuqiao Yao |
Sci. China Ser. F Inf. Sci. | 2 |
| 2005 | DSOM: a novel self-organizing model based on NO dynamic diffusing mechanism
Junsong Yin, Dewen Hu, Zongtan Zhou |
Sci. China Ser. F Inf. Sci. | 4 |
| 2004 | Diffusion and Growing Self-Organizing Map: A Nitric Oxide Based Neural Model
Zongtan Zhou, Dewen Hu |
ISNN (1) | 2 |
| 2004 | A New Computational Model of Biological Vision for Stereopsis
Baoquan Song, Zongtan Zhou, Dewen Hu, Zhengzhi Wang |
ISNN (2) | 2 |