EDBT 2026 Demo / reviewers in the wild / expert
Erwei Yin
dblp:138/2127
· DBLP profile ↗
63ranked-venue papers
1as first author
54since 2021 · last 2027
0000-0002-2147-9888ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 1 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 24 since 2021Human-computer interaction and ubiquitous computing · 10 · 5 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CMCL-Net: A cross-modal contrastive learning network for micro-expression spotting enhanced by electrocardiogram signals
Zhongkai Ma, Jianhang Zhang, Shaokai Zhao, Liang Xie 0012, Erwei Yin |
Expert Syst. Appl. | 8 |
| 2026 | MHED-SLAM: Multi-Scale Hybrid Encoding-Based Decoupled SLAMabstractNeural Radiance Fields (NeRF)-based Visual Simultaneous Localization and Mapping (SLAM) achieve superior scene geometric modeling and robust camera tracking by leveraging neural representations. Existing methods typically relied on multi-resolution hash encoding with truncated signed distance fields (TSDF) to achieve high frame rates. However, unavoidable hash collisions can lead to artifacts, and multi-view color inconsistencies in indoor scenes can result in shape-radiance ambiguity, adversely affecting geometric quality and tracking accuracy. To address these issues, we propose a novel Multi-scale Hybrid Encoding-based Decoupled SLAM (MHED-SLAM). First, to mitigate the adverse effects of hash collisions and reduce the number of learnable parameters, we innovatively fuse a coarse-scale hash tri-plane with a fine-scale hash grid within a single latent volume. Second, to enable precise geometric reconstruction and camera tracking, we decouple the reconstruction and rendering processes, independently learning a TSDF field for reconstruction and a density field for rendering. Third, we devise a Symmetric Kullback-Leibler (SKL) strategy based on ray termination distributions to align the probability distributions derived from the TSDF and density fields for their synchronous convergence. Extensive experimental evaluations demonstrate that our approach surpasses the state-of-the-art (SOTA) methods by utilizing a faster frame rate of 20 Hz and fewer parameters, while achieving higher tracking and reconstruction accuracy. Dengfang Feng, Wenyang Qin, Zhongchen Shi, Wei Chen 0092, Yanhui Duan, Liang Xie 0012, Erwei Yin |
AAAI | 7 |
| 2026 | Self-Enhanced Image Clustering with Cross-Modal Semantic ConsistencyabstractWhile large language-image pre-trained models like CLIP offer powerful generic features for image clustering, existing methods typically freeze the encoder. This creates a fundamental mismatch between the model's task-agnostic representations and the demands of a specific clustering task, imposing a ceiling on performance. To break this ceiling, we propose a self-enhanced framework based on cross-modal semantic consistency for efficient image clustering. Our framework first builds a strong foundation via Cross-Modal Semantic Consistency and then specializes the encoder through Self-Enhancement. In the first stage, we focus on Cross-Modal Semantic Consistency. By mining consistency between generated image-text pairs at the instance, cluster assignment, and cluster center levels, we train lightweight clustering heads to align with the rich semantics of the pre-trained model. This alignment process is bolstered by a novel method for generating higher-quality cluster centers and a dynamic balancing regularizer to ensure well-distributed assignments. In the second stage, we introduce a Self-Enhanced fine-tuning strategy. The well-aligned model from the first stage acts as a reliable pseudo-label generator. These self-generated supervisory signals are then used to feed back the efficient, joint optimization of the vision encoder and clustering heads, unlocking their full potential. Extensive experiments on six mainstream datasets show that our method outperforms existing deep clustering methods by significant margins. Notably, our ViT-B/32 model already matches or even surpasses the accuracy of state-of-the-art methods built upon the far larger ViT-L/14. Jianhua Yin 0001, Erwei Yin, Jianlong Wu |
AAAI | 6 |
| 2026 | DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement
Yilei Wu, Changyan Zheng, Yakun Zhang 0002, Chengshi Zheng, Ye Yan 0001, Erwei Yin |
Appl. Intell. | 8 |
| 2026 | Locomotion in CAVE: Enhancing immersion through full-body motion
Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Meng Gai, Suxia Zhang, Erwei Yin |
Comput. Graph. | 9 |
| 2026 | DAP-Whisper: A robust audio-visual speech recognition system via distribution-aware prompting and consistency-gated modulation
Yakun Zhang 0002, Changyan Zheng, Liang Xie 0012, Jiangbin Zheng 0001, Erwei Yin |
Expert Syst. Appl. | 8 |
| 2026 | DC-AVSR: A dynamic collaborative approach for robust audio-visual speech recognition under multimodal distortions
Yakun Zhang 0002, Changyan Zheng, Liang Xie 0012, Jiangbin Zheng 0001, Erwei Yin |
Neurocomputing | 7 |
| 2026 | Sequential viseme-driven visual speech recognition through dual-stream interactive neural architecture
Yakun Zhang 0002, Changyan Zheng, Liang Xie 0012, Erwei Yin |
Neural Networks | 6 |
| 2026 | MPFNet: A Multi-Prior Fusion Network With a Progressive Training Strategy for Micro-Expression RecognitionabstractMicro-expression recognition (MER), a critical subfield of affective computing, presents greater challenges than macro-expression recognition due to its brief duration and low intensity. While incorporating prior knowledge has been shown to enhance MER performance, existing methods predominantly rely on simplistic, singular sources of prior knowledge, failing to fully exploit multi-source information. This paper introduces the Multi-Prior Fusion Network (MPFNet), leveraging a progressive training strategy to optimize MER tasks. We propose two complementary encoders: the Generic Feature Encoder (GFE) and the Advanced Feature Encoder (AFE), both based on Inflated 3D ConvNets (I3D) with Coordinate Attention (CA) mechanisms, to improve the model's ability to capture spatiotemporal and channel-specific features. Inspired by developmental psychology, we present two variants of MPFNet—MPFNet-P and MPFNet-C—corresponding to two fundamental modes of infant cognitive development: parallel and hierarchical processing. These variants enable the evaluation of different strategies for integrating prior knowledge. Extensive experiments demonstrate that MPFNet significantly improves MER accuracy while maintaining balanced performance across categories, achieving accuracies of 0.811, 0.924, and 0.857 on the SMIC, CASME II, and SAMM datasets, respectively. To the best of our knowledge, our approach achieves state-of-the-art performance on the SMIC and SAMM datasets. The source code is available at:https://github.com/Mac0504/MPFNet. Shaokai Zhao, Dongdong Zhou, Zhiguo Luo, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
IEEE Trans. Affect. Comput. | 8 |
| 2026 | KDB-Gaze: Keypoint-Guided Dual-Branch Learning for Gaze EstimationabstractNear-eye gaze estimation has been a core technology for natural interaction in smart head-mounted devices. Existing gaze estimation approaches have the following limitations: (1) Data-driven methods learn nonlinear mappings directly from images to gaze direction but ignore the physiological characteristics of the iris–pupil system. (2) Model-driven methods rely on hand-designed modeling and perform poorly in complex, noisy scenes. (3) Hybrid-driven methods attempt to combine both paradigms but are restricted by simplified model designs. To address these issues, we propose KDB-Gaze, aKeypoint-GuidedDual-Branch learning for near-eyeGazeestimation. First, a backbone and a 3-Pass path aggregation feature pyramid network effectively capture multi-scale ocular features. Then, two parallel branches refine the features: an implicit spatial modeling branch employs deformable DETR to capture nonlinear features, while an explicit topological constraint branch uses a graph attention network to provide stable physiological guidance. Finally, the dual-branch features are fused and optimized jointly with a multi-task loss for precise prediction. Extensive experiments on the TEyeD, LPW, and NVGaze datasets verify the state-of-the-art performance of the proposed method. Xiaowei Bai, Liang Xie 0012, Qining Wang, Erwei Yin |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | Progressive feature-space alignment for pose-controllable virtual try-on
Erwei Yin |
Vis. Comput. | 2 |
| 2025 | De^2Gaze: Deformable and Decoupled Representation Learning for 3D Gaze Estimationabstract3D Gaze estimation is a challenging task due to two main issues. First, existing methods focus on analyzing dense features (e.g., large pixel regions), which are sensitive to local noise (e.g., light spots, blurs) and result in increased computational complexity. Second, an eyeball model can correspond multiple gaze directions, and the entangled representation between gazes and models increases the learning difficulty. To address these issues, we propose De2Gaze, a lightweight and accurate model-aware 3D gaze estimation method. In De2Gaze, we introduce two key innovations for deformable and decoupled representation learning. Specifically, first, we propose a deformable sparse attention mechanism that can adapt sparse sampling points to attention areas to avoid local noise influences. Second, we propose a spatial decoupling network with a dual-branch decoding architecture to disentangle invariant (e.g., eyeball radius, position) and variable (e.g., gaze, pupil, iris) features from the latent space. Compared to existing methods, De2Gaze requires fewer sparse features, and achieves faster convergence speed, lower computational complexity, and higher accuracy in 3D gaze estimation. Qualitative and quantitative experiments demonstrate that De2Gaze achieves state-of-the-art accuracy and high-quality semantic segmentation for 3D gaze estimation on the TEyeD dataset. Yunfeng Xiao, Xiaowei Bai, Baojun Chen, Liang Xie 0012, Erwei Yin |
CVPR | 7 |
| 2025 | LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech RecognitionabstractVisual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have significantly enhanced the performance of lip reading models. Despite these advancements, existing datasets predominantly feature stable video recordings with limited variability in lip movements. This limitation results in models that are highly sensitive to variations encountered in real-world scenarios. To address this issue, we propose a novel framework, LipGen, which aims to improve model robustness by leveraging speech-driven synthetic visual data, thereby mitigating the constraints of current datasets. Additionally, we introduce an auxiliary task that incorporates viseme classification alongside attention mechanisms. This approach facilitates the efficient integration of temporal information, directing the model’s focus toward the relevant segments of speech, thereby enhancing discriminative capabilities. Our method demonstrates superior performance compared to the current state-of-the-art on the lip reading in the wild (LRW) dataset and exhibits even more pronounced advantages under challenging conditions. Dongliang Zhou, Liang Xie 0012, Jianlong Wu, Erwei Yin |
ICASSP | 7 |
| 2025 | A Multi-Prior Fusion Network for Video-based Micro-Expression RecognitionabstractThe analysis of facial micro-expressions (MEs) has emerged as a significant application and topic within the field of image and video processing. However, challenges persist due to the brief duration and subtle intensity of these spontaneous expressions. This paper presents a novel multi-prior fusion network (MPFNet) for ME recognition based on a progressive training strategy. During the prior learning phase, our model is trained using a dual-stream architecture to capture both generic and advanced ME features. In the classification phase, we merge two pre-trained models with complementary prior knowledge and employ weighted fusion for classification within a meta-learning framework. Additionally, this study employs inflated 3D ConvNets (I3D) as a feature encoder and integrates Coordinate Attention (CA) blocks to enhance the automatic learning of spatiotemporal and channel features of ME video sequences. Extensive experiments conducted on three benchmark datasets validate the effectiveness of our model. Shaokai Zhao, Liang Xie 0012, Erwei Yin, Ye Yan 0001 |
ICASSP | 5 |
| 2025 | Self-supervised Contrastive Pre-training for Dry Electrode EEG Emotion Recognition via Cross Device Representation ConsistencyabstractThe use of dry electrode electroencephalography (EEG) systems holds significant importance in advancing the everyday application of emotion recognition. However, adapting it to real-world applications faces unique challenges due to low signal-to-noise ratios and unreliable emotion labels. To address these challenges, we propose a Cross-Device Representation Consistency (CDRC) pre-training paradigm for dry EEG emotion recognition, where the self-supervised signal is provided by the distance between representations embedded in wet and dry EEG components and trained via contrastive estimation. Specifically, we employ a dual-branch embedding prediction task coupled with contrastive feature alignment module to extract robust and distinctive features from dry electrode EEG signals. We evaluate our model on an available emotional dataset PaDWEED, extensive experiments demonstrate that CDRC performs comparably to fully supervised training and achieves state-of-the-art results compared to several self-supervised approaches. Moreover, the remarkable performance on subject-independent tasks highlights its effectiveness in addressing and mitigating subject variability. Meihong Zhang, Shaokai Zhao, Zhiguo Luo, Liang Xie 0012, Tiejun Liu, Dezhong Yao 0001, Ye Yan 0001, Erwei Yin |
ICASSP | 8 |
| 2025 | Hierarchical-Aware Orthogonal Disentanglement Framework for Fine-Grained Skeleton-Based Action Recognition
Haochen Chang, Pengfei Ren 0001, Liang Xie 0012, Erwei Yin |
ICCV | 6 |
| 2025 | M2EIT: Multi-Domain Mixture of Experts for Robust Neural Inertial Tracking
Changhao Chen, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Erwei Yin |
ICCV | 8 |
| 2025 | Calibration-Free Multi-view 3D Hand Pose Estimation for XR Cockpit Interactions
Hanling Zhan, Baojun Chen, Meng Gai, Liang Xie 0012, Erwei Yin |
ICXR | 8 |
| 2025 | Lite-DIO Is Actually What You Need for Efficient Inertial Localization
Zhongchen Shi, Yanqing Hou, Liang Xie 0012, Erwei Yin |
AAMAS | 7 |
| 2025 | Learning Neural Vocoder from Range-Null Space DecompositionabstractDespite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned challenges. To be specific, we bridge the connection between the classical signal range-null decomposition (RND) theory and vocoder task, and the reconstruction of target spectrogram can be decomposed into the superimposition between the range-space and null-space, where the former is enabled by a linear domain shift from the original mel-scale domain to the target linear-scale domain, and the latter is instantiated via a learnable network for further spectral detail generation. Accordingly, we propose a novel dual-path framework, where the spectrum is hierarchically encoded/decoded, and the cross- and narrow-band modules are elaborately devised for efficient sub-band and sequential modeling. Comprehensive experiments are conducted on the LJSpeech and LibriTTS benchmarks. Quantitative and qualitative results show that while enjoying lightweight network parameters, the proposed approach yields state-of-the-art performance among existing advanced methods. Our code and the pretrained model weights are available at https://github.com/Andong-Li-speech/RNDVoC. Andong Li, Zhihang Sun, Rilin Chen, Erwei Yin, Xiaodong Li 0002, Chengshi Zheng |
IJCAI | 5 |
| 2025 | Unveiling Genuine Emotions: Integrating Micro-Expressions and Physiological Signals for Enhanced Emotion RecognitionabstractIn recent years, multimodal emotion recognition has attracted growing interest due to its potential to improve emotion classification accuracy by integrating information from diverse modalities. This study replicates a scenario in which individuals suppress facial expressions to conceal emotions under intense emotional stimuli, while simultaneously recording micro-expressions (MEs), electroencephalograms (EEG), and peripheral physiological data (PERI) from 75 participants. The resulting multimodal dataset consists of 634 ME video clips across seven emotional categories, as well as 2,890 trials of physiological signals (PS). To assess the dataset’s reliability, we establish a cross-modal contrastive learning framework that incorporates diversity contrastive learning, consistency contrastive learning, and sample-level contrastive learning, designed to capture complementary features across different modalities. Experimental results confirm that multimodal fusion significantly enhances emotion classification accuracy. This study not only offers a solution for emotion analysis but also contributes to understanding the relationship between MEs and PS, along with their underlying mechanisms. Shaokai Zhao, Liang Xie 0012, Erwei Yin |
IJCNN | 5 |
| 2025 | A Transformer-Based Multimodal Framework for Hidden Emotion Recognition through Micro-Expression and EEG FusionabstractWith the development of multimedia technology, emotion recognition has gradually matured, but hidden emotion recognition still faces numerous challenges. Given the unique advantages of micro-expressions (MEs) and electroencephalogram (EEG) signals in capturing subtle emotional cues, we recreated scenarios where individuals suppress facial expressions to conceal emotions in response to intense emotional stimuli. Simultaneous recording of MEs and EEG data from 75 participants resulted in a dataset comprising 634 ME video clips and 2,890 EEG trials across seven emotional categories. To assess the reliability of the dataset, we developed an emotion classification model based on a cross-modal attention mechanism. This framework enables dynamic information exchange between the two modalities through a Cross-Transformer architecture. The model achieved an accuracy of 89.71% for three-class classification and 40.22% for seven-class classification, demonstrating a significant improvement over conventional unimodal approaches. Shaokai Zhao, Liang Xie 0012, Erwei Yin |
ICMR | 5 |
| 2025 | VIHand: Enhancing 3D Hand Pose Estimation with Visual-Inertial Benchmark
Pengfei Ren 0001, Liang Xie 0012, Yue Gao 0005, Erwei Yin |
ACM Multimedia | 8 |
| 2025 | DuAGNet: an unrestricted multimodal speech recognition framework using dual adaptive gating fusion
Jinghan Wu, Yakun Zhang 0002, Meishan Zhang, Changyan Zheng, Liang Xie 0012, Xingwei An, Erwei Yin |
Appl. Intell. | 8 |
| 2025 | FI-HGR: A Robust Hand Gesture Recognition System Based on Wearable Data Glove and Multimodal Fusion AlgorithmabstractHand gesture recognition (HGR) plays a crucial role in human-computer interaction systems within the Internet of Things (IoT). Recent HGR methods often rely on vision-based images or videos, which are limited in terms of occluded fingers and high computational cost due to complex neural networks. In contrast, wearable sensors like inertial measurement units (IMUs) and flexible sensors can handle hand self-obscuration. However, there are two unresolved issues. First, using a single modality is hard to balance high precision and low latency. Second, existing multimodal-based approaches lack deep inter-modal coupling to effectively address IMU drift and mechanical coupling of flexible sensors. To address these problems, we propose FI-HGR (HGR based on flexible and inertial data). FI-HGR comprises a sensor-integrated data glove and a novel Cascaded Complementary-Stochastic Fusion Algorithm (CS-Algorithm). CS-Algorithm employs six Mahony filters to estimate the state quaternion of each IMU, along with an Extended Kalman Filter that continuously corrects IMU drift based on the index finger’s bending angle sensed by a flexible sensor. This design allows a single flexible sensor to calibrate multiple IMUs and introduces a hard constraint, resolving the sensor drift problems that previous methods cannot. Based on the CS-Algorithm outputs, precise finger bending angles are estimated in real time. Subjective and objective experimental results show that our approach effectively addresses IMU drift and mechanical coupling in flexible sensors, reduces gesture tracking error to approximately 3.4∘, and significantly improves both recognition accuracy and operational efficiency. Tao Zhen, Buyuan Zhang, Dezhong Yao 0001, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
IEEE Internet Things J. | 9 |
| 2025 | Bridging semantics across modalities: Decoupled representation learning for audio-visual speech recognition
Linzhi Wu, Yakun Zhang 0002, Changyan Zheng, Tiejun Liu, Liang Xie 0012, Chengshi Zheng, Erwei Yin |
Knowl. Based Syst. | 8 |
| 2025 | PanoGen++: Domain-adapted text-guided panoramic environment generation for vision-and-language navigation
Dongliang Zhou, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Neural Networks | 6 |
| 2025 | MsDUNE: A multi-scale masked temporal fusion framework for speaker-independent lipreading via Dirichlet uncertainty estimation
Jinghan Wu, Xingwei An, Yakun Zhang 0002, Changyan Zheng, Liang Xie 0012, Erwei Yin |
Neural Networks | 7 |
| 2025 | Neural Chinese silent speech recognition with facial electromyography
Liang Xie 0012, Yakun Zhang 0002, Meishan Zhang, Changyan Zheng, Ye Yan 0001, Erwei Yin |
Speech Commun. | 8 |
| 2025 | Lipvis: A Novel Transient Viseme Extraction Framework for Lip ReadingabstractThe accurate extraction of visemes—the minimal distinguishable units in lip reading—lacks a systematic solution. Existing methods primarily rely on audio alignment and phoneme mapping, which suffers from inconsistent categorization and temporal information loss. Suitable evaluation of the extraction quality also remain scarce. This work bridges the research gap by establishing Lipvis as the first viseme extraction framework for lip reading, grounded in the novel conceptualization of transient visemes. Lipvis enables audio-independent extraction of transient visemes with preserved temporal dynamics via a multi-stage pipeline. These visemes are subsequently annotated to assist lip reading, validating their incremental utility. Experiments on both Chinese and English datasets demonstrate that our method enhances lip reading performance through the incorporation of fine-grained transient viseme knowledge. It achieves a word error rate (WER) of 1.02 % on GRID and a character error rate (CER) of 20.93 % on CMLR, outperforming baselines and marking the best-known performance under identical experimental conditions. These compelling results demonstrate Lipvis's capability to extract high-quality, task-relevant visemes, establishing a foundation for advancing viseme-based lipreading research. Yakun Zhang 0002, Liang Xie 0012, Erwei Yin |
IEEE Signal Process. Lett. | 5 |
| 2025 | Identifying Stable EEG Patterns in Manipulation Task for Negative Emotion RecognitionabstractNegative emotion recognition during manipulation task plays crucial role in human-machine interaction, where diverse cognitive variables coexist and influence each other. However, traditional emotion experiments often overemphasize emotion induction while overlooking other practical cognitive tasks, which leads participants to suffer from simplistic emotional experiences and ultimately compromises the real-world applicability of the emotional data collected. To incorporate critical cognitive variables into emotion elicitation, we utilize joystick-based real-time emotion annotation to encourage subjects to continuously feel emotional intensity, to advisedly decide when to manipulate the joystick, and to physically operate it. Consequently, at least two essential cognitive variables—decision-making and action—are integrated into emotion perception. Following this, we develop a novel negative emotion dataset called CRED, which includes a variety of physiological data, particularly Electroencephalograph (EEG). To assess the stability of emotional EEG patterns, we employ strict statistical analysis and a dual-branch transformer (DBT) with the gradient-based attribution method on the proposed CRED. Additionally, two well-known public datasets (SEED and SEED-V) are used to verify the DBT. Compared to traditional methods, DBT improves classification accuracy by approximately 5% on CRED and by around 2% on the public datasets. The experimental results indicate that the occipital lobe plays a crucial role in the discrimination of negative emotions; the critical frequency bands vary between the five emotions in the CRED. Specifically, the low-delta rhythm is associated with anger, while fear is influenced by both theta and alpha rhythms; disgust is found to be significant in the theta rhythm; and for neutral emotions, both low-delta and alpha rhythms are identified as crucial. In summary, our findings demonstrate the existence of stable emotional EEG patterns when additional cognitive variables are involved. Shaokai Zhao, Liang Xie 0012, Zhiguo Luo, Dongdong Zhou, Ye Yan 0001, Erwei Yin |
IEEE Trans. Affect. Comput. | 8 |
| 2025 | AVE Speech: A Comprehensive Multimodal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic SignalsabstractThe global aging population faces considerable challenges, particularly in communication, due to the prevalence of hearing and speech impairments. To address these, we introduce the AVE speech, a comprehensive multimodal dataset for speech recognition tasks. The dataset includes a 100-sentence Mandarin corpus with audio signals, lip-region video recordings, and six-channel electromyography data, collected from 100 participants. Each subject read the entire corpus ten times, with each sentence averaging approximately two seconds in duration, resulting in over 55 hours of multimodal speech data per modality. Experiments demonstrate that combining these modalities significantly improves recognition performance, particularly in cross-subject and high-noise environments. To our knowledge, this is the first publicly available sentence-level dataset integrating these three modalities for large-scale Mandarin speech recognition. We expect this dataset to drive advancements in both acoustic and nonacoustic speech recognition research, enhancing cross-modal learning and human–machine interaction. Dongliang Zhou, Yakun Zhang 0002, Jinghan Wu, Liang Xie 0012, Erwei Yin |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2025 | PVEye: A Large Posture-Variant Eye Tracking Dataset for Head-Mounted AR DevicesabstractEye tracking technology, essential for enhancing user experience in virtual reality (VR) and augmented reality (AR) devices, has been widely incorporated into advanced head-mounted devices like the Apple Vision Pro and PICO 4 Pro, becoming a standard feature. However, dedicated eye tracking datasets for such devices are severely lacking, with existing datasets commonly facing issues like camera skew and low resolution, particularly failing to adequately consider the diversity in wearing postures. To address this gap, we have developed the Posture-Variant Eye Tracking Dataset (PVEye), which includes 11,044,800 high-resolution near-eye images from 104 participants, showcasing a rich variety of wearing postures. This dataset aims to advance the development and application of appearance-based eye tracking methods. Utilizing this dataset, our evaluations demonstrate that the appearance-based method, particularly the NVGaze model, provides improved accuracy and robustness compared to the traditional feature-based method. Crucially, our experiments indicate that variations in wearing posture can significantly impact eye tracking performance, with posture-related errors contributing approximately 45% to the overall error variance. Moreover, the study delves into the specific impact of calibration and other critical factors on eye tracking performance, offering insights for further optimization of tracking effectiveness. Xiaowei Bai, Liang Xie 0012, Yingxi Li, Qining Wang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Two-Stream Vision Swin Transformer for Video-based Eye Movement Detection
Xiaowei Bai, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
CogSci | 7 |
| 2024 | Landmark-Guided Cross-Speaker Lip Reading with Mutual Information RegularizationabstractLip reading, the process of interpreting silent speech from visual lip movements, has gained rising attention for its wide range of realistic applications. Deep learning approaches greatly improve current lip reading systems. However, lip reading in cross-speaker scenarios where the speaker identity changes, poses a challenging problem due to inter-speaker variability. A well-trained lip reading system may perform poorly when handling a brand new speaker. To learn a speaker-robust lip reading model, a key insight is to reduce visual variations across speakers, avoiding the model overfitting to specific speakers. In this work, in view of both input visual clues and latent representations based on a hybrid CTC/attention architecture, we propose to exploit the lip landmark-guided fine-grained visual clues instead of frequently-used mouth-cropped images as input features, diminishing speaker-specific appearance characteristics. Furthermore, a max-min mutual information regularization approach is proposed to capture speaker-insensitive latent representations. Experimental evaluations on public lip reading datasets demonstrate the effectiveness of the proposed approach under the intra-speaker and inter-speaker conditions. Linzhi Wu, Yakun Zhang 0002, Changyan Zheng, Tiejun Liu, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
LREC/COLING | 8 |
| 2024 | Complex-Valued Gabor-Attention Residual Fusion Network for Iris RecognitionabstractIris recognition has gained significant attention in identity verification due to the unique, stable texture patterns in iris. Successfully extracting these patterns is essential for quick and precise identification. Although deep learning methods have automated the iris recognition, they predominantly rely on real-valued networks that overlook the complex-valued representation of iris texture. This means they cannot effectively process phase and amplitude information, and fail to integrate domain-specific knowledge of iris, thereby not fully capturing the intricate details of the iris texture. Inspired by classical manual methods that efficiently harness the complex-valued representation of the iris to extract both amplitude and phase information. We integrate Gabor filters with complex-valued neural networks, propose a Complex-Valued Gabor-Attention Residual Fusion Network (GRFN) tailored for iris recognition, aiming to comprehensively capture the iris texture’s multi-scale and multi-orientation phase and amplitude features. The GRFN incorporates adaptive Gabor Complex-Valued Convolution Kernels (GCVK) to introduce a Gabor attention mechanism focused on iris biometric characteristics. Furthermore, we propose a novel residual feature fusion approach that selects and merges local and global features across multiple directions and scales, mitigating model degradation and enhancing the network’s ability to extract iris texture features effectively. Extensive experiments show that the proposed network outperforms the state-of-the-art performance on two benchmark datasets. Zhuoru Li, Xiaowei Bai, Yingxi Li, Zhenyu Fang, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ECAI | 9 |
| 2024 | GloveTyping: A Hand Gesture Recognition System for Text Input Using a Hierarchical Framework with Attention Mechanism
Tao Zhen, Pengfei Ren 0001, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (5) | 8 |
| 2024 | Exploring Matching Rates: From Keypoint Selection to Camera Relocalization
Chengjiang Long, Yifeng Fei, Qianchen Xia, Erwei Yin, Xin Yang 0011 |
ACM Multimedia | 5 |
| 2024 | Trajectory-based Calibration for Optical See-Through Head-Mounted Displays Without Alignment
Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (6) | 7 |
| 2024 | Challenges and solutions for vision-based hand gesture interpretation: A review
Kun Gao 0002, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Comput. Vis. Image Underst. | 8 |
| 2024 | Flexible Strain Sensor-Based Data Glove for Gesture Interaction in the Metaverse: A ReviewabstractThe rise of the metaverse concept has brought about widespread attention in wearable gesture recognition devices. Data gloves based on flexible strain sensors have been favoured by researchers owing to their low cost, light weight, direct and continuous monitoring of finger movements. In this review, we first compare the advantages and disadvantages of four different approaches based on the vision sensors, myoelectric sensors, inertial and magnetic sensors, and flexible strain sensors in designing data gloves, and demonstrate the superiority of the flexible strain sensor-based data glove used for metaverse applications. Next, some latest commercial data gloves are exampled and the function modules of the data gloves are presented based on the flexible strain sensors. Meanwhile, the potential applications of gesture recognition in the metaverse are summarized in diversified fields. Finally, the existing problems and development prospects of the current data gloves based on flexible strain sensors are concluded. We are optimistic that novel flexible strain sensor-based data gloves will make transformational impact to realize accurate, low-latency, and immersive gesture interaction in the metaverse. Xuanqi Wang, Zekai Liang, Qianchen Xia, Liang Xie 0012, Huijiong Yan, Fanqi Sun, Huicheng Feng, Kai Tao, Erwei Yin |
Int. J. Hum. Comput. Interact. | 12 |
| 2024 | MVINS: Tightly Coupled Mocap-Visual-Inertial Fusion for Global and Drift-Free Pose EstimationabstractAugmented reality (AR), a prominent application within the Internet of Things (IoT) domain, demands high-performance pose estimation. Presently, the visual-inertial navigation system (VINS) is acknowledged as an essential method for providing 6-DoF poses. However, VINS builds the local frame at random during the system initialization stage, making it difficult to establish a connection with the global frame. In addition, VINS is prone to drifting. In this paper, we propose an innovative method that tightly couples markerless motion capture (Mocap) with vision and an IMU to achieve global and drift-free pose estimation for AR glasses. To address the issue of pose initialization and establish a connection between the IMU and Mocap, we introduce a coarse-to-fine initialization strategy, enabling data fusion for Mocap, vision, and the IMU under a unified global frame. Furthermore, we formulate the Mocap factor alongside the visual and inertial factors and integrate them into a factor graph framework to constrain the system states. With a spatiotemporal calibration method, the IMU-Mocap extrinsic parameter and time offset are calibrated online to improve the pose estimation accuracy. Experimental evaluations in real-world experiments demonstrate the capability of our method to accurately estimate drift-free poses in the global frame. Compared to the state-of-the-art VINS-Fusion, ORB-SLAM3, and GVIS, we achieve improvements of 81%, 42%, and 33% in translation accuracy and improvements of 58%, 33%, and 72% in rotation accuracy, respectively. Moreover, we also evaluate our system for the EuRoC dataset, further indicating the effectiveness of the proposed work. Liang Xie 0012, Wei Wang 0076, Zhongchen Shi, Wei Chen 0092, Ye Yan 0001, Erwei Yin |
IEEE Internet Things J. | 7 |
| 2024 | Progressively global-local fusion with explicit guidance for accurate and robust 3d hand pose reconstruction
Kun Gao 0002, Pengfei Ren 0001, Tao Zhen, Liang Xie 0012, Zhongkui Li, Ye Yan 0001, Erwei Yin |
Knowl. Based Syst. | 10 |
| 2024 | Trajectory-based alignment for optical see-through HMD calibrationabstractAbstract In order to align the virtual and real content precisely through augmented reality devices, especially in optical see-through head-mounted displays (OST-HMD), it is necessary to calibrate the device before using it. However, most existing methods estimated the parameters via 3D-2D correspondences based on the 2D alignment, which is cumbersome, time-consuming, theoretically complex, and results in insufficient robustness. To alleviate this issue, in this paper, we propose an efficient and simple calibration method based on the principle of directly calculating the projection transformation between virtual space and the real world via 3D-3D alignment. The proposed method merely needs to record the motion trajectory of the cube-marker in the real and virtual world, and then calculate the transformation matrix between the virtual space and the real world by aligning the two trajectories in the observed view. There are two advantages associated with the proposed method. First, the operation is simple. Theoretically, the user only needs to perform four alignment operations for calibration without changing the rotation variation. Second, the trajectory can be easily distributed throughout the entire observation view, resulting in more robust calibration results. To validate the effectiveness of the proposed method, we conducted extensive experiments on our self-built optical see-through head-mounted display (OST-HMD) device. The experimental results show that the proposed method can achieve better calibration results than other calibration methods. Lingling Chen, Shaohua Zhao, Wei Chen 0092, Zhongchen Shi, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
Multim. Tools Appl. | 7 |
| 2024 | Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted DevicesabstractGaze is a crucial element in human-computer interaction and plays an increasingly vital role in promoting the adoption of head-mounted devices (HMDs). Existing gaze tracking methods for HMDs either demand user calibration or face challenges in balancing accuracy and speed, compromising the overall user experience. In this paper, we introduce a novel strategy for real-time, calibration-free gaze tracking using joint head-eye cues on HMDs. Initially, we create a multimodal gaze tracking dataset named HE-Gaze, encompassing synchronized eye images and 6DoF head movement data, addressing a gap in the current data landscape. Statistical analyses unveil the correlation between head movements and gaze positions. Building on these insights, we introduce the hierarchical head-eye coordinated gaze tracking model (HHE-Tracker), which incorporates two lightweight branches to encode input eye images and head sequences efficiently. It combines encoded head velocity and posture features with eye features across various scales to infer gaze position. HHE-Tracker was implemented on a commercial HMD, and its performance was assessed in unconstrained scenarios. The results demonstrate the HHE-Tracker's capability to accurately estimate gaze positions in real-time. In comparison to the state-of-the-art gaze tracking algorithm, HHE-Tracker exhibits commendable accuracy (3.47$^{\circ }$) and a 40-fold speedup (81FPSon a Snapdragon 845 SoC). Yingxi Li, Xiaowei Bai, Liang Xie 0012, Feng Lu 0005, Feitian Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | Vision-Language Navigation With Beam-Constrained Global NormalizationabstractVision-language navigation (VLN) is a challenging task, which guides an agent to navigate in a realistic environment by natural language instructions. Sequence-to-sequence modeling is one of the most prospective architectures for the task, which achieves the agent navigation goal by a sequence of moving actions. The line of work has led to the state-of-the-art performance. Recently, several studies showed that the beam-search decoding during the inference can result in promising performance, as it ranks multiple candidate trajectories by scoring each trajectory as a whole. However, the trajectory-level score might be seriously biased during ranking. The score is a simple averaging of individual unit scores of the target-sequence actions, and these unit scores could be incomparable among different trajectories since they are calculated by a local discriminant classifier. To address this problem, we propose a global normalization strategy to rescale the scores at the trajectory level. Concretely, we present two global score functions to rerank all candidates in the output beam, resulting in more comparable trajectory scores. In this way, the bias problem can be greatly alleviated. We conduct experiments on the benchmark room-to-room (R2R) dataset of VLN to verify our method, and the results show that the proposed global method is effective, providing significant performance than the corresponding baselines. Our final model can achieve competitive performance on the VLN leaderboard. Liang Xie 0012, Meishan Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Grounded Entity-Landmark Adaptive Pre-training for Vision-and-Language NavigationabstractCross-modal alignment is one key challenge for Vision-and-Language Navigation (VLN). Most existing studies concentrate on mapping the global instruction or single sub-instruction to the corresponding trajectory. However, another critical problem of achieving fine-grained alignment at the entity level is seldom considered. To address this problem, we propose a novel Grounded Entity-Landmark Adaptive (GELA) pre-training paradigm for VLN tasks. To achieve the adaptive pre-training paradigm, we first introduce grounded entity-landmark human annotations into the Room-to-Room (R2R) dataset, named GEL-R2R. Additionally, we adopt three grounded entity-landmark adaptive pre-training objectives: 1) entity phrase prediction, 2) landmark bounding box prediction, and 3) entity-landmark semantic alignment, which explicitly supervise the learning of fine-grained cross-modal alignment between entity phrases and environment landmarks. Finally, we validate our model on two downstream benchmarks: VLN with descriptive instructions (R2R) and dialogue instructions (CVDN). The comprehensive experiments show that our GELA model achieves state-of-the-art results on both tasks, demonstrating its effectiveness and generalizability. Liang Xie 0012, Yakun Zhang 0002, Meishan Zhang, Ye Yan 0001, Erwei Yin |
ICCV | 6 |
| 2023 | Auxiliary Fine-grained Alignment Constraints for Vision-and-Language NavigationabstractVision-and-Language Navigation (VLN) requires a visual agent to navigate in photo-realistic environments following instructions. Fine-grained cross-modal alignment is one critical challenge in VLN because the agent needs to focus on a particular sub-part within the complete instruction for the next movement. However, previous work failed to implement explicit supervision for matching the sub-trajectory to the corresponding sub-instruction. In this paper, we propose Auxiliary Fine-grained Alignment Constraints (AFAC) to facilitate decision-making learning during navigation. AFAC consists of two constraints, i.e., Attention Alignment Constraint (AAC) and Representation Alignment Constraint (RAC), which produce additional supervising signals from the perspective of attention and representation respectively. We test our method on the Landmark-RxR benchmark and achieve state-of-the-art results both in seen and unseen environments. Ruqiang Huang, Yakun Zhang 0002, Yingjie Cen, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICME | 7 |
| 2023 | One-Stage Wireframe Parsing in Fish-Eye Images
Ruqiang Huang, Zhongchen Shi, Wei Chen 0092, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
PRCV (11) | 7 |
| 2022 | Improved Word-level Lipreading with Temporal Shrinkage Network and NetVLADabstractIn most word-level lipreading architectures of recent years, temporal feature extraction module tend to employ Multi-scale Temporal Convolution Network (MS-TCN). In our experiments, we have noticed it is hard for MS-TCN to deal with noise information that may contain in image sequences. In order to solve the problems, we propose a lipreading architecture based on temporal shrinkage network and NetVLAD. We first propose Temporal Shrinkage Unit according to Residual Shrinkage Network and then replace temporal convolution unit with it. The improved network which named Multi-scale Temporal Shrinkage Network (MS-TSN) could focus more on relevant information. Following with MS-TSN that deals with noise frames, NetVLAD is proposed to integrate local information into global feature. Compared with Global Average Pooling, NetVLAD could extract key features by clustering. Our experiments on Lipreading in the Wild (LRW) show that the architecture we propose achieves an accuracy of 89.41%, attaining new state-of-the-art in word-level lipreading. In addition, we build a new Mandarin Chinese lipreading dataset named MCLR-100 and verify our proposed architecture on it. Tao Luo 0010, Yakun Zhang 0002, Mingwu Song, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICMI | 7 |
| 2022 | Real-time Gaze Tracking with Head-eye Coordination for Head-mounted DisplaysabstractHigh-accuracy, low-latency gaze tracking is becoming one of the indispensable features in augmented reality (AR) head-mounted devices (HMDs). Researchers have proposed different approaches to predict gaze positions from eye images. However, since only the eye modality is focused, these appearance-based algorithms are still struggle to trade off the accuracy and running speed in HMDs. In this paper, we propose a lightweight multi-modal network (HE-Tracker) to regress gaze positions. By fusing head-movement features with eye features, HE-Tracker achieves comparable accuracy (3.655° in all subjects) and $27 \times$ speedup (48 fps in the specialized AR HMD) compared to the state-of-the-art gaze tracking algorithm. We further demonstrate that when applying our head-eye coordination strategy to other baseline models, all these models achieve at least 6.36% performance improvement without a pronounced effect on running speed. Moreover, we construct HE-Gaze, the first multi-modal dataset with eye images and head-movement data for near-eye gaze tracking. This dataset is currently made of 757,360 frames and 15 persons, providing an opportunity to foster research in multi-modal gaze tracking approaches. Our dataset is available at DOWNLOAD LINK1. Lingling Chen, Yingxi Li, Xiaowei Bai, Yongqiang Hu, Mingwu Song, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ISMAR | 9 |
| 2022 | Recent advances in wireless epicortical and intracortical neuronal recording systems
Zekai Liang, Xichen Yuan, Honglai Xu, Erwei Yin, Zhejun Guo, Longchun Wang, Huicheng Feng, Honglong Chang |
Sci. China Inf. Sci. | 6 |
| 2022 | Auto calibration of multi-camera system for human pose estimationabstractAbstract The multi‐camera calibration is an essential step for many spatially aware applications, such as robotic navigation, augmented reality, and 3D human pose estimation. Traditional calibration methods use off‐the‐shelf checkerboards or triangles as the known world coordinate system and their corresponding corners are set as control points, which heavily depends on specific calibration patterns and is not suitable for calibration pattern‐denied environments. In this paper, an automatic calibration method is proposed to calibrate the multi‐camera system without the aid of a known calibration pattern. The key idea of the proposed method is that the authors consider the human body, which is always available, as the counterpart of the calibration pattern. The authors’ approach starts with binocular camera calibration, in which the extrinsic and intrinsic parameters are calculated in order and followed by a joint optimisation. With the results of each pair of binocular camera calibration, the multi‐camera system calibration is carried out in three steps: (i) parameters initialisation, (ii) extrinsic parameters optimisation, and (iii) jointly optimising intrinsic and extrinsic parameters. Since the authors’ approach does not require additional calibration patterns except for one visible person, it is flexible and easy to be implemented. Real experiments are conducted in different scenes, camera angles, and camera settings. Human pose estimation with the multi‐camera system is additionally performed for exhaustive experiments. The experimental results demonstrate that the authors’ method shows superior performance than the traditional method with the aid of a specific calibration pattern. Lingling Chen, Liang Xie 0012, Jian Yin 0025, Shuwei Gan, Ye Yan 0001, Erwei Yin |
IET Comput. Vis. | 7 |
| 2021 | A Fusion Framework to Enhance sEMG-Based Gesture Recognition Using TD and FD Features
Yao Luo, Tao Luo 0010, Qianchen Xia, Huijiong Yan, Liang Xie 0012, Ye Yan 0001, Erwei Yin |
ICONIP (6) | 7 |
| 2019 | An Tactile ERP-Based Brain-Computer Interface for CommunicationabstractA classical visual event relative potential (ERP) brain–computer interface (BCI) system relies on visual stimuli to choose commands. Users obtain most information about their surroundings visually as well. This large amount of information can aggravate visual burden and fatigue. In our study, we proposed a novel approach to evoke ERP with a tactile stimulus. To achieve this approach, we first designed a wireless stimulus module with vibrators to provide a tactile stimulus for the system. The vibrators were located on the subject’s arm to imitate the joint motion of a robotic arm. Then, the ERP feature and the parameters of classifiers were obtained through offline experimental data analysis. Based on the analysis, the suitable electrode channels, stimulus onset asynchrony (SOA), and filter upper limit were different for different subjects. According to those outcomes, a unique classifier was designed for each subject. Finally, 10 healthy BCI-naive subjects participated in online experiments to evaluate the performance of our tactile BCI system; they achieved an accuracy range from 78.67% to 100% with an average of 89.1% and an instantaneous transmission rate (ITR) range from 7.77 to 28.70 bits/min with an average of 14.77 bits/min. The accuracy of different subjects and SOAs remained relatively stable, the ITR fluctuated mainly due to the different SOAs, and we achieved balance between ITR and accuracy. Yadong Liu 0001, Jingjun Wang, Erwei Yin, Yang Yu 0014, Zongtan Zhou, Dewen Hu |
Int. J. Hum. Comput. Interact. | 3 |
| 2019 | Towards a Hybrid BCI Gaming Paradigm Based on Motor Imagery and SSVEPabstractBrain-computer interfaces (BCIs) not only can allow individuals to voluntarily control external devices, helping to restore lost motor functions of the disabled, but can also be used by healthy users for entertainment and gaming applications. In this study, we proposed a hybrid BCI paradigm to explore a feasible and natural way to play games by using electroencephalogram (EEG) signals in a practical environment. In this paradigm, we combined motor imagery (MI) and steady-state visually evoked potentials (SSVEPs) to generate multiple commands. A classic game, Tetris, was chosen as the control object. The novelty of this study includes the effective usage of a “dwell time” approach and fusion rules to design BCI games. To demonstrate the feasibility of the proposed hybrid paradigm, ten subjects were chosen to participate in online control experiments. The experimental results showed that all subjects successfully completed the predefined tasks with high accuracy. This proposed hybrid BCI paradigm could potentially provide those who suffer disability or paralysis with additional entertainment options, such as brain-actuated games, that could improve their happiness and quality of life.Abbreviations: BCI: brain-computer interface; EEG: electroencephalogram; MI: motor imagery; SSVEP: steady-state visually evoked potential; ERP: event-related potential; SMR: sensorimotor rhythm; VEP: visual evoked potential; TCP/IP: transmission control protocol/internet protocol; GUI: graphical user interface; ERD/ERS: event-related desynchronization/synchronization; CIC: control intention classifier; LRC: left/right classifier; CSP: common spatial pattern; LDA: linear discriminant analysis; ROC: receiver operating characteristic; TPR: true positive rate; FPR: false positive rate; CCA: canonical correlation analysis. Zhihua Wang 0002, Yang Yu 0014, Ming Xu 0022, Yadong Liu 0001, Erwei Yin, Zongtan Zhou |
Int. J. Hum. Comput. Interact. | 5 |
| 2019 | Hierarchical feature fusion framework for frequency recognition in SSVEP-based BCIs
Yangsong Zhang 0001, Erwei Yin, Fali Li, Yu Zhang 0009, Daqing Guo, Dezhong Yao 0001, Peng Xu 0001 |
Neural Networks | 2 |
| 2019 | Sparse Group Representation Model for Motor Imagery EEG ClassificationabstractA potential limitation of a motor imagery (MI) based brain-computer interface (BCI) is that it usually requires a relatively long time to record sufficient electroencephalogram (EEG) data for robust classifier training. The calibration burden during data acquisition phase will most probably cause a subject to be reluctant to use a BCI system. To alleviate this issue, we propose a novel sparse group representation model (SGRM) for improving the efficiency of MI-based BCI by exploiting the intersubject information. Specifically, preceded by feature extraction using common spatial pattern, a composite dictionary matrix is constructed with training samples from both the target subject and other subjects. By explicitly exploiting within-group sparse and group-wise sparse constraints, the most compact representation of a test sample of the target subject is then estimated as a linear combination of columns in the dictionary matrix. Classification is implemented by calculating the class-specific representation residual based on the significant training samples corresponding to the nonzero representation coefficients. Accordingly, the proposed SGRM method effectively reduces the required training samples from the target subject due to auxiliary data available from other subjects. With two public EEG data sets, extensive experimental comparisons are carried out between SGRM and other state-of-the-art approaches. Superior classification performance of our method using 40 trials of the target subject for model calibration (Averaged accuracy = 78.2%, Kappa = 0.57 and Averaged accuracy = 77.7%, Kappa = 0.55 for the two data sets, respectively) indicates its promising potential for improving the practicality of MI-based BCI. Yong Jiao, Yu Zhang 0009, Xun Chen 0001, Erwei Yin, Jing Jin 0001, Xingyu Wang 0004, Andrzej Cichocki |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Towards correlation-based time window selection method for motor imagery BCIs
Jiankui Feng, Erwei Yin, Jing Jin 0001, Rami Saab, Ian Daly, Xingyu Wang 0004, Dewen Hu, Andrzej Cichocki |
Neural Networks | 2 |
| 2017 | Detect visual field using eye tracking and steady-state visual evoked potentialabstractThis paper makes the subjects' sight locked in a certain area using an eye tracker, getting Steady-state visual evoked potential (SSVEP) from flickering stimuli with a fixed frequency but at random positions, in order to observe the impact of stimulus at different positions and their distances on the electroencephalogram (EEG). The result suggests that if human have to select the positions of stimuli of SSVEP-BCI, it is an agreeable strategy to separate them at least 4 ° for avoiding the possible mistakes. We hope that it could help in setting distances between stimuli or updating pattern selection algorithms in the future BCI system and other paradigms. Yadong Liu 0001, Zongtan Zhou, Dewen Hu, Erwei Yin |
SMC | 5 |
| 2017 | Toward a Hybrid BCI: Self-Paced Operation of a P300-based Speller by Merging a Motor Imagery-Based "Brain Switch" into a P300 Spelling ApproachabstractThis study presents the self-paced operation of a brain–computer interface (BCI) speller, which can be voluntarily turned on/off by merging a motor imagery (MI)-based brain switch into a P300-based BCI speller. From an off state (idle state), the users can generate a “control signal” by consciously changing the cognitive state differential from the idle state to turn on a P300-based spelling system when he or she wants to spell words. With the system turned on, the user can spell words, and then, the spelling system can be voluntarily turned off and switched to the initial state using a command. In this paradigm, the participants tried to perform the two different cognitive tasks sequentially, rather than simultaneously, and multiple EEG components were processed sequentially. The practicability and effectiveness of the proposed approach were validated by eleven participants, and all of them achieved a satisfactory performance. For the P300 speller, they achieved an average PITR of 42.61 bits/min. The preliminary results indicated that the proposed hybrid BCI system with different mental strategies operating sequentially is feasible and has potential applications for practical self-paced control. Yang Yu 0014, Zongtan Zhou, Jun Jiang 0001, Erwei Yin, Kunjia Liu, Jingjun Wang, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 4 |
| 2016 | A P300-Based Brain-Computer Interface for Chinese Character InputabstractThe majority of previously developed assistive communication brain–computer interface systems have primarily focused on languages that are written in alphabetic scripts. However, languages that are written in logographic scripts, such as those in Chinese hanzi (or sinograms), pose a challenge for the implementation of visual spelling systems because it is impossible to simultaneously display thousands of items in a stimulus matrix of a reasonable size. In this study, a P300 visual spelling system that uses a novel method to input Chinese sinograms developed with a Hanyu Pinyin-based method is presented. This method transcribes a Chinese Pinyin into initial consonant and vowel components according to its Mandarin pronunciation. In this paradigm, each sinogram is input by selecting the initial consonant and then the vowel components and subsequently selecting the sinogram itself. Ten healthy subjects participated in the study and achieved an average offline accuracy of 92.6% with a mean information transfer rate of 39.2 bits/min and an average online input speed of one sinogram per 43.9 s. The preliminary results presented here indicated that the online input of Chinese text using a Pinyin-based visual speller is feasible. Yang Yu 0014, Zongtan Zhou, Erwei Yin, Jun Jiang 0001, Yadong Liu 0001, Dewen Hu |
Int. J. Hum. Comput. Interact. | 3 |
| 2016 | An Auditory-Tactile Visual Saccade-Independent P300 Brain-Computer InterfaceabstractMost P300 event-related potential (ERP)-based brain-computer interface (BCI) studies focus on gaze shift-dependent BCIs, which cannot be used by people who have lost voluntary eye movement. However, the performance of visual saccade-independent P300 BCIs is generally poor. To improve saccade-independent BCI performance, we propose a bimodal P300 BCI approach that simultaneously employs auditory and tactile stimuli. The proposed P300 BCI is a vision-independent system because no visual interaction is required of the user. Specifically, we designed a direction-congruent bimodal paradigm by randomly and simultaneously presenting auditory and tactile stimuli from the same direction. Furthermore, the channels and number of trials were tailored to each user to improve online performance. With 12 participants, the average online information transfer rate (ITR) of the bimodal approach improved by 45.43% and 51.05% over that attained, respectively, with the auditory and tactile approaches individually. Importantly, the average online ITR of the bimodal approach, including the break time between selections, reached 10.77 bits/min. These findings suggest that the proposed bimodal system holds promise as a practical visual saccade-independent P300 BCI. Erwei Yin, Timothy J. Zeyl, Rami Saab, Dewen Hu, Zongtan Zhou, Tom Chau |
Int. J. Neural Syst. | 1 |