EDBT 2026 Demo / reviewers in the wild / expert
Zhu Liang Yu
dblp:29/3856 · also Zhuliang Yu
· DBLP profile ↗
79ranked-venue papers
13as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Band-Specific Graph Learning for EEG-Based Emotion RecognitionabstractElectroencephalogram (EEG)-based emotion recognition has attracted increasing attention due to the high temporal resolution of EEG signals. However, accurate EEG-based emotion recognition remains challenging because EEG signals inherently exhibit low signal-to-noise ratios, non-stationarity and considerable inter-individual variability. In addition, brain functional connectivity presents distinct band-dependent characteristics, whereas many existing graph-based methods employ shared or static graphs across spectral bands, limiting their ability to characterize band-specific inter-channel interactions. To address these issues, we propose BSGNet, a band-specific graph learning framework for EEG-based emotion recognition. Specifically, instead of assuming a unified connectivity structure, BSGNet learns a set of band-specific channel adjacency matrices to capture functional interactions associated with different spectral bands. The learned graphs are further integrated into a STGCN backbone to jointly capture spatial dependencies among EEG channels and temporal dynamics of emotional responses. Experimental results on two public EEG emotion recognition datasets show that BSGNet achieves improved average performance over representative deep learning and graph-based baselines, supporting the effectiveness of band-specific connectivity modeling for EEG representation learning. Changchun Li, Chao Mao, Tianyou Yu, Ke Liu 0008, Jun Zhang 0026, Zhenghui Gu, Zhu Liang Yu |
IEEE Signal Process. Lett. | 7 |
| 2025 | Self-Correcting Robot Manipulation via Gaussian-Splatted ForesightabstractLanguage-conditioned robotic manipulation in unstructured environments presents significant challenges for intelligent robotic systems. However, due to partial observation or imprecise action prediction, failure may be unavoidable for learned policies. Moreover, operational failures can lead to the robotic arm entering an untrained state, potentially causing destructive results. Consequently, the ability to detect and self-correct failures is crucial for the development of practical robotic systems. To address this challenge, we propose a foresight-driven failure detection and self-correction module for robot manipulation. By leveraging 3D Gaussian Splatting, we represent the current scene with multiple Gaussians. Subsequently, we train a prediction network to forecast the Gaussian representation of future scenes conditioned on planned actions. Failure is detected when the predicted future significantly deviates from the real observation after action execution. In such cases, the end-effector rolls back to the previous action to avoid an untrained state. Integrating this approach with the PerACT framework, we develop a self-correcting robot manipulation policy. Evaluations on ten RLBench tasks with 166 variations demonstrate the superior performance of the proposed method, which outperforms state-of-the-art methods by 12.0% success rate on average. Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu |
AAAI | 6 |
| 2025 | A Multi-Stage Framework for Chess Puzzle Difficulty PredictionabstractAccurately estimating the difficulty of the chess puzzle is important for adaptive training systems, personalized recommendations, and large-scale content curation.Unlike engine evaluations optimized for perfect play, this task involves modeling human-perceived solving difficulty, typically expressed by Glicko-2 ratings.We present a multi-stage framework developed for the FedCSIS 2025 Challenge.The method trains four rating-banded neural regression models in different Elo ranges to capture localized difficulty patterns and reduce bias from unbalanced data.Their predictions are combined with statistical attributes, including success probabilities, failure distributions, and solution length, through a feature-based regression stage to improve cross-range generalization.A final calibration step adjusts the output to statistically plausible rating levels, mitigating systematic prediction biases without adding computational complexity.An additional mask selection procedure was explored as part of the competition extension to identify 10% of the puzzles that are most likely to benefit from the refined evaluation.The proposed solution ranked 5 th on the public leaderboard and 6 th in the final standings.These results demonstrate that a lightweight and interpretable regression pipeline can achieve competitive precision in modeling human-perceived chess puzzle difficulty. Ling Cen, Jiahao Cen, Malin Song 0001, Zhu Liang Yu |
FedCSIS | 4 |
| 2025 | Rethinking 3D Robotic Perception: Elastic Voxel Representation with Splatting DistillationabstractLanguage-guided robotic manipulation is advancing rapidly with Vision-Language-Action (VLA) models, yet faces fundamental challenges in 3D perception. This paper addresses two critical challenges: the scale elasticity requirement for simultaneously processing coarse environmental context and fine manipulation details, and the scarcity of action-annotated training data. We present Splat-Actor, a novel robotic manipulation framework that introduces two key innovations. First, we develop an elastic voxel encoder that combines multi-scale processing with selective tokenization, enabling efficient 3D spatial reasoning while adaptively focusing on informative regions. Second, we propose a depth-constrained feature distillation framework that leverages Gaussian Splatting to bridge 2D and 3D representations, transferring rich semantic features from pre-trained vision models to enhance 3D understanding. Extensive experiments across 10 manipulation tasks with 166 variations demonstrate that Splat-Actor achieves a 6.8% improvement over state-of-the-art methods while maintaining the computational efficiency. Shaohui Pan, Yong Xu 0007, Ruotao Xu, Zihan Zhou 0007, Si Wu 0002, Zhu Liang Yu, Patrick Le Callet |
ICME | 6 |
| 2025 | Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic GuidanceabstractOpen-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues. Code and appendix are available at github.com/wang678/LLM-UPC. Yufeng Wang 0004, Jinwu Hu, Ziteng Huang, Kunyang Lin, Zitian Zhang, Peihao Chen, Yu Hu 0004, Qianyue Wang, Zhu Liang Yu, Bin Sun 0001, Xiaofen Xing, Mingkui Tan |
IJCAI | 9 |
| 2025 | Cross-Activity sEMG-Driven Joint Angle Estimation via Hybrid Attention Fusion: Bridging Traditional Features and Deep Spatial RepresentationsabstractThe growing prevalence of stroke necessitates advanced lower-limb exoskeleton control. This paper proposes HybridFusionAtt, a novel model for continuous joint angle estimation using surface electromyography (sEMG). Unlike conventional approaches, our framework uniquely integrates traditional time-domain features with CNN-extracted high-dimensional spatial features through an attention mechanism, where traditional features dynamically guide feature fusion as attention queries. The model was validated using data collected from eight participants performing four activities of daily living (walking, stair climbing, stair descending, and obstacle crossing). The proposed model achieves average R2values for knee and hip joint angle prediction of 0.8682 (walking), 0.8482 (obstacle crossing), 0.9294 (stair climbing), and 0.8676 (stair descending). Experimental results show that the proposed model significantly outperforms traditional LSTM and CNN-LSTM models in terms of accuracy and robustness, particularly in handling non-periodic actions such as obstacle crossing. The model achieves high performance by effectively fusing features and adaptively focusing on key features, enabling it to maintain robustness even under noisy conditions and significant individual differences. This demonstrates the model’s broad application potential, especially in rehabilitation and prosthetic control systems. Xiaoyan Deng, Yinke Wen, Jiatong Wu, Zhu Liang Yu |
IROS | 6 |
| 2025 | Open-World Drone Active Tracking with Goal-Centered RewardsabstractDrone Visual Active Tracking aims to autonomously follow a target object by controlling the motion system based on visual observations, providing a more practical solution for effective tracking in dynamic environments. However, accurate Drone Visual Active Tracking using reinforcement learning remains challenging due to the absence of a unified benchmark and the complexity of open-world environments with frequent interference. To address these issues, we pioneer a systematic solution. First, we propose DAT, the first open-world drone active air-to-ground tracking benchmark. It encompasses 24 city-scale scenes, featuring targets with human-like behaviors and high-fidelity dynamics simulation. DAT also provides a digital twin tool for unlimited scene generation. Additionally, we propose a novel reinforcement learning method called GC-VAT, which aims to improve the performance of drone tracking targets in complex scenarios. Specifically, we design a Goal-Centered Reward to provide precise feedback across viewpoints to the agent, enabling it to expand perception and movement range through unrestricted perspectives. Inspired by curriculum learning, we introduce a Curriculum-Based Training strategy that progressively enhances the tracking performance in complex environments. Besides, experiments on simulator and real-world images demonstrate the superior performance of GC-VAT, achieving a Tracking Success Rate of approximately 72% on the simulator. The benchmark and code are available at https://github.com/SHWplus/DAT_Benchmark. Haowei Sun, Jinwu Hu, Zhirui Zhang, Haoyuan Tian, Xinze Xie, Yufeng Wang 0004, Xiaohua Xie, Zhu Liang Yu, Mingkui Tan |
NeurIPS | 9 |
| 2025 | Emotion agent: Unsupervised deep reinforcement learning with distribution-prototype reward for continuous emotional EEG analysis
Li Zhang 0041, Qile Liu, Zhu Liang Yu |
Neurocomputing | 5 |
| 2025 | Enhancing EEG-Based Cross-Subject Emotion Recognition via Adaptive Source Joint Domain AdaptationabstractEEG emotion recognition is crucial in both human-machine interaction and healthcare. However, recognizing emotions across different subjects remains challenging due to individual variability. While existing multi-source domain adaptation methods have been utilized for cross-subject EEG emotion decoding, they often struggle with irrelevant or weakly relevant source domains, leading to negative transfer. Additionally, variations within subdomains are often neglected in these studies. We propose a joint domain adaptation method, Adaptive Source Joint Domain Adaptation (ASJDA) to address these issues. ASJDA utilizes an unsupervised adaptive source selection strategy to select a subset of source domains by evaluating the Jensen-Shannon divergence between the source and target domains, choosing those most relevant to the target. Subsequently, it implements joint domain adaptation with these chosen sources at both the domain and category subdomain levels. Our proposed method outperforms existing state-of-the-art methods, achieving cross-subject accuracies of 96.81% in SEED, 89.69% in SEED-IV, and 69.31% in DEAP. This work significantly advances the state of the art in EEG emotion recognition by effectively addressing the challenges of cross-subject variability. Ke Liu 0008, Wenrui Zhu, Zhu Liang Yu, Hong Yu 0007, Bin Xiao 0002, Wei Wu 0022 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | EEG-Based Cross-Subject Emotion Recognition Using Sparse Bayesian Learning With Enhanced Covariance AlignmentabstractEEG (Electroencephalography)-based emotion recognition has emerged as a crucial area of research due to its potential applications in mental health, brain-computer interfaces (BCIs), and affective computing. However, the inherent variability in EEG signals across individuals, coupled with limited dataset sizes, significantly hinders the development of robust and generalizable emotion recognition models. To overcome these challenges, we propose the Sparse Bayesian Learning with Enhanced Covariance Alignment (SBLECA) algorithm. SBLECA formulates cross-subject emotion recognition as an end-to-end decoding problem, integrating spatiotemporal filtering and classification within a sparse Bayesian learning (SBL) framework. Crucially, SBLECA incorporates a novel covariance alignment technique to mitigate inter-subject variability in EEG patterns. Rigorous evaluations on two publicly available emotion datasets demonstrate that SBLECA consistently outperforms state-of-the-art methods. Furthermore, SBLECA offers valuable insights into the neural correlates of emotion through interpretable visualizations of learned spatial and temporal filters. SBLECA holds promise as a valuable EEG decoding tool to advance the development and translation of neurotechnologies and biomarkers for brain disorders. Feifei Qi, Weichen Huang, Yuanqing Li 0001, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | EmotionMIL: An End-to-End Multiple Instance Learning Framework for Emotion Recognition From EEG SignalsabstractEmotion recognition from EEG signals offers significant advantages in affective computing, as EEG more accurately reflects internal emotional states than other modalities, such as facial expressions or peripheral physiological signals. Modeling and capturing subtle affective changes over time is crucial for real-world applications to achieve better human-computer interaction. However, training such models usually requires segment-level emotion labels, which are costly and may not be feasible. Assigning the overall label to all EEG segments within a trial can lead to inaccurate model training and degraded performance, as emotions evolve continuously. This highlights the need for models capable of learning from trial-wise emotion labels while capturing temporal dynamics of emotional responses within each segment because trial-wise post-stimulus labels are more accessible. To this end, we propose EmotionMIL, an end-to-end EEG-based emotion recognition framework that leverages recent advances in deep multiple instance learning (MIL). This framework enables robust emotion recognition from weakly labeled EEG signals and identifies the most prominent emotional responses. EmotionMIL captures the temporal dynamics of emotions using a retentive self-attention mechanism, which adaptively assigns weights to EEG segments based on their relevance in predicting the overall emotion label. A pseudo-bag augmentation strategy is also introduced to enhance the model's generalization ability by generating additional pseudo-bags from the original ones. Evaluated on three benchmark datasets—DEAP, DREAMER, and SEED—EmotionMIL outperforms state-of-the-art non-MIL and MIL models in both subject-dependent and subject-independent tasks, achieving superior accuracy and F1-score. Ablation study further validates the model design, while visualization results demonstrate that EmotionMIL effectively identifies both spatial EEG patterns and temporal emotional dynamics. These findings underscore EmotionMIL's potential for robust, interpretable emotion recognition, paving the way for real-world applications in emotion-aware systems. The code is available athttps://github.com/yuty2009/emotionmil. Feifei Qi, Lingli Wang, Yanbin He, Jingang Yu, Wei Wu 0022, Zhu Liang Yu, Yuanqing Li 0001, Zhenghui Gu, Tianyou Yu |
IEEE Trans. Affect. Comput. | 7 |
| 2025 | ADMM-ESINet: A Deep Unrolling Network for EEG Extended Source ImagingabstractElectroencephalography (EEG) source imaging (ESI) methods aim to reconstruct cortical sources from scalp EEG signals, a crucial task for understanding the normal brain as well as brain disorders. Traditional model-driven ESI methods face challenges in real-time reconstruction, while deep neural network (DNN)-based ESI methods often struggle with generalization to new data. To address these issues, we propose ADMM-ESINet, a novel deep unfolding neural network for robust and efficient reconstruction of EEG extended sources. ADMM-ESINet leverages a structured sparsity constraint within a regularization framework and employs the Alternating Direction Method of Multipliers (ADMM) to achieve iterative solutions. By unrolling the ADMM algorithm into a cascaded network architecture, ADMM-ESINet effectively integrates prior knowledge, enabling end-to-end, real-time ESI. Crucially, both the regularization parameters and the spatial transform operator are learned directly from the training data. Numerical results demonstrate that ADMM-ESINet surpasses traditional DNN-based methods in generalization ability and accurately reconstructs the location, extent, and temporal dynamics of extended sources, establishing ADMM-ESINet as a promising method for real-time ESI. Ke Liu 0008, Jun Zhang 0026, Zhenghui Gu, Zhu Liang Yu, Yu Zhang 0009, Bin Xiao 0002, Wei Wu 0022 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | DMSACNN: Deep Multiscale Attentional Convolutional Neural Network for EEG-Based Motor DecodingabstractOBJECTIVE: Accurate decoding of electroencephalogram (EEG) signals has become more significant for the brain-computer interface (BCI). Specifically, motor imagery and motor execution (MI/ME) tasks enable the control of external devices by decoding EEG signals during imagined or real movements. However, accurately decoding MI/ME signals remains a challenge due to the limited utilization of temporal information and ineffective feature selection methods. METHODS: This paper introduces DMSACNN, an end-to-end deep multiscale attention convolutional neural network for MI/ME-EEG decoding. DMSACNN incorporates a deep multiscale temporal feature extraction module to capture temporal features at various levels. These features are then processed by a spatial convolutional module to extract spatial features. Finally, a local and global feature fusion attention module is utilized to combine local and global information and extract the most discriminative spatiotemporal features. MAIN RESULTS: DMSACNN achieves impressive accuracies of 78.20%, 96.34% and 70.90% for hold-out analysis on the BCI-IV-2a, High Gamma and OpenBMI datasets, respectively, outperforming most of the state-of-the-art methods. CONCLUSION AND SIGNIFICANCE: These results highlight the potential of DMSACNN in robust BCI applications. Our proposed method provides a valuable solution to improve the accuracy of the MI/ME-EEG decoding, which can pave the way for more efficient and reliable BCI systems. Ke Liu 0008, Zhu Liang Yu, Bin Xiao 0002, Guoyin Wang 0001, Wei Wu 0022 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | M3D: Manifold-Based Domain Adaptation With Dynamic Distribution for Non-Deep Transfer Learning in Cross-Subject and Cross-Session EEG-Based Emotion RecognitionabstractEmotion decoding using Electroencephalography (EEG)-based affective brain-computer interfaces (aBCIs) is crucial for affective computing but is hindered by EEG's non-stationarity, individual variability, and the high cost of large-scale labeled data. Deep learning-based approaches, while effective, require substantial computational resources and large datasets, limiting their practicality. To address these challenges, we propose Manifold-based Domain Adaptation with Dynamic Distribution (M3D), a lightweight non-deep transfer learning framework. M3D includes four main modules: manifold feature transformation, dynamic distribution alignment, classifier learning, and ensemble learning. The data undergoes a transformation onto an optimal Grassmann manifold space, enabling dynamic alignment of the source and target domains. This process prioritizes both marginal and conditional distributions according to their significance, ensuring enhanced adaptation efficiency across various types of data. In the classifier learning, the principle of structural risk minimization is integrated to develop robust classification models. This is complemented by dynamic distribution alignment, which refines the classifier iteratively. Additionally, the ensemble learning module aggregates the classifiers obtained at different stages of the optimization process, which leverages the diversity of the classifiers to enhance the overall prediction accuracy. The proposed M3D framework is evaluated on three benchmark EEG emotion recognition datasets using two validation protocols (cross-subject single-session and cross-subject cross-session), as well as on a clinical EEG dataset of Major Depressive Disorder (MDD). Experimental results demonstrate that M3D outperforms traditional non-deep learning methods, achieving an average improvement of 6.67%, while achieving deep learning-comparable performance with significantly lower data and computational requirements. These findings highlight the potential of M3D to enhance the practicality and applicability of aBCIs in real-world scenarios. Yingwei Qiu, Li Zhang 0041, Zhu Liang Yu |
IEEE J. Biomed. Health Informatics | 6 |
| 2024 | Spatial Adaptive Filter Network With Scale-Sharing Convolution for Image DemoiréingabstractRemoving moiré patterns is a challenging task as it is a spatially varying degradation that varies in shape, color and scale. Existing image restoration models often rely on static convolutional neural networks (CNNs)-based architectures, and hence potentially suboptimal for addressing the diverse manifestations of moiré patterns across different images and spatial positions. To this end, we propose a spatially adaptive neural network for image demoiréing. This network introduces a dual-branch filter prediction module engineered to predict pixel-wise adaptive filters that can process moiré patterns of varying orientations and color-shift issues. To further tackle the challenge presented by scale variability, a scale-sharing convolution module is proposed, utilizing pixel-wise adaptive filters with multiple dilations to handle moiré patterns of different sizes but similar shapes effectively. Upon extensive evaluations of three benchmark datasets, our model consistently outperforms existing methods, yielding a PSNR improvement of over 0.37dB across all evaluated datasets and providing additional benefits in terms of model size. Yong Xu 0007, Zhiyu Wei, Ruotao Xu, Zihan Zhou 0007, Zhu Liang Yu |
IEEE Signal Process. Lett. | 5 |
| 2024 | MSVTNet: Multi-Scale Vision Transformer Neural Network for EEG-Based Motor Imagery DecodingabstractOBJECT: Transformer-based neural networks have been applied to the electroencephalography (EEG) decoding for motor imagery (MI). However, most networks focus on applying the self-attention mechanism to extract global temporal information, while the cross-frequency coupling features between different frequencies have been neglected. Additionally, effectively integrating different neural networks poses challenges for the advanced design of decoding algorithms. METHODS: This study proposes a novel end-to-end Multi-Scale Vision Transformer Neural Network (MSVTNet) for MI-EEG classification. MSVTNet first extracts local spatio-temporal features at different filtered scales through convolutional neural networks (CNNs). Then, these features are concatenated along the feature dimension to form local multi-scale spatio-temporal feature tokens. Finally, Transformers are utilized to capture cross-scale interaction information and global temporal correlations, providing more distinguishable feature embeddings for classification. Moreover, auxiliary branch loss is leveraged for intermediate supervision to ensure the effective integration of CNNs and Transformers. RESULTS: The performance of MSVTNet was assessed through subject-dependent (session-dependent and session-independent) and subject-independent experiments on three MI datasets, i.e., the BCI competition IV 2a, 2b and OpenBMI datasets. The experimental results demonstrate that MSVTNet achieves state-of-the-art performance in all analyses. CONCLUSION: MSVTNet shows superiority and robustness in enhancing MI decoding performance. Ke Liu 0008, Zhu Liang Yu, Weibo Yi, Hong Yu 0007, Guoyin Wang 0001, Wei Wu 0022 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Electromagnetic Source Imaging via a Data-Synthesis-Based Convolutional Encoder-Decoder NetworkabstractElectromagnetic source imaging (ESI) requires solving a highly ill-posed inverse problem. To seek a unique solution, traditional ESI methods impose various forms of priors that may not accurately reflect the actual source properties, which may hinder their broad applications. To overcome this limitation, in this article, a novel data-synthesized spatiotemporally convolutional encoder-decoder network (DST-CedNet) method is proposed for ESI. The DST-CedNet recasts ESI as a machine learning problem, where discriminative learning and latent-space representations are integrated in a CedNet to learn a robust mapping from the measured electroencephalography/magnetoencephalography (E/MEG) signals to the brain activity. In particular, by incorporating prior knowledge regarding dynamical brain activities, a novel data synthesis strategy is devised to generate large-scale samples for effectively training CedNet. This stands in contrast to traditional ESI methods where the prior information is often enforced via constraints primarily aimed for mathematical convenience. Extensive numerical experiments as well as analysis of a real MEG and epilepsy EEG dataset demonstrate that the DST-CedNet outperforms several state-of-the-art ESI methods in robustly estimating source signals under a variety of source configurations. Gexin Huang, Ke Liu 0008, Jiawen Liang, Zhenghui Gu, Feifei Qi, Yuanqing Li 0001, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Disentangling Writer and Character Styles for Handwriting GenerationabstractTraining machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall writing style, neglecting subtle style inconsistencies between characters written by the same person. For example, while a person's handwriting typically exhibits general uniformity (e.g., glyph slant and aspect ratios), there are still small style variations in finer details (e.g., stroke length and curvature) of characters. In light of this, we propose to disentangle the style representations at both writer and character levels from individual handwritings to synthesize realistic stylized online handwritten characters. Specifically, we present the style-disentangled Transformer (SDT), which employs two complementary contrastive objectives to extract the style commonalities of reference samples and capture the detailed style patterns of each sample, respectively. Extensive experiments on various language scripts demonstrate the effectiveness of SDT. Notably, our empirical findings reveal that the two learned style representations provide information at different frequency magnitudes, underscoring the importance of separate style extraction. Our source code is public at: https://github.com/dailenson/SDT. Gang Dai 0002, Yifan Zhang 0004, Zhu Liang Yu, Zhuoman Liu, Shuangping Huang |
CVPR | 5 |
| 2023 | Hierarchical dynamic movement primitive for the smooth movement of robots based on deep reinforcement learning
Yinlong Yuan, Zhu Liang Yu, Liang Hua, Xiaohu Sang |
Appl. Intell. | 2 |
| 2023 | Sparse Bayesian Learning for End-to-End EEG DecodingabstractDecoding brain activity from non-invasive electroencephalography (EEG) is crucial for brain-computer interfaces (BCIs) and the study of brain disorders. Notably, end-to-end EEG decoding has gained widespread popularity in recent years owing to the remarkable advances in deep learning research. However, many EEG studies suffer from limited sample sizes, making it difficult for existing deep learning models to effectively generalize to highly noisy EEG data. To address this fundamental limitation, this paper proposes a novel end-to-end EEG decoding algorithm that utilizes a low-rank weight matrix to encode both spatio-temporal filters and the classifier, all optimized under a principled sparse Bayesian learning (SBL) framework. Importantly, this SBL framework also enables us to learn hyperparameters that optimally penalize the model in a Bayesian fashion. The proposed decoding algorithm is systematically benchmarked on five motor imagery BCI EEG datasets ( N=192) and an emotion recognition EEG dataset ( N=45), in comparison with several contemporary algorithms, including end-to-end deep-learning-based EEG decoding algorithms. The classification results demonstrate that our algorithm significantly outperforms the competing algorithms while yielding neurophysiologically meaningful spatio-temporal patterns. Our algorithm therefore advances the state-of-the-art by providing a novel EEG-tailored machine learning tool for decoding brain activity. Feifei Qi, David P. Wipf, Tianyou Yu, Yuanqing Li 0001, Yu Zhang 0009, Zhu Liang Yu, Wei Wu 0022 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2022 | CNN-Based Hyperspectral Pansharpening With Arbitrary ResolutionabstractTraditional hyperspectral (HS) pansharpening aims at fusing a HS image with its panchromatic (PAN) counterpart, to bring the spatial resolution of the HS image to that of the PAN image. However, in many practical applications, arbitrary resolution HS (ARHS) pansharpening is required, where the HS and PAN images need to be integrated to generate a pansharpened HS image with arbitrary resolution (usually higher than that of the PAN image). Such an innovative task brings forth new challenges for the pansharpening technique, mainly including how to reconstruct HS images beyond the training scale and how to guarantee spectral fidelity at any spatial resolutions. To tackle the challenges, we present a novel convolutional neural network (CNN)-based method for ARHS pansharpening called ARHS-CNN. It is based on a two-step relay optimization process, which is associated with a multilevel enhancement subnetwork and a rescaling subnetwork. With a careful design following the thread, our ARHS-CNN is able to pansharpen HS images to any spatial resolutions using just a single CNN model trained on a limited number of scales while meantime to keep spectral fidelity at those resolutions, which wins an obvious advantage over traditional pansharpening methods. Experimental results obtained on several datasets verify the excellent performance of our ARHS-CNN method. Lin He 0001, Jun Li 0009, Antonio Plaza, Jocelyn Chanussot, Zhu Liang Yu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Multilayer Neural Dynamics-Based Adaptive Control of Multirotor UAVs for Tracking Time-Varying TasksabstractTo realize the robust control of multirotor unmanned aerial vehicle (UAV) systems, adaptive multilayer neural dynamics (AMND) controllers are proposed and analyzed. The proposed AMND controllers with the strong anti-perturbation property can drive multirotor UAVs to track time-varying tasks and deal with parameter uncertainty problems. First, the design method of the general multilayer neural dynamics (MLND) controllers is introduced and analyzed. Second, based on the design method, the attitude angles, height, and position controllers of a UAV system are designed. Third, according to the adaptive control theory, a novel AMND controller is designed, which can self-tune the parameters of the UAV. Finally, the proposed AMND method applies to a real-world hexrotor UAV system to illustrate its reliability. Mathematical analysis, computer simulations, and experiments verify the reliability, stability, and effectiveness of the proposed controllers which are used to track time-varying tasks. Lunan Zheng, Feiqi Deng, Zhu Liang Yu, Yamei Luo, Zhijun Zhang 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2022 | Dynamic User Activity and Data Detection for Grant-Free NOMA via Weighted ℓ2, 1 MinimizationabstractGrant-free non-orthogonal multiple access (NOMA) has recently received wide attention for reducing signaling overhead and transmission latency in massive machine-type communications (mMTC). In grant-free NOMA systems, user activity and data (UAD) has to be detected, which is challenging in practice. As an emerging technique, compressive sensing (CS) shows great promise in solving this problem by exploiting the inherent sparsity nature of user activity. This paper proposes to use the weighted$\ell _{2, 1}$minimization (WL21M) to jointly detect UAD in realistic dynamic scenarios. At first, the average recoverability of the WL21M is analyzed. This analysis reveals the fact that the WL21M can improve the detection performance by means of an appropriate weighting and the incorporation of intrinsic temporal correlation. Motivated by the analysis, a collaborative hierarchical match pursuit (C-HiMP) algorithm is proposed for dynamic UAD detection. In the C-HiMP, a sequence of WL21M problems are solved in the subspaces spanned by all of the components in the hierarchical estimated support sets, where the weights are collaboratively updated by the solutions in previous time slots so that an attractive self-correction capacity is obtained. Simulation results demonstrate that the proposed C-HiMP can obtain significant performance improvements, in terms of detection accuracy and computational complexity, compared with several state-of-the-art CS-based detection algorithms. Jun Zhang 0026, Zhijing Yang, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2022 | Compressive Sensing-Based Power Allocation Optimization for Energy Harvesting IoT NodesabstractIn this paper, we address the problem of optimizing the power allocation at each time slot for energy-harvesting sensors in an IoT system, where each sensor transmits its observation via a coherent multiple access channel, and thus the observation vector received at the fusion center (FC) becomes a compressed version of the original observations. Our goal is to minimize the number of transmissions required by the high-fidelity reconstruction of the original observations in the FC. In this scenario, the power allocation, coupling with the channel effect, constitute an effective measurement matrix in compressive sensing. Because its performance decides the requirement on the number of transmissions, the goal can be achieved through constructing the measurement matrix as good as possible, or equivalently, solving the power optimization allocation problem, subject to the available energy constraints at each sensor. However, the sensors can only obtain unreliable and intermittent available energy, so that the traditional performance metric, i.e., mutual coherence (MC), of the measurement matrix cannot be directly used to guide the optimization, because an “equal-norm columns” assumption is implicitly required, but not satisfied in our scenario due to the available energy constraints. Moreover, this optimization problem is also non-convex. To overcome these obstacles, we first carry out a distortion analysis based on the generalized MC, which abandons the “equal-norm columns” assumption. The theoretical results indicate that the measurement matrix construction can be formulated as an optimization problem that not only minimizes the MC, but also minimizes the maximum and maximizes the minimum of the column norms of the effective matrix. We further transform this problem into a sequence of surrogate convex problems and iteratively find the solution. Numerical results show that the proposed framework improves the tradeoffs between reconstruction accuracy and the number of transmissions over various power allocation strategies. In some cases, where other strategies achieve a probability of exact recovery of below 0.7, the proposed framework can achieve a more than 0.9 probability. Jun Zhang 0026, Guangfei Xie, Guojun Han, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2021 | fMRI-SI-STBF: An fMRI-informed Bayesian electromagnetic spatio-temporal extended source imaging
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Cuntai Guan |
Neurocomputing | 2 |
| 2021 | Deep Unfolding With Weighted ℓ₂ Minimization for Compressive SensingabstractCompressive sensing (CS) aims to accurately reconstruct high-dimensional signals from a small number of measurements by exploiting signal sparsity and structural priors. However, signal priors utilized in existing CS reconstruction algorithms rely mainly on hand-crafted design, which often cannot offer the best sparsity-undersampling tradeoff because high-order structural priors of signals are hard to be captured in this manner. In this article, a new recovery guarantee of the unified CS reconstruction model-weighted ℓ1minimization (WL1M) is derived, which indicates universal priors could hardly lead to the optimal selection of the weights. Motivated by the analysis, we propose a deep unfolding network for the general WL1M model. The proposed deep unfolding-based WL1M (D-WL1M) integrates universal priors with learning capability so that all of the parameters, including the crucial weights, can be learned from training data. We demonstrate the proposed D-WL1M outperforms several state-of-the-art CS-based methods and deep learning-based methods by a large margin via the experiments on the Caltech-256 image data set. Jun Zhang 0026, Yuanqing Li 0001, Zhu Liang Yu, Zhenghui Gu, Yu Cheng 0010, Huoqing Gong |
IEEE Internet Things J. | 3 |
| 2021 | SDM3d: shape decomposition of multiple geometric priors for 3D pose estimation
Mengxi Jiang, Zhu Liang Yu, Cuihua Li |
Neural Comput. Appl. | 2 |
| 2021 | EEG extended source imaging with structured sparsity and L1-norm residual
Furong Xu, Ke Liu 0008, Zhu Liang Yu, Xin Deng 0003, Guoyin Wang 0001 |
Neural Comput. Appl. | 3 |
| 2021 | Spatiotemporal-Filtering-Based Channel Selection for Single-Trial EEG ClassificationabstractAchieving high classification performance in electroencephalogram (EEG)-based brain-computer interfaces (BCIs) often entails a large number of channels, which impedes their use in practical applications. Despite the previous efforts, it remains a challenge to determine the optimal subset of channels in a subject-specific manner without heavily compromising the classification performance. In this article, we propose a new method, called spatiotemporal-filtering-based channel selection (STECS), to automatically identify a designated number of discriminative channels by leveraging the spatiotemporal information of the EEG data. In STECS, the channel selection problem is cast under the framework of spatiotemporal filter optimization by incorporating a group sparsity constraints, and a computationally efficient algorithm is developed to solve the optimization problem. The performance of STECS is assessed on three motor imagery EEG datasets. Compared with state-of-the-art spatiotemporal filtering algorithms using full EEG channels, STECS yields comparable classification performance with only half of the channels. Moreover, STECS significantly outperforms the existing channel selection methods. These results suggest that this algorithm holds promise for simplifying BCI setups and facilitating practical utility. Feifei Qi, Wei Wu 0022, Zhu Liang Yu, Zhenghui Gu, Zhenfu Wen, Tianyou Yu, Yuanqing Li 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Imaging brain extended sources from EEG/MEG based on variation sparsity using automatic relevance determination
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 2 |
| 2020 | Hyperspectral Image Spectral-Spatial-Range Gabor FilteringabstractSpectral-spatial Gabor filtering, which is based on 3-D local harmonic analysis, has been a powerful spectral-spatial feature extraction tool for hyperspectral image (HSI) classification. However, existing spectral-spatial Gabor approaches are prone to oversmoothing, neglecting the existences of edges and negatively affecting the classification. In this article, we propose a new HSI Gabor filtering concept, called spectral-spatial-range Gabor filtering, which intends to restrain edge interference from disturbing local spectral-spatial harmonic components. Contributions and novelties of our work can be identified as follows: 1) an HSI filtering framework is created, which can accommodate various Gabor filtering procedures and hence offer the potential to guide the design of new Gabor filters; 2) following such a unified filtering framework and taking into consideration both local spectral-spatial harmonic characteristics and range domain variations, we develop a new concept of spectral-spatial-range Gabor filtering; and 3) utilizing this proposed Gabor prototype and elaborating mathematical derivations, we achieve a novel discriminative spectral-spatial-range Gabor filtering method, which can deal with discriminative local harmonics and edge interference simultaneously along the spectral-spatial-range domain, obtaining highly discriminative Gabor features while yielding linear computational complexity. Our novel method is evaluated on four real HSI data sets and achieves excellent performances. Lin He 0001, Chenying Liu 0001, Jun Li 0009, Yuanqing Li 0001, Shutao Li 0001, Zhu Liang Yu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Exemplar-Based Recursive Instance Segmentation With Application to Plant Image AnalysisabstractInstance segmentation is a challenging computer vision problem which lies at the intersection of object detection and semantic segmentation. Motivated by plant image analysis in the context of plant phenotyping, a recently emerging application field of computer vision, this paper presents the Exemplar-Based Recursive Instance Segmentation (ERIS) framework. A three-layer probabilistic model is firstly introduced to jointly represent hypotheses, voting elements, instance labels and their connections. Afterwards, a recursive optimization algorithm is developed to infer the maximum a posteriori (MAP) solution, which handles one instance at a time by alternating among the three steps of detection, segmentation and update. The proposed ERIS framework departs from previous works mainly in two respects. First, it is exemplar-based and model-free, which can achieve instance-level segmentation of a specific object class given only a handful of (typically less than 10) annotated exemplars. Such a merit enables its use in case that no massive manually-labeled data is available for training strong classification models, as required by most existing methods. Second, instead of attempting to infer the solution in a single shot, which suffers from extremely high computational complexity, our recursive optimization strategy allows for reasonably efficient MAP-inference in full hypothesis space. The ERIS framework is substantialized for the specific application of plant leaf segmentation in this work. Experiments are conducted on public benchmarks to demonstrate the superiority of our method in both effectiveness and efficiency in comparison with the state-of-the-art. Jin-Gang Yu, Yansheng Li 0001, Changxin Gao, Hongxia Gao, Gui-Song Xia, Zhu Liang Yu, Yuanqing Li 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Robustness Analysis of a Power-Type Varying-Parameter Recurrent Neural Network for Solving Time-Varying QM and QP Problems and ApplicationsabstractVarying-parameter recurrent neural network, being a special kind of neural-dynamic methodology, has revealed powerful abilities to handle various time-varying problems, such as quadratic minimization (QM) and quadratic programming (QP) problems. In this paper, a novel power-type varying-parameter recurrent neural network (PT-VP-RNN) is proposed to solve the perturbed time-varying QM and QP problems. First, based on the generalization of time-varying QM and QP problems, the design process of the PT-VP-RNN is presented in detail. Second, the robustness performance of the proposed PT-VP-RNN is theoretically analyzed and proved. What is more, two numerical examples are simulated to illustrate the robustness convergence performance of PT-VP-RNN even in a large disturbance condition. Finally, two practical application examples (i.e., a robot tracking example and a venture investment example) further verify the effectiveness, accuracy, and widespread applicability of the proposed PT-VP-RNN. Zhijun Zhang 0003, Lingdong Kong, Lunan Zheng, Pengchao Zhang, Xilong Qu, Bolin Liao, Zhu Liang Yu |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2019 | A SSVEP-Based BCI for Controlling a 4-DOF Robotic ManipulatorabstractIt is difficult to control high degree of freedom (DOF) robotic manipulator based on brain computer interface (BCI). This paper proposes a BCI system which utilizes steady-state visual evoked potential (SSVEP) for controlling a 4-DOF robotic manipulator. In order to elicit SSVEP responses, a stimulator panel circumferentially equipped with 12 LEDs as flickering sources is designed to provide visual stimuli. Applying canonical correlation analysis (CCA) method, elicited SSVEP responses can be decoded as movement commands. Furthermore, the novel commands mapping is designed for controlling end effector to move in 3D space in an efficient and accurate manner. Online experiments were carried out and three subjects participated in this study to perform move-grasp tasks in the simulation environment. Seven out of nine tasks were successfully completed within 50 trials with an average time of 156.71 seconds. The results show that it is feasible and efficient to control 4-DOF robotic manipulator with the proposed SSVEP-based BCI. Canguang Lin, Xiaoyan Deng, Zhu Liang Yu, Zhenghui Gu |
SMC | 3 |
| 2019 | A novel multi-step reinforcement learning method for solving reward hacking
Yinlong Yuan, Zhu Liang Yu, Zhenghui Gu, Xiaoyan Deng, Yuanqing Li 0001 |
Appl. Intell. | 2 |
| 2019 | Reweighted sparse representation with residual compensation for 3D human pose estimation from a single RGB image
Mengxi Jiang, Zhu Liang Yu, Yan Zhang 0059, Qicong Wang, Cuihua Li |
Neurocomputing | 2 |
| 2019 | MMAN: Multi-modality aggregation network for brain segmentation from MR images
Jingcong Li 0001, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 2 |
| 2019 | A novel multi-step Q-learning method to improve data efficiency for deep reinforcement learning
Yinlong Yuan, Zhu Liang Yu, Zhenghui Gu, Yao Yeboah, Wei Wu 0022, Xiaoyan Deng, Jingcong Li 0001, Yuanqing Li 0001 |
Knowl. Based Syst. | 2 |
| 2019 | Classification of symmetric positive definite matrices based on bilinear isometric Riemannian embedding
Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Pattern Recognit. | 2 |
| 2018 | Cartoon-to-Photo Facial Translation with Generative Adversarial NetworksabstractCartoon-to-photo facial translation could be widely used in different applications, such as law enforcement and anime remaking. Nevertheless, current general-purpose image-to-image models \ygyan{usually} %can only produce blurry or unrelated results in this task. In this paper, we propose a Cartoon-to-Photo facial translation with Generative Adversarial Networks (\name) for inverting cartoon faces to generate photo-realistic and related face images. In order to produce convincing faces with intact facial parts, we exploit global and local discriminators to capture global facial features and three local facial regions, respectively. Moreover, we use a specific content network to capture and preserve face characteristic and identity between cartoons and photos. As a result, the proposed approach can generate convincing high-quality faces that satisfy both the characteristic and identity constraints of input cartoon faces. Compared with recent works on unpaired image-to-image translation, our proposed method is able to generate more realistic and correlative images. Junhong Huang, Mingkui Tan, Yuguang Yan, Chunmei Qing, Qingyao Wu, Zhu Liang Yu |
ACML | 6 |
| 2018 | A clustering method based on extreme learning machine
Jinhong Huang, Zhu Liang Yu, Zhenghui Gu |
Neurocomputing | 2 |
| 2018 | A EOG-based switch and its application for "start/stop" control of a wheelchair
Yuanqing Li 0001, Shenghong He, Qiyun Huang, Zhenghui Gu, Zhu Liang Yu |
Neurocomputing | 5 |
| 2018 | Deep learning based on Batch Normalization for P300 signal detection
Mingfei Liu, Wei Wu 0022, Zhenghui Gu, Zhu Liang Yu, Feifei Qi, Yuanqing Li 0001 |
Neurocomputing | 4 |
| 2018 | Variation sparse source imaging based on conditional mean for electromagnetic extended sources
Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001, Srikantan S. Nagarajan |
Neurocomputing | 2 |
| 2018 | Deterministic construction of sparse binary matrices via incremental integer optimization
Jun Zhang 0026, Zhu Liang Yu, Ling Cen, Zhenghui Gu, Zhiping Lin 0001, Yuanqing Li 0001 |
Inf. Sci. | 2 |
| 2018 | A Novel clustering method based on hybrid K-nearest-neighbor graph
Yikun Qin, Zhu Liang Yu, Chang-Dong Wang 0001, Zhenghui Gu, Yuanqing Li 0001 |
Pattern Recognit. | 2 |
| 2018 | A Novel Three-Dimensional P300 Speller Based on Stereo Visual StimuliabstractGoal: P300 spellers are among the most popular types of brain-computer interfaces (BCIs) and are extremely useful assistive devices that enable severely disabled patients to communicate. However, P300 speller performances should be further improved to translate laboratory designs into practical applications. We aimed to design a new speller paradigm that could evoke higher event-related potentials (ERPs) than traditional P300 spellers, thus improving the performance of BCI systems. Methods: We proposed a new P300 speller paradigm based on three-dimensional (3-D) stereo visual stimuli. In this paradigm, flashing buttons are presented in 3-D stereo form. We designed two experiments, one that tested a traditional two-dimensional (2-D) speller and another that tested the proposed 3-D speller. Twelve healthy volunteers participated in our experiments. We compared the ERPs elicited by the 2-D speller and the 3-D speller, and we also compared the classification accuracy, information transfer rate (ITR), and user workload between the two paradigms. Results: The 3-D P300 speller elicited higher amplitudes of P300 waveforms than the traditional 2-D P300 speller. The online experimental results showed that the classification accuracy and the ITR were significantly improved with the 3-D P300 speller. We also found that the user workload of the 3-D P300 speller was significantly lower than that of the 2-D P300 speller. Conclusion : The proposed 3-D P300 speller based on stereo visual stimuli outperformed a traditional 2-D P300 speller. This finding indicates that our 3-D paradigm offers a new method that will improve the performance of P300 BCI systems. Jun Qu, Fei Wang 0026, Zhenping Xia, Tianyou Yu, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 6 |
| 2017 | Deep Learning Method for Sleep Stage Classification
Ling Cen, Zhu Liang Yu, Yun Tang 0003, Wen Shi 0008, Tilmann Kluge, Wee Ser |
ICONIP (2) | 2 |
| 2017 | Batch image alignment via subspace recovery based on alternative sparsity pursuitabstractThe problem of robust alignment of batches of images can be formulated as a low-rank matrix optimization problem, relying on the similarity of well-aligned images. Going further, observing that the images to be aligned are sampled from a union of low-rank subspaces, we propose a new method based on subspace recovery techniques to provide more robust and accurate alignment. The proposed method seeks a set of domain transformations which are applied to the unaligned images so that the resulting images are made as similar as possible. The resulting optimization problem can be linearized as a series of convex optimization problems which can be solved by alternative sparsity pursuit techniques. Compared to existing methods like robust alignment by sparse and low-rank models, the proposed method can more effectively solve the batch image alignment problem, and extract more similar structures from the misaligned images. Xianhui Lin, Zhu Liang Yu, Zhenghui Gu, Jun Zhang 0026, Zhaoquan Cai 0001 |
Comput. Vis. Media | 2 |
| 2017 | An online semi-supervised P300 speller based on extreme learning machine
Zhenghui Gu, Zhu Liang Yu, Yuanqing Li 0001 |
Neurocomputing | 3 |
| 2016 | Sufficient conditions for sparse recovery by weighted ℓ1-constrained quadratic programmingabstractIn this paper, we study the performance guarantee of weighted ℓ1-constrained quadratic programming in recovering the support of a sparse signal from a few linear measurements. A new sufficient condition for the success of weighted ℓ1-constrained quadratic programming is derived. Further, we demonstrate its applications in two typical weighted ℓ1models: Firstly, a theoretical result on the modified-BPDN can be generated directly, which indicates that if a partial support information is available and more than half of the information is accurate, then modified-BPDN relaxes the sufficient condition of BPDN ensuring the support recovery for any sparse signals. Secondly, the proposed condition offers a criterion to compute the optimal weights of weighted ℓ1minimization for nonuniform sparse models. Some simulations are carried out to validate its effectiveness. Jun Zhang 0026, Yuanqing Li 0001, Zhu Liang Yu, Zhenghui Gu |
IJCNN | 3 |
| 2016 | A real-time visual object tracking system based on Kalman filter and MB-LBP feature matching
Zebin Cai, Zhenghui Gu, Zhu Liang Yu, Hao Liu 0040, Ke Zhang 0016 |
Multim. Tools Appl. | 3 |
| 2016 | Multimodal BCIs: Target Detection, Multidimensional Control, and Awareness Evaluation in Patients With Disorder of ConsciousnessabstractDespite rapid advances in the study of brain–computer interfaces (BCIs) in recent decades, two fundamental challenges, namely, improvement of target detection performance and multidimensional control, continue to be major barriers for further development and applications. In this paper, we review the recent progress in multimodal BCIs (also called hybrid BCIs), which may provide potential solutions for addressing these challenges. In particular, improved target detection can be achieved by developing multimodal BCIs that utilize multiple brain patterns, multimodal signals, or multisensory stimuli. Furthermore, multidimensional object control can be accomplished by generating multiple control signals from different brain patterns or signal modalities. Here, we highlight several representative multimodal BCI systems by analyzing their paradigm designs, detection/control methods, and experimental results. To demonstrate their practicality, we report several initial clinical applications of these multimodal BCI systems, including awareness evaluation/detection in patients with disorder of consciousness (DOC). As an evolving research area, the study of multimodal BCIs is increasingly requiring more synergetic efforts from multiple disciplines for the exploration of the underlying brain mechanisms, the design of new effective paradigms and means of neurofeedback, and the expansion of the clinical applications of these systems. Yuanqing Li 0001, Jiahui Pan 0003, Jinyi Long, Tianyou Yu, Fei Wang 0026, Zhu Liang Yu, Wei Wu 0022 |
Proc. IEEE | 6 |
| 2015 | STRAPS: A Fully Data-Driven Spatio-Temporally Regularized Algorithm for M/EEG Patch Source ImagingabstractFor M/EEG-based distributed source imaging, it has been established that the L2-norm-based methods are effective in imaging spatially extended sources, whereas the L1-norm-based methods are more suited for estimating focal and sparse sources. However, when the spatial extents of the sources are unknown a priori, the rationale for using either type of methods is not adequately supported. Bayesian inference by exploiting the spatio-temporal information of the patch sources holds great promise as a tool for adaptive source imaging, but both computational and methodological limitations remain to be overcome. In this paper, based on state-space modeling of the M/EEG data, we propose a fully data-driven and scalable algorithm, termed STRAPS, for M/EEG patch source imaging on high-resolution cortices. Unlike the existing algorithms, the recursive penalized least squares (RPLS) procedure is employed to efficiently estimate the source activities as opposed to the computationally demanding Kalman filtering/smoothing. Furthermore, the coefficients of the multivariate autoregressive (MVAR) model characterizing the spatial-temporal dynamics of the patch sources are estimated in a principled manner via empirical Bayes. Extensive numerical experiments demonstrate STRAPS's excellent performance in the estimation of locations, spatial extents and amplitudes of the patch sources with varying spatial extents. Ke Liu 0008, Zhu Liang Yu, Wei Wu 0022, Zhenghui Gu, Yuanqing Li 0001 |
Int. J. Neural Syst. | 2 |
| 2015 | Analysis of fMRI data based on sparsity of source components in signal dictionary
Bao Feng, Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001 |
Neurocomputing | 2 |
| 2015 | Energy-Efficient ECG Compression on Wireless Biosensors via Minimal Coherence Sensing and Weighted ℓ1 Minimization ReconstructionabstractLow energy consumption is crucial for body area networks (BANs). In BAN-enabled ECG monitoring, the continuous monitoring entails the need of the sensor nodes to transmit a huge data to the sink node, which leads to excessive energy consumption. To reduce airtime over energy-hungry wireless links, this paper presents an energy-efficient compressed sensing (CS)-based approach for on-node ECG compression. At first, an algorithm called minimal mutual coherence pursuit is proposed to construct sparse binary measurement matrices, which can be used to encode the ECG signals with superior performance and extremely low complexity. Second, in order to minimize the data rate required for faithful reconstruction, a weighted ℓ1 minimization model is derived by exploring the multisource prior knowledge in wavelet domain. Experimental results on MIT-BIH arrhythmia database reveals that the proposed approach can obtain higher compression ratio than the state-of-the-art CS-based methods. Together with its low encoding complexity, our approach can achieve significant energy saving in both encoding process and wireless transmission. Jun Zhang 0026, Zhenghui Gu, Zhu Liang Yu, Yuanqing Li 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2012 | Discriminative dictionary learning for EEG signal classification in Brain-computer interfaceabstractModern Brain-computer interface (BCI) technique is essentially based on the classification of the brain signals. The sparse representation classification (SRC) method has been studied for classifying EEG signals of the motor imagery based BCI. The dictionary used in the SRC method is the simple combination of feature vectors which are extracted from the EEG signal-trials by common spatial pattern (CSP) algorithm. In this paper, we propose a method to learn a new dictionary with smaller size and more discriminative ability for the classification. The proposed method, discriminative dictionary learning (DDL), is based on minimizing an objective function containing a reconstructive term and a discriminative term. We apply an iterative scheme to the optimization and transform it to a series of mixed ℓ1-ℓ2 optimizations, which are solved based on separable surrogate functions (SSF) technique. We evaluate the proposed method using the dataset from BCI competition III. The experimental results show that the proposed method outperforms the SRC method. Ya Yang, Zhu Liang Yu |
ICARCV | 3 |
| 2011 | Speech Emotion Recognition System Based on L1 Regularized Linear Regression and Decision Fusion
Ling Cen, Zhu Liang Yu, Minghui Dong |
ACII (2) | 2 |
| 2010 | Feature extraction with multiscale autoregression of multichannel time series for P300 speller BCIabstractP300 is one of the most studied components of event related potentials which reflects the responses of brain to events in the external environment. In this paper, we present a new method that utilizes multiresolution autoregression of multichannel time series (MAMTS) for feature extraction of P300 wave. First, it adopts multiresolution autoregression on dyadic tree to depict the characteristic of electroencephalogram (EEG) signal. Then the corresponding autoregression noise of multichannel time series is extracted as the feature. The experiment results verified the effectiveness of this new feature for P300 speller brain compute interface (BCI). Lin He 0001, Zhenghui Gu, Yuanqing Li 0001, Zhu Liang Yu |
ICASSP | 4 |
| 2010 | Robust adaptive beamformer with a large controlled mainlobeabstractMany advanced adaptive beamformers are robust against arbitrary array steering vector (ASV) mismatches within a presumed uncertainty set. Adaptive array tolerating significant steering direction error usually requires a large size of ASV uncertainty set. In such case, however, the output signal-to-interference-plus-noise ratios (SINRs) of robust methods degrade quickly with the increasing size of the uncertainty set. In this paper, we propose a new compact ASV uncertainty set which is modelled explicitly by the uncertainty on steering direction and the other arbitrary ASV errors. A robust adaptive beamformer is derived based on this new ASV uncertainty set. To eliminate the non-convex constraint on array magnitude response, we force the real part of array response to exceed unity regarding the ASVs within the uncertainty set. Furthermore, using the worst-case optimization technique, the resultant beamformer is formulated as a quadratic optimization problem with semi-infinite second-order cone (SOC) constraints. Numerical studies show that a large robust response region is easy to achieve and the resultant beamformer achieves high performance on SINR enhancement. Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001, Wee Ser, Meng Hwa Er |
ICASSP | 1 |
| 2010 | Linear sparse array synthesis via convex optimizationabstractDue to the constraint on half-wavelength inter-element spacing of a uniformly spaced array, sparse arrays are usually designed to be non-uniformly spaced. Through proper design, unequally-spaced sparse arrays can have higher spatial resolution, lower sidelobe and less number of sensors required in comparison with uniformly spaced arrays. However, the synthesis of a sparse array is a non-convex process as the array beampattern is an exponential or trigonometric function of sensor positions. In this paper, we propose a synthesis scheme, where the design of sparse arrays is formulated as a convex optimization problem. The array weights are optimized to achieve minimum Peak Sidelobe Level (PSL) as well as maximizing the sparsity of the array by optimizing an objective function that includes two terms, one measures the PSL and the other measures the sparsity of array. Sparse array is then obtained by removing those sensors with weights approximately equal to zero. The proposed design scheme eliminates the need of optimization according to sensor positions, which consequently solves the problem in non-convex optimization that cannot guarantee to find the optimum solution with reasonable computation time. Numerical studies show that it can be successfully applied to sparse array synthesis with low computational complexity. Moreover, lower PSL and higher resolution of mainlobe can be achieved in comparison with a uniformly spaced array. Ling Cen, Wee Ser, Wei Cen, Zhu Liang Yu |
ISCAS | 4 |
| 2010 | Robust response control with linear inequality matrix constraints for adaptive beamformerabstractA novel robust adaptive beamformer, with new robust constraints on array magnitude response was proposed by utilizing the autocorrelation sequence of array weight vector and the worst-case optimization technique. The proposed adaptive beamformer was formulated as a linear programming problem with second-order cone semi-infinite constraints, which can be eliminated by using the sampling technique. In this paper, we transform these semi-infinite second-order cone constraints into some norm constraints and linear matrix inequality (LMI) constraints. The advantage of this new formulation of the problem is that the sampling of the angles is avoided. The exact optimal result of the problem can be obtained instead of the approximated one provided by sampling technique. The resultant beamformer possesses superior robustness against arbitrary array imperfections and high performance on signal-to-interference-plus-noise ratio (SINR) enhancement even with a large controlled robust response region. Zhu Liang Yu, Zhenghui Gu, Yuanqing Li 0001, Wee Ser, Meng Hwa Er |
ISCAS | 1 |
| 2009 | Voxel selection in fMRI data analysis: A sparse representation methodabstractThis paper proposes an iterative sparse representation-based algorithmfor voxel selection in functionalmagnetic resonance imaging (fMRI) data. The output of the algorithm is a sparse weight vector, of which the magnitude of each entry represents the significance of its corresponding voxel with respect to mental tasks or stimulus. To demonstrate the validity of our algorithm and illustrate its application, we apply this algorithm to the Pittsburgh Brain Activity Interpretation Competition (PBAIC) 2007 fMRI data set for selecting the voxels which are the most relevant to the tasks of the subjects. Compared with three baseline methods, general linear model (GLM)-based statistical parametric mapping (SPM), correlation method and mutual information method, our method shows satisfactory performance for voxel selection. Yuanqing Li 0001, Zhu Liang Yu, Praneeth Namburi, Cuntai Guan |
ICASSP | 2 |
| 2008 | An improved genetic algorithm for aperiodic array synthesisabstractIn this paper, a novel algorithm on beam pattern synthesis for linear aperiodic arrays with arbitrary geometrical configuration is proposed. The algorithm is based on an improved genetic algorithm (IGA) that simultaneously adjusts the weight coefficients and inter-sensor spacings of a linear aperiodic array. A novel section-based crossover and a self-supervised mutation process are developed to improve the convergence performance. The results from simulation illustrate that with the IGA, the peak sidelobe level (PSL) of the synthesized beam pattern has been successfully lowered. In addition, the computational cost of the proposed algorithm can be as low as being about 10% of that of a recently reported genetic algorithm based synthesis method. The robustness of the proposed IGA has been illustrated clearly from the statistic of multiple independent runs too. The excellent performance of the IGA makes it a promising optimization algorithm where expensive cost functions are involved. Ling Cen, Wee Ser, Zhu Liang Yu, Susanto Rahardja |
ICASSP | 3 |
| 2008 | Novel Adaptive Antenna Array Based on Robust Semidefinite ProgrammingabstractIn this paper, a novel robust adaptive beamformer is proposed based on semidefinite programming (SDP) and worst- case optimization. With the SDP formulation, the array output power and magnitude response can be expressed as linear functions. New constraints on magnitude response are introduced in the adaptive array. The proposed method can flexibly control the robust response region with a specific beamwidth and response ripple. In practical applications, the array suffers from having not only steering direction error, but also many other array imperfections. To make the adaptive beamformer robust against all kinds of array imperfections, the worst-case optimization technique is proposed to reconstruct the beamformer. By minimizing the array output power with respect to the worst-case effect of array imperfections, the resultant beamformer possesses superior robustness against arbitrary array imperfections. Since the constraints on magnitude response are inequality constraints, most of them are inactive in the optimization process so that few degrees of freedom (DOFs) of the adaptive beamformer are consumed. Consequently, the resultant beamformer has high performance on signal-to-interference-plus-noise ratio (SINR) improvement. Simple implementation, flexible performance control as well as significant SINR enhancement support the practicability of the proposed method. Zhu Liang Yu, Wee Ser, Meng Hwa Er |
ICC | 1 |
| 2008 | Robust Adaptive Beamformer with LMI Constraints on Magnitude ResponseabstractIn this paper, a novel robust adaptive antenna array with new linear matrix inequality (LMI) constraints on the array magnitude response is proposed. Most of the advanced robust adaptive beamformers, which are robust against arbitrary array steering vector (ASV) errors within a presumed uncertainty set, have poor performance facing a large uncertainty on direction-of-arrival (DOA). Since the DOA uncertainty may result in a large error in the ASV, a big uncertainty set is required to make the adaptive array robust against a large DOA error. The adverse effect is the degraded output signal-to-interference- plus-noise ratio (SINR). In this paper, a compact model of ASV uncertainty set is described by the uncertainties on DOA and other arbitrary ASV errors explicitly. Based on this new ASV uncertainty set, a novel robust adaptive beamformer is derived. In order to eliminate the semi-infinite constraint in the proposed beamformer, new LMI constraints are derived. The proposed beamformer possesses superior robustness against arbitrary array imperfections as well as a large DOA error. A large robust response region is easy to achieve and the resultant beamformers still have high performance on SINR enhancement. Zhu Liang Yu, Wee Ser, Meng Hwa Er |
ICC | 1 |
| 2008 | Speech Emotion Recognition Using Canonical Correlation Analysis and Probabilistic Neural NetworkabstractIn this paper, automatic identification of emotional states from human speech is addressed. While several papers have been published in the literature on speech emotion recognition, the features used are taken or modified from those used for speech recognition purposes. However, not all features used for speech recognition are of equal importance for emotion recognition. This paper addresses this issue and proposes a systematic method on feature selection for emotion recognition from speech signals. The idea is to work on a well-selected small feature set and use it to remove irrelevant information. Specifically, the proposed method uses the similar idea of the Canonical Correlation Analysis (CCA) to estimate the linear relationship between the various features and the emotional states. The outcome is a set of features that are of most relevance to the emotions. Experiments have been conducted using the LDC database and with the use of the Probabilistic Neural Network (PNN) as the classification method. The results obtained show that, comparable accuracies can be obtained for the emotional states tested with the use of only about 30% of the features considered. This implies that the computational load can be reduced greatly too. Ling Cen, Wee Ser, Zhu Liang Yu |
ICMLA | 3 |
| 2008 | A Hybrid PNN-GMM classification scheme for speech emotion recognitionabstractWith the increasing demand for spoken language interfaces in human-computer interactions, automatic recognition of emotional states from human speeches has become of increasing importance. In this paper, we propose a novel hybrid scheme that combines the probabilistic neural network (PNN) and the Gaussian mixture model (GMM) for identifying emotions from speech signals. In order to handle mismatches more effectively, the universal background model (UBM) is incorporated into the GMM, and the resultant model is denoted as UBM-GMM. In the hybrid scheme, the strengths of the PNN and the UBM-GMM are combined through a novel conditional-probability based fusion algorithm. Experimental results show that the proposed scheme is able to achieve higher recognition accuracy than that obtained by using PNN or UBM-GMM alone. Wee Ser, Ling Cen, Zhu Liang Yu |
ICPR | 3 |
| 2008 | Robust adaptive beamformers with linear matrix inequality constraintsabstractIn this paper, a novel robust adaptive beamformers with linear matrix inequality (LMI) constraints on magnitude response is proposed. The recently proposed robust adaptive beamformer has a drawback that some of the constraints are semi-infinite. Although the sampling technique provides a good approximation of the optimization problem, the cost is the heavy computation load. In order to overcome this problem, new LMI constraints are derived in this paper to replace the semi- infinite constraints so that the sampling is avoid and an exact solution can be obtained. Zhu Liang Yu, Wee Ser, Meng Hwa Er |
ISCAS | 1 |
| 2008 | Spectral factorization for integer-interval sampled sequence and its applications in array processing
Zhu Liang Yu, Meng Hwa Er, Wee Ser, Zhenghui Gu |
Signal Process. | 1 |
| 2008 | Robust response control for adaptive beamformers against arbitrary array imperfections
Zhu Liang Yu, Wee Ser, Meng Hwa Er, Zhenghui Gu |
Signal Process. | 1 |
| 2006 | A Robust Capon Beamformer with New Uncertainty Constraint on Steering VectorabstractA robust Capon beamformer (RCB) with a new constraint on the uncertainty of nominal array steering vector (ASV) is proposed in this paper. The new constraint is constructed by replacing the nominal ASV with a projected one onto the signal-plus-interference subspace. The proposed RCB achieves higher output signal-to-noise-plus-interference ratio (SINR) compared with the conventional RCBs. Theoretical analysis and simulation results show the effectiveness of the proposed method. Zhu Liang Yu, Meng Hwa Er |
ICASSP (4) | 1 |
| 2006 | QR-RLS Based Minimum Variance Distortionless Responses BeamformerabstractIn this paper, a QR-RLS based Minimum Variance Distortionless Responses (MVDR) method, and its systolic array processor, are proposed. The QR-RLS based MVDR has many advantages, such as numerical stability, computational efficiency and pipelined structure in implementation. We also point out that the conventional method, MVDR using QR-RLS method by directly forcing the desired signal to zero is not correct. Numerical experiments are carried out to illustrate the effectiveness of the proposed method. Zhu Liang Yu, Wee Ser, Susanto Rahardja |
ICASSP (3) | 1 |
| 2006 | A fast fading channel identification method for OFDM communication systemabstractA fading channel identification and tracking method for orthogonal frequency division multiplexing (OFDM) communication system is proposed in this paper. The fading channel is modeled as an auto-regression (AR) process with unknown AR parameters. With the AR model of channel, a state-space representation of the the system is formulated. The system parameters are identified using subspace based method and Kalman filtering. With training symbols and feedback decisions, the system model is simplified to have fixed parameters contrast with other subspace based identification methods. Moreover, QR based subspace method is proposed to solve this problem efficiently. Numerical simulation results show that the proposed method can be applied in fast fading environment. Zhu Liang Yu, Wee Ser |
IWCMC | 1 |
| 2006 | A robust minimum variance beamformer with new constraint on uncertainty of steering vector
Zhu Liang Yu, Meng Hwa Er |
Signal Process. | 1 |
| 2004 | Multiscale corner detection for gray level images using Plessey methodabstractThis paper proposes an improved Plessey corner detection method for gray level images using multiscale analysis. Plessey corner detector is well known for its good performance. But, as we understand, Plessey corner detection method has three drawbacks. Firstly, it works only in the spatial domain, so it can only detect corners belonging to a specific scale. On the other hand, the proposed detector works in the scale-space domain, thus, it can detect the corner points belonging to different scales. Secondly, three parameters are needed to be set manually in the original Plessey method, while the proposed method requires to set only one parameter. Thirdly, the delocalization is a well-known inherent problem for the Plessey corner operator. And it can be increased with the scale at which it operates. The proposed algorithm partially solves this problem by detecting the corners from smaller scale to larger scales. Our proposed multiscale scheme can also be applied to other spatial corner detectors to improve their performances. Better simulation results are shown and compared with the original Plessey and SUSAN corner detectors. Xinting Gao, Zhu Liang Yu, Farook Sattar, Ronda Venkateswarlu |
ICARCV | 2 |
| 2004 | A robust adaptive blind multichannel identification algorithm for acoustic applicationsabstractWe propose a robust adaptive blind multichannel identification algorithm in the frequency domain. It utilizes the fast Fourier transform (FFT) to reduce computational complexity when the channel impulse response (IR) is long. Moreover, the Newton-LMS algorithm is obtained in the frequency domain with small computational load to improve the convergence speed. The advantage of the proposed method is its robustness to input noise, especially when the channel IR is long, e.g., the room acoustic IR with a length up to hundreds or thousands taps. The conventional methods cannot obtain an estimate with acceptable accuracy and low computational load. The situation becomes worse when the input signal-to-noise ratio (SNR) is low. Simulation results show that the proposed method is suitable for estimating long multichannel IRs in practical environments. Zhu Liang Yu, Meng Hwa Er |
ICASSP (2) | 1 |
| 2004 | Blind multichannel identification for speech dereverberation and enhancementabstractA multichannel wideband signal dereverberation and enhancement method is proposed in this paper. It uses the blindly estimated impulse responses (IR) relating the signal source and each sensor to form a multiple input inverse filter (MINT) for speech dereverberation and enhancement with extended generalized sidelobe canceller (GSC). With the replacement of MINT for fixed beamformer and modification of blocking matrix in the conventional GSC, the resulting extended GSC not only dereverberates the distorted target signal, but also suppresses the interference/noise. Computer simulation results show the effectiveness of the proposed method. Zhu Liang Yu, Meng Hwa Er |
ICASSP (4) | 1 |
| 2004 | A robust algorithm for linearly constrained adaptive beamformingabstractA new approach to robust adaptive beamforming for wideband array signals is proposed. General steering vector errors, such as direction-of-arrival mismatch and array positional error, are modeled by "time-delay errors" and compensated for by self-adjusted interpolation filtering. The proposed method effectively overcomes the target-signal cancellation problem without suffering from loss in the degree of freedom for interference rejection, as verified by simulations. Qiyue Zou, Zhu Liang Yu, Zhiping Lin 0001 |
IEEE Signal Process. Lett. | 2 |