VLDB 2026 Research / reviewers in the wild / expert
Yueming Wang 0001
dblp:01/3962-1
· DBLP profile ↗
58ranked-venue papers
5as first author
25since 2021 · last 2026
0000-0001-7742-0722ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Synthesis to Clinical Assistance: A Strategy-Aware Agent Framework for Autism Intervention based on Real Clinical DatasetabstractJunhong Lai, Shuzhong Lai, Yanhao Yu, Wanlin Chen, Chenyu Yan, Haifeng Li, Lin Yao, Yueming Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junhong Lai, Shuzhong Lai, Yanhao Yu, Wanlin Chen, Chenyu Yan, Lin Yao 0002, Yueming Wang 0001 |
ACL (1) | 8 |
| 2025 | Self-Attentive Spatio-Temporal Calibration for Precise Intermediate Layer Matching in ANN-to-SNN DistillationabstractSpiking Neural Networks (SNNs) are promising for low-power computation due to their event-driven mechanism but often suffer from lower accuracy compared to Artificial Neural Networks (ANNs). ANN-to-SNN knowledge distillation can improve SNN performance, but previous methods either focus solely on label information, missing valuable intermediate layer features, or use a layer-wise approach that neglects spatial and temporal semantic inconsistencies, leading to performance degradation. To address these limitations, we propose a novel method called self-attentive spatio-temporal calibration (SASTC). SASTC uses self-attention to identify semantically aligned layer pairs between ANN and SNN, both spatially and temporally. This enables the autonomous transfer of relevant semantic information. Extensive experiments show that SASTC outperforms existing methods, effectively solving the mismatching problem. Superior accuracy results include 95.12% on CIFAR-10, 79.40% on CIFAR-100 with 2 time steps, and 68.69% on ImageNet with 4 time steps for static datasets, and 97.92% on DVS-Gesture and 83.60% on DVS-CIFAR10 for neuromorphic datasets. This marks the first time SNNs have outperformed ANNs on both CIFAR-10 and CIFAR-100, shedding the new light on the potential applications of SNNs. Di Hong, Yueming Wang 0001 |
AAAI | 2 |
| 2025 | Cauchy Diffusion: A Heavy-tailed Denoising Diffusion Probabilistic Model for Speech SynthesisabstractDenoising diffusion probabilistic models (DDPMs) have gained popularity in devising neural vocoders and obtained outstanding performance. However, existing DDPM-based neural vocoders struggle to handle the prosody diversities due to their susceptibility to mode-collapse issues confronted with imbalanced data. We introduced Cauchy Diffusion, a model incorporating the Cauchy noises to address this challenge. The heavy-tailed Cauchy distribution exhibits better resilience to imbalanced speech data, potentially improving prosody modeling. Our experiments on the LJSpeech and VCTK datasets demonstrate that Cauchy Diffusion achieved state-of-the-art speech synthesis performance. Compared to existing neural vocoders, our Cauchy Diffusion notably improved speech diversity while maintaining superior speech quality. Remarkably, Cauchy Diffusion surpassed neural vocoders based on generative adversarial networks (GANs) that are explicitly optimized to improve diversity. Qi Lian, Yueming Wang 0001 |
AAAI | 3 |
| 2025 | DeCorrNet: Enhancing Neural Decoding Performance by Eliminating Correlations in NoiseabstractNeural decoding, which transforms neural signals into motor commands, plays a key role in brain-computer interfaces (BCIs). Existing neural decoding approaches mainly rely on the assumption of independent noises, which could perform poorly in case the assumption is invalid. However, correlations in noises have been commonly observed in neural signals. Specifically, noise in different neural channels can be similar or highly related, which could degrade the performance of those neural decoders. To tackle this problem, we propose the DeCorrNet, which explicitly removes noise correlation in neural decoding. DeCorrNet could incorporate diverse neural decoders as an ensemble module to enhance the neural decoding performance. Experiments with benchmark BCI datasets demonstrated the superiority of DeCorrNet and achieved state-of-the-art results. Xianhan Tan, Yueming Wang 0001 |
AAAI | 3 |
| 2025 | Complementary-Disentangled Neural Generalization: a Robust Framework for Stable Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) enable communication between the brain and the external environment, showing significant potential in restoration, rehabilitation and movement enhancement. However, neural drift causes BCI performance to degrade substantially over time, compromising their long-term reliability. A fundamental limitation of current methods is their failure to account for a key insight from neural preference theory that the magnitude of neural drift depends on specific motor parameters (e.g., velocity, direction, and speed), ultimately compromising performance. To overcome the limitation, we introduce a novel framework named ComplementaryDisentangled Neural Generalization (CDNG) inspired by neural preference theory. Specifically, we first conduct pre-experiments about the neural decoding preference, revealing that neural drifts differ across velocity, speed and direction. Then we adopt CDNG, which captures invariant neural representations through an ensemble of three specialized neural decoders after disentangling velocity into speed and direction. Extensive experiments on several datasets demonstrate that our method achieves state-of-the-art performance and significantly enhances cross-day generalization. Jiyu Wei, Dazhong Rong, Di Hong, Zhanjie Zhang, Xinyun Zhu, Qinming He, Yueming Wang 0001 |
BIBM | 7 |
| 2025 | Improving Unsupervised Task-driven Models of Ventral Visual Stream via Relative Position Predictivity
Dazhong Rong, Jiyu Wei, Di Hong, Yaoyao Hao, Qinming He, Yueming Wang 0001 |
CogSci | 8 |
| 2025 | Bridging the Gap Between Brain and Machine in Interpreting Visual Semantics: Towards Self-Adaptive Brain-to-Text Decoding
Yueming Wang 0001, Gang Pan 0001 |
ICCV | 3 |
| 2025 | Flow Matching for Few-Trial Neural Adaptation with Stable Latent DynamicsabstractThe primary goal of brain-computer interfaces (BCIs) is to establish a direct linkage between neural activities and behavioral actions via neural decoders. Due to the nonstationary property of neural signals, BCIs trained on one day usually obtain degraded performance on other days, hindering the user experience. Existing studies attempted to address this problem by aligning neural signals across different days. However, these neural adaptation methods may exhibit instability and poor performance when only a few trials are available for alignment, limiting their practicality in real-world BCI deployment. To achieve efficient and stable neural adaptation with few trials, we propose Flow-Based Distribution Alignment (FDA), a novel framework that utilizes flow matching to learn flexible neural representations with stable latent dynamics, thereby facilitating source-free domain alignment through likelihood maximization. The latent dynamics of FDA framework is theoretically proven to be stable using Lyapunov exponents, allowing for robust adaptation. Further experiments across multiple motor cortex datasets demonstrate the superior performance of FDA, achieving reliable results with fewer than five trials. Our FDA approach offers a novel and efficient solution for few-trial neural data adaptation, offering significant potential for improving the long-term viability of real-world BCI applications. Puli Wang, Yueming Wang 0001, Gang Pan 0001 |
ICML | 3 |
| 2025 | ASD-Chat: An Innovative Dialogue Intervention System for Children with Autism based on LLM and VB-MAPPabstractThe main problem of children with autism spectrum disorder (ASD) is communication barriers, resulting in social impairment in their daily lives. However, previous research on autism dialogue intervention computer-assisted systems overlooks two important aspects: paradigm design lacks a theoretical foundation in clinical intervention methods, and evaluation of system effectiveness is based solely on behavioral performance, neglecting physiological signal. Based on the aforementioned points, we proposed ASD-Chat, a social intervention system based on VB-MAPP (Verbal Behavior Milestones Assessment and Placement Program) powered by LLM (Large Language Model) for dialogue generation. We used the same paradigm for 12 ASD children to engage in thematic dialogues with both ASD-Chat and professional interventionists. Behavioral performance and physiological signal analysis results indicate that the ASD-Chat system achieves competitive intervention effects comparable to those of professional interventionists. This indicates that the dialogue paradigm and prototype system we designed have the potential for long-term intervention in the future. Chengyun Deng, Shuzhong Lai, Mengyi Bao, Lin Yao 0002, Yueming Wang 0001 |
IJCNN | 8 |
| 2025 | CRRL: Learning Channel-invariant Neural Representations for High-performance Cross-day DecodingabstractBrain-computer interfaces have shown great potential in motor and speech rehabilitation, but still suffer from low performance stability across days, mostly due to the instabilities in neural signals. These instabilities, partially caused by neuron deaths and electrode shifts, leading to channel-level variabilities among different recording days. Previous studies mostly focused on aligning multi-day neural signals of onto a low-dimensional latent manifold to reduce the variabilities, while faced with difficulties when neural signals exhibit significant drift. Here, we propose to learn a channel-level invariant neural representation to address the variabilities in channels across days. It contains a channel-rearrangement module to learn stable representations against electrode shifts, and a channel reconstruction module to handle the missing neurons. The proposed method achieved the state-of-the-art performance with cross-day decoding tasks over two months, on multiple benchmark BCI datasets. The proposed approach showed good generalization ability that can be incorporated to different neural networks. Xianhan Tan, Binli Luo, Yueming Wang 0001 |
NeurIPS | 4 |
| 2025 | LaSNN: Layer-wise ANN-to-SNN distillation for effective and efficient training in deep spiking neural networks
Di Hong, Yueming Wang 0001 |
Neurocomputing | 3 |
| 2025 | MindGPT: Interpreting What You See With Non-Invasive Brain RecordingsabstractDecoding of seen visual contents with non-invasive brain recordings has important scientific and practical values. Efforts have been made to recover the seen images from brain signals. However, most existing approaches cannot faithfully reflect the visual contents due to insufficient image quality or semantic mismatches. Compared with reconstructing pixel-level visual images, speaking is a more efficient and effective way to explain visual information. Here we introduce a non-invasive neural decoder, termed MindGPT, which interprets perceived visual stimuli into natural languages from functional Magnetic Resonance Imaging (fMRI) signals in an end-to-end manner. Specifically, our model builds upon a visually guided neural encoder with a cross-attention mechanism. By the collaborative use of data augmentation techniques, this architecture permits us to guide latent neural representations towards a desired language semantic direction in a self-supervised fashion. Through doing so, we found that the neural representations of the MindGPT are explainable, which can be used to evaluate the contributions of visual properties to language semantics. Our experiments show that the generated word sequences truthfully represented the visual information (with essential details) conveyed in the seen stimuli. The results also suggested that with respect to language decoding tasks, the higher visual cortex (HVC) is more semantically informative than the lower visual cortex (LVC), and using only the HVC can recover most of the semantic information. The source code for the MindGPT model is publicly available at https://github.com/JxuanC/MindGPT. Jiaxuan Chen 0002, Yueming Wang 0001, Gang Pan 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Dynamic Instance-Level Graph Learning Network of Intracranial Electroencephalography Signals for Epileptic Seizure PredictionabstractBrain-computer interface (BCI) technology is emerging as a valuable tool for diagnosing and treating epilepsy, with deep learning-based feature extraction methods demonstrating remarkable progress in BCI-aided systems. However, accurately identifying causal relationships in temporal dynamics of epileptic intracranial electroencephalography (iEEG) signals remains a challenge. This paper proposes a Dynamic Instance-level Graph Learning Network (DIGLN) for seizure prediction using iEEG signals. The DIGLN comprises two core components: a grouped temporal neural network that extracts node features and a graph structure learning method to capture the causality from intra-channel to inter-channel. Furthermore, we propose a graphical interactive writeback technique to enable DIGLN to capture the causality from inter-channel to intra-channel. Consequently, our DIGLN enables patient-specific dynamic instance-level graph learning, facilitating the modelling of evolving signals and functional connectivities through end-to-end data-driven learning. Experimental results on the Freiburg iEEG dataset demonstrate the superior performance of DIGLN, surpassing other deep learning-based seizure prediction methods. Visualization results further confirm DIGLN's capability to learn interpretable and diverse connections. Qi Lian, Yueming Wang 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | A Habenula Neural Biomarker Simultaneously Tracks Weekly and Daily Symptom Variations During Deep Brain Stimulation Therapy for DepressionabstractOBJECTIVE: Deep brain stimulation (DBS) targeting the lateral habenula (LHb) is a promising therapy for treatment-resistant depression (TRD) but its clinical effect has been variable, which can be improved by adaptive DBS (aDBS) guided by a neural biomarker of depression symptoms. Existing neural biomarkers, however, cannot simultaneously track slow and fast symptom dynamics, do not sufficiently respond to stimulation parameters, and lack neurobiological interpretability, which hinder their use in developing aDBS. METHODS: We conducted a study on one TRD patient who achieved remission following a 41-week LHb DBS treatment, during which we assessed slow symptom variations using weekly clinical ratings and fast variations using daily self-reports. We recorded daily LHb local field potentials (LFP) concurrently with the reports during the entire treatment process. We then used machine learning methods to identify a personalized depression neural biomarker from spectral and temporal LFP features. RESULTS: The neural biomarker was identified from classification of high and low depression symptom states with a cross-validated accuracy of 0.97. It further simultaneously tracked both weekly (slow) and daily (fast) depression symptom variation dynamics, achieving test data explained variance of 0.74 and 0.63 respectively and responded to DBS frequency alterations. Finally, it can be neurobiologically interpreted as indicating LHb excitatory and inhibitory balance changes during DBS treatment. CONCLUSION: By collecting and analyzing a unique personalized dataset of weekly and daily LFP recordings and symptom evaluations, we identified a high-performance neural biomarker for depression during LHb DBS. SIGNIFICANCE: Our results hold promise to facilitate future aDBS for treating TRD. Shaohua Hu, Junming Zhu, Hemmings Wu, Hailan Hu, Yueming Wang 0001 |
IEEE J. Biomed. Health Informatics | 10 |
| 2024 | Bridging the Semantic Latent Space between Brain and Machine: Similarity Is All You NeedabstractHow our brain encodes complex concepts has been a longstanding mystery in neuroscience. The answer to this problem can lead to new understandings about how the brain retrieves information in large-scale data with high efficiency and robustness. Neuroscience studies suggest the brain represents concepts in a locality-sensitive hashing (LSH) strategy, i.e., similar concepts will be represented by similar responses. This finding has inspired the design of similarity-based algorithms, especially in contrastive learning. Here, we hypothesize that the brain and large neural network models, both using similarity-based learning rules, could contain a similar semantic embedding space. To verify that, this paper proposes a functional Magnetic Resonance Imaging (fMRI) semantic learning network named BrainSem, aimed at seeking a joint semantic latent space that bridges the brain and a Contrastive Language-Image Pre-training (CLIP) model. Given that our perception is inherently cross-modal, we introduce a fuzzy (one-to-many) matching loss function to encourage the models to extract high-level semantic components from neural signals. Our results claimed that using only a small set of fMRI recordings for semantic space alignment, we could obtain shared embedding valid for unseen categories out of the training set, which provided potential evidence for the semantic representation similarity between the brain and large neural networks. In a zero-shot classification task, our BrainSem achieves an 11.6% improvement over the state-of-the-art. Jiaxuan Chen 0007, Yueming Wang 0001, Gang Pan 0001 |
AAAI | 3 |
| 2024 | Speed-enhanced Subdomain Alignment for Long-term Stable Neural Decoding in Brain-computer InterfacesabstractBrain-computer interfaces (BCIs) offer a means to convert neural signals into control signals, providing a potential restoration of movement for people with paralysis. Despite their promise, BCIs face a significant challenge in maintaining decoding accuracy over time due to neural nonstationarities. While current recalibration techniques address this issue to a degree, they either fail to adequately exploit the limited labeled data, fail to perform conditional alignment in regression tasks, or overlook the signal correlation between data from two days. This paper proposes a novel Speed-enhanced Subdomain Alignment (SeSA) framework, integrating semi-supervised learning with domain adaptation techniques in regressive neural decoding. Specifically, SeSA carries out two alignments (i.e., global alignment and conditional speed alignment) to achieve recalibration. Our comprehensive set of experiments, both qualitative and quantitative, substantiate the superior recalibration performance and robustness of our proposed SeSA. Jiyu Wei, Dazhong Rong, Xinyun Zhu, Qinming He, Yueming Wang 0001 |
BIBM | 5 |
| 2024 | Mind Artist: Creating Artistic Snapshots with Human ThoughtabstractWe introduce Mind Artist (MindArt), a novel and efficient neural decoding architecture to snap artistic photographs from our mind in a controllable manner. Recently, progress has been made in image reconstruction with non-invasive brain recordings, but it's still difficult to generate realistic images with high semantic fidelity due to the scarcity of data annotations. Unlike previous methods, this work casts the neural decoding into optimal transport (OT) and representation decoupling problems. Specifically, under discrete OT theory, we design a graph matching-guided neural representation learning framework to seek the underlying correspondences between conceptual semantics and neural signals, which yields a natural and meaningful self-supervisory task. Moreover, the proposed MindArt, structured with multiple stand-alone modal branches, enables the seamless incorporation of semantic representation into any visual style information, thus leaving it to have multi-modal reconstruction and training-free semantic editing capabilities. By doing so, the reconstructed images of MindArt have phenomenal realism both in terms of semantics and appearance. We compare our MindArt with leading alternatives, and achieve SOTA performance in different decoding tasks. Importantly, our approach can directly generate a series of stylized “mind snapshots” w/o extra optimizations, which may open up more potential applications. Code is available at https://github.com/JxuanC/MindArt. Jiaxuan Chen 0007, Yueming Wang 0001, Gang Pan 0001 |
CVPR | 3 |
| 2024 | A Human-Machine Joint Learning Framework to Boost Endogenous BCI TrainingabstractBrain-computer interfaces (BCIs) provide a direct pathway from the brain to external devices and have demonstrated great potential for assistive and rehabilitation technologies. Endogenous BCIs based on electroencephalogram (EEG) signals, such as motor imagery (MI) BCIs, can provide some level of control. However, mastering spontaneous BCI control requires the users to generate discriminative and stable brain signal patterns by imagery, which is challenging and is usually achieved over a very long training time (weeks/months). Here, we propose a human-machine joint learning framework to boost the learning process in endogenous BCIs, by guiding the user to generate brain signals toward an optimal distribution estimated by the decoder, given the historical brain signals of the user. To this end, we first model the human-machine joint learning process in a uniform formulation. Then a human-machine joint learning framework is proposed: 1) for the human side, we model the learning process in a sequential trial-and-error scenario and propose a novel "copy/new" feedback paradigm to help shape the signal generation of the subject toward the optimal distribution and 2) for the machine side, we propose a novel adaptive learning algorithm to learn an optimal signal distribution along with the subject's learning process. Specifically, the decoder reweighs the brain signals generated by the subject to focus more on "good" samples to cope with the learning process of the subject. Online and psuedo-online BCI experiments with 18 healthy subjects demonstrated the advantages of the proposed joint learning process over coadaptive approaches in both learning efficiency and effectiveness. Lin Yao 0002, Yueming Wang 0001, Dario Farina, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have manifested remarkable advantages in power consumption and event-driven property during the inference process. To take full advantage of low power consumption and improve the efficiency of these models further, the pruning methods have been explored to find sparse SNNs without redundancy connections after training. However, parameter redundancy still hinders the efficiency of SNNs during training. In the human brain, the rewiring process of neural networks is highly dynamic, while synaptic connections maintain relatively sparse during brain development. Inspired by this, here we propose an efficient evolutionary structure learning (ESL) framework for SNNs, named ESL-SNNs, to implement the sparse SNN training from scratch. The pruning and regeneration of synaptic connections in SNNs evolve dynamically during learning, yet keep the structural sparsity at a certain level. As a result, the ESL-SNNs can search for optimal sparse connectivity by exploring all possible parameters across time. Our experiments show that the proposed ESL-SNNs framework is able to learn SNNs with sparse structures effectively while reducing the limited accuracy. The ESL-SNNs achieve merely 0.28% accuracy loss with 10% connection density on the DVS-Cifar10 dataset. Our work presents a brand-new approach for sparse training of SNNs from scratch with biologically plausible evolutionary mechanisms, closing the gap in the expressibility between sparse training and dense training. Hence, it has great potential for SNN lightweight training and inference with low power consumption and small memory usage. Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Yueming Wang 0001, Gang Pan 0001, Huajin Tang |
AAAI | 4 |
| 2023 | HybridSNN: Combining Bio-Machine Strengths by Boosting Adaptive Spiking Neural NetworksabstractSpiking neural networks (SNNs), inspired by the neuronal network in the brain, provide biologically relevant and low-power consuming models for information processing. Existing studies either mimic the learning mechanism of brain neural networks as closely as possible, for example, the temporally local learning rule of spike-timing-dependent plasticity (STDP), or apply the gradient descent rule to optimize a multilayer SNN with fixed structure. However, the learning rule used in the former is local and how the real brain might do the global-scale credit assignment is still not clear, which means that those shallow SNNs are robust but deep SNNs are difficult to be trained globally and could not work so well. For the latter, the nondifferentiable problem caused by the discrete spike trains leads to inaccuracy in gradient computing and difficulties in effective deep SNNs. Hence, a hybrid solution is interesting to combine shallow SNNs with an appropriate machine learning (ML) technique not requiring the gradient computing, which is able to provide both energy-saving and high-performance advantages. In this article, we propose a HybridSNN, a deep and strong SNN composed of multiple simple SNNs, in which data-driven greedy optimization is used to build powerful classifiers, avoiding the derivative problem in gradient descent. During the training process, the output features (spikes) of selected weak classifiers are fed back to the pool for the subsequent weak SNN training and selection. This guarantees HybridSNN not only represents the linear combination of simple SNNs, as what regular AdaBoost algorithm generates, but also contains neuron connection information, thus closely resembling the neural networks of a brain. HybridSNN has the benefits of both low power consumption in weak units and overall data-driven optimizing strength. The network structure in HybridSNN is learned from training samples, which is more flexible and effective compared with existing fixed multilayer SNNs. Moreover, the topological tree of HybridSNN resembles the neural system in the brain, where pyramidal neurons receive thousands of synaptic input signals through their dendrites. Experimental results show that the proposed HybridSNN is highly competitive among the state-of-the-art SNNs. Jiangrong Shen, Jian K. Liu, Yueming Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Tracking Functional Changes in Nonstationary Signals with Evolutionary Ensemble Bayesian Model for Robust Neural DecodingabstractNeural signals are typical nonstationary data where the functional mapping between neural activities and the intentions (such as the velocity of movements) can occasionally change. Existing studies mostly use a fixed neural decoder, thus suffering from an unstable performance given neural functional changes. We propose a novel evolutionary ensemble framework (EvoEnsemble) to dynamically cope with changes in neural signals by evolving the decoder model accordingly. EvoEnsemble integrates evolutionary computation algorithms in a Bayesian framework where the fitness of models can be sequentially computed with their likelihoods according to the incoming data at each time slot, which enables online tracking of time-varying functions. Two strategies of evolve-at-changes and history-model-archive are designed to further improve efficiency and stability. Experiments with simulations and neural signals demonstrate that EvoEnsemble can track the changes in functions effectively thus improving the accuracy and robustness of neural decoding. The improvement is most significant in neural signals with functional changes. Xinyun Zhu, Gang Pan 0001, Yueming Wang 0001 |
NeurIPS | 4 |
| 2022 | Subdomain contraction in deep networks for robust representation learning
Zhentao Pan, Gang Pan 0001, Yueming Wang 0001 |
Neurocomputing | 4 |
| 2022 | Jointly Optimizing Expressional and Residual Models for 3D Facial Expression RemovalabstractThis article proposes a facial expression removal method to recover a 3D neutral face from a single 3D expressional or non-neutral face. We treat a 3D non-neutral face as the sum of its neutral one and the residual. This can be satisfied if the correspondence between 3D vertices of expressional faces and those of neutral faces is established. We propose a non-rigid deformation method to establish the correspondence between 3D faces. Then, according to algebra inequality, the minimization of a neutral face model can be replaced by the minimization of its upper bound, i.e., the errors of an expressional face model and a residual model. Thus, we co-optimize the representation errors of the latter two models and build the relationship between the representation coefficients of the two models. Given an expressional face as the input, its corresponding neutral face can be inferred by the associative representation parameters in these two models. In the testing stage, we use an iterative joint fitting scheme to obtain a more accurate recovery. Extensive experiments are conducted to evaluate our method. The results show that our method obtains considerably better performance than existing methods in terms of average root mean square errors and recognition rates, and also better visual effects. Yueming Wang 0001, Zhenfang Hu, Zhaohui Wu 0001, Gang Pan 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Hierarchical Dynamical Model for Multiple Cortical Neural DecodingabstractMotor brain machine interfaces (BMIs) interpret neural activities from motor-related cortical areas in the brain into movement commands to control a prosthesis. As the subject adapts to control the neural prosthesis, the medial prefrontal cortex (mPFC), upstream of the primary motor cortex (M1), is heavily involved in reward-guided motor learning. Thus, considering mPFC and M1 functionality within a hierarchical structure could potentially improve the effectiveness of BMI decoding while subjects are learning. The commonly used Kalman decoding method with only one simple state model may not be able to represent the multiple brain states that evolve over time as well as along the neural pathway. In addition, the performance of Kalman decoders degenerates in heavy-tailed nongaussian noise, which is usually generated due to the nonlinear neural system or influences of movement-related noise in online neural recording. In this letter, we propose a hierarchical model to represent the brain states from multiple cortical areas that evolve along the neural pathway. We then introduce correntropy theory into the hierarchical structure to address the heavy-tailed noise existing in neural recordings. We test the proposed algorithm on in vivo recordings collected from the mPFC and M1 of two rats when the subjects were learning to perform a lever-pressing task. Compared with the classic Kalman filter, our results demonstrate better movement decoding performance due to the hierarchical structure that integrates the past failed trial information over multisite recording and the combination with correntropy criterion to deal with noisy heavy-tailed neural recordings. Xi Liu 0006, Xiang Zhang 0025, Yifan Huang 0001, Yueming Wang 0001, Yiwen Wang 0002 |
Neural Comput. | 6 |
| 2021 | Dynamic Spatiotemporal Pattern Recognition With Recurrent Spiking Neural NetworkabstractOur real-time actions in everyday life reflect a range of spatiotemporal dynamic brain activity patterns, the consequence of neuronal computation with spikes in the brain. Most existing models with spiking neurons aim at solving static pattern recognition tasks such as image classification. Compared with static features, spatiotemporal patterns are more complex due to their dynamics in both space and time domains. Spatiotemporal pattern recognition based on learning algorithms with spiking neurons therefore remains challenging. We propose an end-to-end recurrent spiking neural network model trained with an algorithm based on spike latency and temporal difference backpropagation. Our model is a cascaded network with three layers of spiking neurons where the input and output layers are the encoder and decoder, respectively. In the hidden layer, the recurrently connected neurons with transmission delays carry out high-dimensional computation to incorporate the spatiotemporal dynamics of the inputs. The test results based on the data sets of spiking activities of the retinal neurons show that the proposed framework can recognize dynamic spatiotemporal patterns much better than using spike counts. Moreover, for 3D trajectories of a human action data set, the proposed framework achieves a test accuracy of 83.6% on average. Rapid recognition is achieved through the learning methodology-based on spike latency and the decoding process using the first spike of the output neurons. Taken together, these results highlight a new model to extract information from activity patterns of neural computation in the brain and provide a novel approach for spike-based neuromorphic computing. Jiangrong Shen, Jian K. Liu, Yueming Wang 0001 |
Neural Comput. | 3 |
| 2020 | Recognizing Scoring in Basketball Game from AER Sequence by Spiking Neural NetworksabstractThe automatic score detection and recognition in basketball game has important application potentials, for examples, basketball technique analysis and 24 second control in the game. Although existing studies have been conducted on broadcast videos, most of them usually learned a machine learning algorithm on long videos recorded by traditional cameras. Address Event Representation (AER) sensor provides a possibility to deal with the problem by a human sensing manner. It represents the visual information as a series of spike-based events and records event sequences. Compared to traditional videos, AER events can fully utilize their addresses and timestamp information, forming precise spatio-temporal features with significantly less storage cost. More importantly, it issues spikes which can be naturally processed by human-style spiking neural networks (SNNs). In this paper, we propose to recognize scoring in basketball game from AER sequences. A new model is designed to extract dynamic features and discriminate different event streams using SNN. To handle the imbalance problem between positive and negative samples, we use an imbalanced Tempotron algorithm in our SNN model. Meanwhile, an AER sequence dataset of basketball games is collected. The experimental results demonstrate that our method achieves better performance compared with existing models. Jiangrong Shen, Jian K. Liu, Yueming Wang 0001 |
IJCNN | 4 |
| 2020 | Binless Kernel Machine: Modeling Spike Train Transformation for Cognitive Neural ProsthesesabstractModeling spike train transformation among brain regions helps in designing a cognitive neural prosthesis that restores lost cognitive functions. Various methods analyze the nonlinear dynamic spike train transformation between two cortical areas with low computational eficiency. The application of a real-time neural prosthesis requires computational eficiency, performance stability, and better interpretation of the neural firing patterns that modulate target spike generation. We propose the binless kernel machine in the point-process framework to describe nonlinear dynamic spike train transformations. Our approach embeds the binless kernel to eficiently capture the feedforward dynamics of spike trains and maps the input spike timings into reproducing kernel Hilbert space (RKHS). An inhomogeneous Bernoulli process is designed to combine with a kernel logistic regression that operates on the binless kernel to generate an output spike train as a point process. Weights of the proposed model are estimated by maximizing the log likelihood of output spike trains in RKHS, which allows a global-optimal solution. To reduce computational complexity, we design a streaming-based clustering algorithm to extract typical and important spike train features. The cluster centers and their weights enable the visualization of the important input spike train patterns that motivate or inhibit output neuron firing. We test the proposed model on both synthetic data and real spike train data recorded from the dorsal premotor cortex and the primary motor cortex of a monkey performing a center-out task. Performances are evaluated by discrete-time rescaling Kolmogorov-Smirnov tests. Our model outperforms the existing methods with higher stability regardless of weight initialization and demonstrates higher eficiency in analyzing neural patterns from spike timing with less historical input (50%). Meanwhile, the typical spike train patterns selected according to weights are validated to encode output spike from the spike train of single-input neuron and the interaction of two input neurons. Cunle Qian, Xuyun Sun, Yueming Wang 0001, Xiaoxiang Zheng, Yiwen Wang 0002, Gang Pan 0001 |
Neural Comput. | 3 |
| 2019 | Dynamic Ensemble Modeling Approach to Nonstationary Neural Decoding in Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-art neural signal decoders such as Kalman filter assume fixed relationship between neural activities and motor movements, thus will fail if this assumption is not satisfied. We propose a dynamic ensemble modeling (DyEnsemble) approach that is capable of adapting to changes in neural signals by employing a proper combination of decoding functions. The DyEnsemble method firstly learns a set of diverse candidate models. Then, it dynamically selects and combines these models online according to Bayesian updating mechanism. Our method can mitigate the effect of noises and cope with different task behaviors by automatic model switching, thus gives more accurate predictions. Experiments with neural data demonstrate that the DyEnsemble method outperforms Kalman filters remarkably, and its advantage is more obvious with noisy signals. Yueming Wang 0001, Gang Pan 0001 |
NeurIPS | 3 |
| 2019 | Activity-dependent neuron model for noise resistance
Yueming Wang 0001, Gang Pan 0001 |
Neurocomputing | 5 |
| 2019 | Deep Attention Network for Egocentric Action RecognitionabstractRecognizing a camera wearer's actions from videos captured by an egocentric camera is a challenging task. In this paper, we employ a two-stream deep neural network composed of an appearance-based stream and a motion-based stream to recognize egocentric actions. Based on the insight that human action and gaze behavior are highly coordinated in object manipulation tasks, we propose a spatial attention network to predict human gaze in the form of attention map. The attention map helps each of the two streams to focus on the most relevant spatial region of the video frames to predict actions. To better model the temporal structure of the videos, a temporal network is proposed. The temporal network incorporates bi-directional long short-term memory to model the long-range dependencies to recognize egocentric actions. The experimental results demonstrate that our method is able to predict attention maps that are consistent with human attention and achieve competitive action recognition performance with the state-of-the-art methods on the GTEA Gaze and GTEA Gaze+ datasets. Minlong Lu, Ze-Nian Li, Yueming Wang 0001, Gang Pan 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Epileptic State Segmentation with Temporal-Constrained ClusteringabstractAutomatic seizure identification plays an important role in epilepsy evaluation. Most existing methods regard seizure identification as a classification problem and rely on labelled training set. However, labelling seizure onset is very expensive and seizure data for each individual is especially limited, classifier-based methods are usually impractical in use. Clustering methods could learn useful information from unlabelled data, while they may lead to unstable results given epileptic signals with high noises. In this paper, we propose to use Gaussian temporal-constrained k-medoids method for seizure state segmentation. Using temporal information, the noises could be effectively suppressed and robust clustering performance is achieved. Besides, a new criterion called signed total variation (STV) which describes temporal integrity and consistency is proposed for temporal-constrained clustering evaluation. Experimental results show that, compared with the existing methods, the k-medoids method with Gaussian temporal constraint achieves the best results on both F1-score and STV. Kang Lin, Shaozhe Feng, Qi Lian, Gang Pan 0001, Yueming Wang 0001 |
ICASSP | 6 |
| 2018 | Jointly Learning Network Connections and Link Weights in Spiking Neural NetworksabstractSpiking neural networks (SNNs) are considered to be biologically plausible and power-efficient on neuromorphic hardware. However, unlike the brain mechanisms, most existing SNN algorithms have fixed network topologies and connection relationships. This paper proposes a method to jointly learn network connections and link weights simultaneously. The connection structures are optimized by the spike-timing-dependent plasticity (STDP) rule with timing information, and the link weights are optimized by a supervised algorithm. The connection structures and the weights are learned alternately until a termination condition is satisfied. Experiments are carried out using four benchmark datasets. Our approach outperforms classical learning methods such as STDP, Tempotron, SpikeProp, and a state-of-the-art supervised algorithm. In addition, the learned structures effectively reduce the number of connections by about 24%, thus facilitate the computational efficiency of the network. Jiangrong Shen, Yueming Wang 0001, Huajin Tang, Hang Yu 0010, Zhaohui Wu 0001, Gang Pan 0001 |
IJCAI | 3 |
| 2016 | Robust object recognition via weakly supervised metric and template learning
Zhenfang Hu, Baoyuan Wang, Jieyi Zhao, Yueming Wang 0001 |
Neurocomputing | 5 |
| 2016 | Robust discriminative non-negative matrix factorization
Ruiqing Zhang, Zhenfang Hu, Gang Pan 0001, Yueming Wang 0001 |
Neurocomputing | 4 |
| 2016 | Sparse Principal Component Analysis via Rotation and TruncationabstractSparse principal component analysis (sparse PCA) aims at finding a sparse basis to improve the interpretability over the dense basis of PCA, while still covering the data subspace as much as possible. In contrast to most existing work that addresses the problem by adding sparsity penalties on various objectives of PCA, we propose a new method, sparse PCA via rotation and truncation (SPCArt), which finds a rotation matrix and a sparse basis such that the sparse basis approximates the basis of PCA after the rotation. The algorithm of SPCArt consists of three alternating steps: 1) rotating the PCA basis; 2) truncating small entries; and 3) updating the rotation matrix. Its performance bounds are also given. The SPCArt is efficient, with each iteration scaling linearly with the data dimension. Parameter choice is simple, due to explicit physical explanations. We give a unified view to several existing sparse PCA methods and discuss the connections with SPCArt. Some ideas from SPCArt are extended to GPower, a popular sparse PCA algorithm, to address its limitations. Experimental results demonstrate that SPCArt achieves the state-of-the-art performance, along with a good tradeoff among various criteria, including sparsity, explained variance, orthogonality, balance of sparsity among loadings, and computational speed. Zhenfang Hu, Gang Pan 0001, Yueming Wang 0001, Zhaohui Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Effective insect recognition using a stacked autoencoder with maximum correntropy criterionabstractThroughout the history, insects had been intimately connected to humanity, in both positive and negative ways. Insects play an important part in crop pollination, on the other hand, some of them spread diseases that kill millions of people every year. Effective control of harmful insects while having little impact to beneficial insects and environment is extremely important. Recently, an intelligent trap that uses laser sensors was presented to control the population of target insects. The device could record and analyze sensor signals when an insect passes through the trap and make quick decisions whether to catch it or not. The effectiveness of the trap relies on the correct choice of classification algorithm to perform the insect detection. In this paper, we propose to use a deep neural network with maximum correntropy criterion (MCC) for reliable classification of insects in real-time. Experimental results show that, deep networks are effective for learning stable features from brief insect passage signals. By replacing the mean square error cost with MCC, the robustness of autoencoders against noise is improved significantly and robust features could be learned. The method is tested on five species of insects and a total of 5325 passages. High classification accuracy of 92.1% is achieved. Compared with previously applied methods, better classification performance is obtained using only 10% of the computation time. Therefore, our method is efficient and reliable for online insect detection. Goktug T. Cinar, Vinícius M. A. de Souza, Gustavo Batista, Yueming Wang 0001, José C. Príncipe |
IJCNN | 5 |
| 2015 | Accelerometer-Based Gait Recognition by Sparse Representation of Signature Points With ClustersabstractGait, as a promising biometric for recognizing human identities, can be nonintrusively captured as a series of acceleration signals using wearable or portable smart devices. It can be used for access control. Most existing methods on accelerometer-based gait recognition require explicit step-cycle detection, suffering from cycle detection failures and intercycle phase misalignment. We propose a novel algorithm that avoids both the above two problems. It makes use of a type of salient points termed signature points (SPs), and has three components: 1) a multiscale SP extraction method, including the localization and SP descriptors; 2) a sparse representation scheme for encoding newly emerged SPs with known ones in terms of their descriptors, where the phase propinquity of the SPs in a cluster is leveraged to ensure the physical meaningfulness of the codes; and 3) a classifier for the sparse-code collections associated with the SPs of a series. Experimental results on our publicly available dataset of 175 subjects showed that our algorithm outperformed existing methods, even if the step cycles were perfectly detected for them. When the accelerometers at five different body locations were used together, it achieved the rank-1 accuracy of 95.8% for identification, and the equal error rate of 2.2% for verification. Yuting Zhang 0001, Gang Pan 0001, Kui Jia, Minlong Lu, Yueming Wang 0001, Zhaohui Wu 0001 |
IEEE Trans. Cybern. | 5 |
| 2014 | Robust feature learning by stacked autoencoder with maximum correntropy criterionabstractUnsupervised feature learning with deep networks has been widely studied in the recent years. Despite the progress, most existing models would be fragile to non-Gaussian noises and outliers due to the criterion of mean square error (MSE). In this paper, we propose a robust stacked autoencoder (R-SAE) based on maximum correntropy criterion (MCC) to deal with the data containing non-Gaussian noises and outliers. By replacing MSE with MCC, the anti-noise ability of stacked autoencoder is improved. The proposed method is evaluated using the MNIST benchmark dataset. Experimental results show that, compared with the ordinary stacked autoencoder, the R-SAE improves classification accuracy by 14% and reduces the reconstruction error by 39%, which demonstrates that R-SAE is capable of learning robust features on noisy data. Yueming Wang 0001, Xiaoxiang Zheng, Zhaohui Wu 0001 |
ICASSP | 2 |
| 2014 | High-fidelity compression of extracellular recordings from motor cortexabstractIn invasive brain-machine interfaces (BMI), the recorded high-quality neural signals produce a large data volume. This calls for effective compression. In this paper, we focus on extracellular recording of motor cortex. First the characteristics of the signals are studied, one of which is that peaks of DCT coefficients at high frequency may correspond to spike firing patterns. Based on these characteristics, we propose a high-fidelity compression framework for these signals. The DCT coefficients of the signal are divided into two parts according to amplitude, rather than frequency. The Low-Amplitude-Component (LAC) is encoded by a phase called Symbol Encoding, which helps to reduce overall distortion. The High-Amplitude-Component (HAC), containing major information and spikes, is encoded by another phase called Hybrid Encoding. It combines the Huffman encoding and a novel Zero-Length-Encoding. Experiments show that the algorithm achieves a compression ratio of 18% without obvious distortion. Moreover, spikes are reserved more than 92%, outperforming existing work. Our algorithm enables low-cost storage devices to store long-time neural signals. Rachel Zhang, Gang Pan 0001, Yueming Wang 0001, Zhenfang Hu |
IJCNN | 3 |
| 2014 | Decoding motor cortical activities of Monkey: A datasetabstractMotor brain-machine interface (BMI) has great potentials in neural motor prostheses and has received increasing attention during the past decades in the neural engineering field. It requires an approach to decode neural activities that represents desired movements. Much of the progress in decoding algorithms has been driven by the availability of neural data, e.g. spike trains, in some research groups having animal laboratories and capable of performing surgery and building BMI systems. However, researchers in the neural signal processing field often face a dilemma of lacking neural data. To continue the innovation in decoding algorithms, this paper introduces a public neural dataset, the ZJU Neural Decoding Dataset (ZJUNDD). We give the detailed paradigm of the BMI system on monkey, including the experimental setup and the collection of 96-channel motor cortical activities. The dataset contains spike rates of neurons obtained by a consistent spike sorting method. To improve the data quality and reduce outliers, the spike data are carefully selected according to the quality of hand movements of the monkey. A standard protocol is provided for the assessment of decoding algorithms on the dataset, including the partition of training and testing sets, and the evaluation metrics. We also build an online evaluation system in order to enable a fair comparison between decoding approaches. Further, we benchmark several existing algorithms, which provides a basic performance of the methods. To the best of our knowledge, this is the first public dataset of spike trains for the decoding research of motor cortical activities. Luoqing Zhou, Yueming Wang 0001, Gang Pan 0001, Yiwen Wang 0002, Xiaoxiang Zheng, Zhaohui Wu 0001 |
IJCNN | 3 |
| 2014 | L1-norm latent SVM for compact features in object detection
Gang Pan 0001, Yueming Wang 0001, Yuting Zhang 0001, Zhaohui Wu 0001 |
Neurocomputing | 3 |
| 2014 | Abidirectional brain-computer interface for effective epilepsy controlabstractBrain-computer interfaces (BCIs) can provide direct bidirectional communication between the brain and a machine. Recently, the BCI technique has been used in seizure control. Usually, a closed-loop system based on BCI is set up which delivers a therapic electrical stimulus only in response to seizure onsets. In this way, the side effects of neurostimulation can be greatly reduced. In this paper, a new BCI-based responsive stimulation system is proposed. With an efficient morphology-based seizure detector, seizure events can be identified in the early stages which trigger electrical stimulations to be sent to the cortex of the brain. The proposed system was tested on rats with penicillin-induced epileptic seizures. Online experiments show that 83% of the seizures could be detected successfully with a short average time delay of 3.11 s. With the therapy of the BCI-based seizure control system, most seizures were suppressed within 10 s. Compared with the control group, the average seizure duration was reduced by 30.7%. Therefore, the proposed system can control epileptic seizures effectively and has potential in clinical applications. Fei-Qiang Ma, Ting-Ting Ge, Yueming Wang 0001, Junming Zhu, Jian-Min Zhang, Xiaoxiang Zheng, Zhaohui Wu 0001 |
J. Zhejiang Univ. Sci. C | 4 |
| 2013 | Generating fluent tubes in video synopsisabstractVideo synopsis is one of the effective techniques to build a short video representation preserving the essential activities for a long video. Existing methods usually have the problem that a continuous activity (tube) from a single moving object is separated to a few small pieces. In this paper, two schemes are proposed to generate fluent tubes for video synopsis. The Gaussian mixture model and a texture method are combined to detect more compact foreground with shadow removed. The foreground constitutes a set of initial trajectories. A particle filter tracker is used to concatenate two trajectories if they belong to the same foreground activity, which generates more fluent tubes for video synopsis. Experimental results on 4 videos show that our method produces better accuracies and visual effects in video synopsis. Minlong Lu, Yueming Wang 0001, Gang Pan 0001 |
ICASSP | 2 |
| 2013 | Efficient computation of histograms on densely overlapped polygonal regions
Yuting Zhang 0001, Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001 |
Neurocomputing | 2 |
| 2013 | Combining velocity and Location-Specific Spatial Clues in Trajectories for Counting Crowded Moving ObjectsabstractTrajectory-clustering-based methods have shown a good performance in counting moving objects in densely crowded scenes. However, they still fall into trouble in complex scenes, such as with the close proximity of moving objects, freely moving parts of objects, and different object size in different locations of the scene. This paper proposes a new method combining velocity and location-specific spatial clues in trajectories to deal with these problems. We first extract the velocities of a trajectory over its life-time. To alleviate confusion around the boundary regions between close objects, extracted velocity information is utilized to eliminate unreal-world feature points on objects' boundaries. Then, a function is introduced to measure the similarity of the trajectories integrating both of the spatial and the velocity clues. This function is employed in the Mean-Shift clustering procedure to reduce the effect of freely moving parts of the objects. To address the problem of various object sizes in different regions of the scene, we suggest a technique to learn the location-specific size distribution of objects in different locations of a scene. The experimental results show that our proposed method achieves a good performance. Compared with other trajectory-clustering-based methods, it decreases the counting error rate by about 10%. Mahdi Hashemzadeh, Gang Pan 0001, Yueming Wang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | Establishing Point Correspondence of 3D Faces Via Sparse Facial Deformable ModelabstractEstablishing a dense vertex-to-vertex anthropometric correspondence between 3D faces is an important and fundamental problem in 3D face research, which can contribute to most applications of 3D faces. This paper proposes a sparse facial deformable model to automatically achieve this task. For an input 3D face, the basic idea is to generate a new 3D face that has the same mesh topology as a reference face and the highly similar shape to the input face, and whose vertices correspond to those of the reference face in an anthropometric sense. Two constraints: 1) the shape constraint and 2) correspondence constraint are modeled in our method to satisfy the three requirements. The shape constraint is solved by a novel face deformation approach in which a normal-ray scheme is integrated to the closest-vertex scheme to keep high-curvature shapes in deformation. The correspondence constraint is based on an assumption that if the vertices on 3D faces are corresponded, their shape signals lie on a manifold and each face signal can be represented sparsely by a few typical items in a dictionary. The dictionary can be well learnt and contains the distribution information of the corresponded vertices. The correspondence information can be conveyed to the sparse representation of the generated 3D face. Thus, a patch-based sparse representation is proposed as the correspondence constraint. By solving the correspondence constraint iteratively, the vertices of the generated face can be adjusted to correspondence positions gradually. At the early iteration steps, smaller sparsity thresholds are set that yield larger representation errors but better globally corresponded vertices. At the later steps, relatively larger sparsity thresholds are used to encode local shapes. By this method, the vertices in the new face approach the right positions progressively until the final global correspondence is reached. Our method is automatic, and the manual work is needed only in training procedure. The experimental results on a large-scale publicly available 3D face data set, BU-3DFE, demonstrate that our method achieves better performance than existing methods. Gang Pan 0001, Yueming Wang 0001, Zhenfang Hu, Xiaoxiang Zheng, Zhaohui Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | FlyingBuddy2: a brain-controlled assistant for the handicappedabstractThe motor impaired people have much limit in moving. The devices augmenting their mobility will be much helpful for improving their living experiences. This poster develops a brain-controlled assistive system, called FlyingBuddy2, to aid the handicapped in mobility. It uses the brain EEG signals to directly control a quadrotor. Signals from an EEG headset are transmitted wirelessly to a computer, then the decoded brain signals are converted to trigger the quadrotor to move in 3D space. Three applications are developed: thinking to play games, thinking to see, and thinking to take pictures. Yipeng Yu, Weidong Hua, Shijian Li, Yueming Wang 0001, Gang Pan 0001 |
UbiComp | 6 |
| 2011 | A deformation model to reduce the effect of expressions in 3D face recognition
Yueming Wang 0001, Gang Pan 0001, Jianzhuang Liu |
Vis. Comput. | 1 |
| 2010 | Robust 3D Face Recognition by Local Shape Difference BoostingabstractThis paper proposes a new 3D face recognition approach, Collective Shape Difference Classifier (CSDC), to meet practical application requirements, i.e., high recognition performance, high computational efficiency, and easy implementation. We first present a fast posture alignment method which is self-dependent and avoids the registration between an input face against every face in the gallery. Then, a Signed Shape Difference Map (SSDM) is computed between two aligned 3D faces as a mediate representation for the shape comparison. Based on the SSDMs, three kinds of features are used to encode both the local similarity and the change characteristics between facial shapes. The most discriminative local features are selected optimally by boosting and trained as weak classifiers for assembling three collective strong classifiers, namely, CSDCs with respect to the three kinds of features. Different schemes are designed for verification and identification to pursue high performance in both recognition and computation. The experiments, carried out on FRGC v2 with the standard protocol, yield three verification rates all better than 97.9 percent with the FAR of 0.1 percent and rank-1 recognition rates above 98 percent. Each recognition against a gallery with 1,000 faces only takes about 3.6 seconds. These experimental results demonstrate that our algorithm is not only effective but also time efficient. Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Image Segmentation by MAP-ML EstimationsabstractImage segmentation plays an important role in computer vision and image analysis. In this paper, image segmentation is formulated as a labeling problem under a probability maximization framework. To estimate the label configuration, an iterative optimization scheme is proposed to alternately carry out the maximum a posteriori (MAP) estimation and the maximum likelihood (ML) estimation. The MAP estimation problem is modeled with Markov random fields (MRFs) and a graph cut algorithm is used to find the solution to the MAP estimation. The ML estimation is achieved by computing the means of region features in a Gaussian model. Our algorithm can automatically segment an image into regions with relevant textures or colors without the need to know the number of regions in advance. Its results match image edges very well and are consistent with human perception. Comparing to six state-of-the-art algorithms, extensive experiments have shown that our algorithm performs the best. Shifeng Chen, Liangliang Cao, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
IEEE Trans. Image Process. | 3 |
| 2009 | Automatic facial expression recognition on a single 3D face by exploring shape deformationabstractFacial expression recognition has many applications in multimedia processing and the development of 3D data acquisition techniques makes it possible to identify expressions using 3D shape information. In this paper, we propose an automatic facial expression recognition approach based on a single 3D face. The shape of an expressional 3D face is approximated as the sum of two parts, a basic facial shape component (BFSC) and an expressional shape component (ESC). The BFSC represents the basic face structure and neutral-style shape and the ESC contains shape changes caused by facial expressions. To separate the BFSC and ESC, our method firstly builds a reference face for each input 3D non-neutral face by a learning method, which well represents the basic facial shape. Then, based on the BFSC and the original expressional face, a facial expression descriptor is designed. The surface depth changes are considered in the descriptor. Finally, the descriptor is input into an SVM to recognize the expression. Unlike previous methods which recognize a facial expression with the help of manually labeled key points and/or a neutral face, our method works on a single 3D face without any manual assistance. Extensive experiments are carried out on the BU-3DFE database and comparisons with existing methods are conducted. The experimental results show the effectiveness of our method. Boqing Gong, Yueming Wang 0001, Jianzhuang Liu, Xiaoou Tang |
ACM Multimedia | 2 |
| 2008 | 3D Face Recognition by Local Shape Difference Boosting
Yueming Wang 0001, Xiaoou Tang, Jianzhuang Liu, Gang Pan 0001, Rong Xiao 0003 |
ECCV (1) | 1 |
| 2007 | 3D Face Recognition in the Presence of Expression: A Guidance-based Constraint Deformation ApproachabstractThree-dimensional human face recognition in the presence of expression is a big challenge, since the shape distortion caused by facial expression greatly weakens the rigid matching. This paper proposes a guidance-based constraint deformation(GCD) model to cope with the shape distortion by expression. The basic idea is that, the face model with non-neutral expression is deformed toward its neutral one under certain constraint so that the distortion is reduced while inter-class discriminative information is preserved. The GCD model exploits the neutral 3D face shape to guide the deformation, meanwhile applies a rigid constraint on it. Both steps are smoothly unified in the Poisson equation framework. The GCD approach only needs one neutral model for each person in the gallery. The experimental results, carried out on the large 3D face databases-FRGC v2.0, demonstrate that our method significantly outperforms ICP method for both identification and authentication mode. It shows the GCD model is promising for coping with the shape distortion in 3D face recognition. Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001 |
CVPR | 1 |
| 2006 | Hallucinating 3D Faces
Shiqi Peng, Gang Pan 0001, Shi Han, Yueming Wang 0001 |
ACCV (2) | 4 |
| 2006 | Exploring Facial Expression Effects in 3D Face Recognition Using Partial ICP
Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001, Yigang Wang |
ACCV (1) | 1 |
| 2006 | Super-Resolution of 3D Face
Gang Pan 0001, Shi Han, Zhaohui Wu 0001, Yueming Wang 0001 |
ECCV (2) | 4 |
| 2004 | 3d face recognition using local shape map
Zhaohui Wu 0001, Yueming Wang 0001, Gang Pan 0001 |
ICIP | 2 |
| 2003 | Pose-invariant detection of facial features from range dataabstractThis paper, firstly, presents a new signature representation for point in range data, called curgram. Curgram serves to describe the structural neighborhood of a point and establish a signature for each of the given 3D data points rather than information of this point position. It yields invariance under rigid motions and mirror imaging. Secondly, its application to detection of facial features from range data is performed by incorporating the metric for curgram into the conventional appearance-based face detection framework. Experimental results with fifteen face range images have demonstrated the validity and effectiveness of the proposed method. Gang Pan 0001, Yueming Wang 0001, Zhaohui Wu 0001 |
SMC | 2 |