EDBT 2026 Demo / reviewers in the wild / expert
Gang Pan 0001
dblp:86/4183-1
· DBLP profile ↗
238ranked-venue papers
14as first author
136since 2021 · last 2026
0000-0002-4049-6181ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 139 · 5 first-author · 97 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 8 first-author · 45 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 3 first-author · 12 since 2021Human-computer interaction and ubiquitous computing · 22 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 15 · 6 since 2021Systems, architecture and hardware · 13 · 10 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EMOD: A Unified EEG Emotion Representation Framework Leveraging V-A Guided Contrastive LearningabstractEmotion recognition from EEG signals is essential for affective computing and has been widely explored using deep learning. While recent deep learning approaches have achieved strong performance on single EEG emotion datasets, their generalization across datasets remains limited due to the heterogeneity in annotation schemes and data formats. Existing models typically require dataset-specific architectures tailored to input structure and lack semantic alignment across diverse emotion labels. To address these challenges, we propose EMOD: A Unified EEG Emotion Representation Framework Leveraging Valence–Arousal (V–A) Guided Contrastive Learning. EMOD learns transferable and emotion-aware representations from heterogeneous datasets by bridging both semantic and structural gaps. Specifically, we project discrete and continuous emotion labels into a unified V–A space and formulate a soft-weighted supervised contrastive loss that encourages emotionally similar samples to cluster in the latent space. To accommodate variable EEG formats, EMOD employs a flexible backbone comprising a Triple-Domain Encoder followed by a Spatial-Temporal Transformer, enabling robust extraction and integration of temporal, spectral, and spatial features. We pretrain EMOD on 8 public EEG datasets and evaluate its performance on three benchmark datasets. Experimental results show that EMOD achieves the state-of-the-art performance, demonstrating strong adaptability and generalization across diverse EEG-based emotion recognition scenarios. Yuning Chen, Sha Zhao, Shijian Li, Gang Pan 0001 |
AAAI | 4 |
| 2026 | S³: Spiking Neurons as an Isolating Segmenter for Brain Signal DecodingabstractRecent brain decoding studies have primarily emphasized the development of brain decoders, while largely neglecting the segmentation step. Existing methods typically adopt fixed-length segmentation, which might overlook subject- or task-level variability and disrupt temporal patterns within brain signals. To address this gap, we propose S3, which leverages spiking neurons as an isolating segmenter for brain signal decoding. S3 segments brain signals adaptively, considering subject- and task-level variability while preserving intrinsic temporal patterns of brain signals. It exploits the unique reset mechanism of spiking neurons to isolate previous irrelevant temporal patterns during the generation of each segmentation point. To optimize S3 for enhancing task performance in the absence of segmentation labels, we develop an optimization method where segmentation pseudo-labels are created with a stochastic-greedy algorithm to optimize them, while circumventing gradient blockade between S3 and task performance. Experiments on 10 downstream tasks across 13 public datasets demonstrate that S3 consistently outperforms existing methods, validating its effectiveness, generalizability and interpretability. Sha Zhao, Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 8 |
| 2026 | SPEAK: Spiking Neurons as an Entropy-Aware Tokenizer for Large Language ModelsabstractMing Chen, Wenyao Li, Chao Liang, Shi Gu, Peng Lin, De Ma, Huajin Tang, Qian Zheng, Gang Pan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shi Gu, De Ma, Huajin Tang, Gang Pan 0001 |
ACL (1) | 9 |
| 2026 | MASI: Memory-Adaptive Inference Framework for Spiking Neural Networks on Edge DevicesabstractThe rapid development of the Internet of Things (IoT) applications necessitates resource-efficient computing paradigms that can unify heterogeneous sensing modalities. Spiking Neural Networks (SNNs) meet this need with their event-driven and energy-efficient processing nature. However, deploying SNNs on mobile and embedded platforms is hindered by strict and fluctuating memory budgets. While prior work explores lightweight model design and system-level memory management, these methods either sacrifice accuracy or incur high runtime overhead due to timestep-dependent dynamics. To tackle these challenges, we propose a memory-adaptive framework MASI that enables efficient on-device SNN inference by combining (1) a fine-grained memory-adaptive layer slicing strategy, (2) a timestep-agnostic scheduler that maximizes memory utilization with minimal fragmentation, and (3) a timestep-aware early-exit mechanism that reduces redundant calculations. Evaluated on diverse workloads and edge devices, MASI can dynamically adapt to runtime memory availability, approximately reducing memory usage by 20.67% and inference latency by 58.53% on average with negligible accuracy loss compared to other feasible on-device implementations under memory constraints. Di Yu 0001, Helin Zheng, Changze Lv, Xin Du 0002, Linshan Jiang, Xiang Liu 0017, Gang Pan 0001, Shuiguang Deng |
WWW | 7 |
| 2026 | Neuromorphic neural decoding model towards high-performance neuroprosthetics
Gang Pan 0001 |
Neurocomputing | 4 |
| 2026 | EEGDiffuser: Label-guided EEG signals synthesis via diffusion model for BCI applications
Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Shijian Li, Gang Pan 0001 |
Neurocomputing | 6 |
| 2026 | Resource-efficient fully automatic spike sorter via scale-adaptive neuromorphic model
Hang Yu 0010, Jiayu Yu, Gang Pan 0001 |
Neurocomputing | 4 |
| 2026 | Cross-subject EEG-based emotion recognition leveraging multi-source domain adaptation with curriculum leaning strategy
Sha Zhao, Yitian Liu, Shijian Li, Gang Pan 0001 |
Neurocomputing | 4 |
| 2026 | CausalCOMRL: Context-based offline meta-reinforcement learning with causal representationabstractContext-based offline meta-reinforcement learning (OMRL) methods have achieved appealing success by leveragingpre-collected offline datasets to develop task representations that guide policy learning. However, current context-based OMRL methods often introduce spurious correlations, where task components are incorrectly correlated due to confounders. These correlations can degrade policy performance when the confounders in the test taskdiffer from those in the training task. To address this problem, we propose CausalCOMRL, a context-based OMRL method that integrates causal representation learning. This approach uncovers causal relationships among the task components and incorporates the causal relationships into task representations, enhancing the generalizability of RL agents. We further improve the distinction of task representations from different tasks by using mutual information optimization and contrastive learning. Utilizing these causal task representations, we employSAC to optimize policies on meta-RL benchmarks. Experimental results show that CausalCOMRL achieves better performance than other methods on most benchmarks. Zhengzhe Zhang, Wenjia Meng, Haoliang Sun, Gang Pan 0001 |
Neural Networks | 4 |
| 2026 | DepAsync: An Asynchronous SNN Accelerator Based on Core-DependencyabstractSpiking Neural Networks (SNNs) are widely used in brain-inspired computing and neuroscience research. Several many-core accelerators have been built to improve the running speed and energy efficiency of SNNs. However, current accelerators generally need explicit synchronization among all cores after each timestep of SNNs, which poses a challenge to overall efficiency. This paper proposes DepAsync, an asynchronous architecture that eliminates inter-core synchronization, facilitating fast and energy-efficient SNN inference with commendable scalability. The main idea is to exploit the dependency of neuromorphic cores predetermined at compile time. We design a DepAsync scheduler for each core to trace the running state of its dependencies and control the core to safely forward to the next timestep without waiting for other cores to complete their tasks. This approach prevents the necessity for global synchronization, allowing DepAsync to minimize core waiting time facing inherent core and time imbalance in SNN workloads. The comprehensive evaluations using five SNN workloads show that DepAsync achieves 2.47x speedup and 1.55x energy efficiency compared to the state-of-the-art synchronization architectures. Zhuo Chen 0044, De Ma, Xiaofei Jin, Qinghui Xing, Ouwen Jin, Xin Du 0002, Shuibing He, Gang Pan 0001 |
IEEE Trans. Computers | 8 |
| 2026 | NPEva: An NoC-Based Neuromorphic Processors Performance Evaluation Framework for Benchmarking Spiking Neural NetworksabstractSpiking Neural Network (SNN) applications place diverse demands on neuromorphic processors’ computational and communication capabilities. Specifically, the sparse parallel computation of SNNs requires storage systems capable of extensive parallel data access and processing. Additionally, the time-dependent nature of SNN computations demands a communication framework to manage spike transmission and synchronization. Addressing these requires continuous early-stage evaluation of potential designs, which can be time-intensive. To this end, this paper introduces a performance evaluation model for NoC-based neuromorphic processors, named NPEva, which includes a Computation Model (CpMo) and a Communication Model (CoMo). This model facilitates rapid, high-dimensional exploration of design spaces across a wide range of microarchitectural parameters. CpMo quickly assesses computation-related latency, power consumption, and area by extracting relevant parameters from hierarchical storage organizations and neuron computation models. CoMo employs communication upper-bound latency analysis to evaluate packet latencies to reduce redundant synchronization time. Comparisons with TrueNorth’s publicly released data reveal that the model’s evaluation results are within an 8% error margin. The CoMo offers enhanced precision in estimating latency upper-bounds compared to other models like worst-contention delay (WCD) and worst-case traversal time (WCTT), especially when varying the number of VCs. With VC settings of 2, 4, 8, and 16, CoMo reduced latency by 0.55× to 2.53× compared to WCTT and by 0.83× to 25.57× compared to WCD. Additionally, using the NPEva model for exploring the design space of Liquid State Machines (LSM) networks has shown that optimized hardware can improve cost efficiency by 2.94× with only a 1% increase in latency. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event FeaturesabstractEmotion recognition is essential for improving user experience and interaction quality in human-centered applications. While recent studies have leveraged both event and traditional cameras to enhance eye-based emotion recognition, their practical deployment is hindered by the scarcity of event cameras and the complexity of dual-modality frameworks. Personalization, which is critical for handling individual differences in emotional expression, is also affected by these factors, resulting in reduced performance and adaptation efficiency. To address these challenges, we propose a lightweight and personalized single-eye emotion recognition network, called LPSEER. LPSEER introduces a novel hybrid neural architecture that integrates a convolutional neural network (CNN) and a spiking neural network (SNN) to capture spatiotemporal features from video frames and events, respectively. Additionally, we design a memorybased event feature inference (MEFI) module that recalls event features from video frames, eliminating the reliance on event cameras during inference and personalization while retaining the discriminative advantages of event-based representations. Experimental results demonstrate that LPSEER achieves state-of- the-art recognition accuracy while maintaining the smallest model size and lowest computational cost. Further experiments confirm the strong generalization capabilities and the ability to achieve faster, more accurate personalization. These advantages collectively enable lightweight, accurate, and efficient emotion recognition for real-world human-centered applications. Qianhui Liu, Jiqing Zhang, Yang Wang 0106, Malu Zhang, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Cost-Efficient Open Vocabulary 3D Scene Understanding Based on Semantic ProbabilityabstractTraditional 3D scene understanding methods heavily depend on 3D annotation and training, which allow for the identification of seen classes but struggle to recognize unseen classes. In this paper, we leverage the open vocabulary inference capabilities of pre-trained models, enabling the encoding of open vocabulary concepts. However, unlike existing open vocabulary 3D scene understanding methods, we propose a framework based on semantic probability. This innovation significantly reduces computational cost and is compatible with state-of-the-art two-stage 2D pre-trained models. Specifically, we align the text features from the CLIP model with the pixel features from the 2D pre-trained models, inferring semantic probability of image pixels based on similarity and projecting it onto 3D points. Subsequently, we introduce a point cloud pairs semantic fusion method to merge the point clouds, reducing the semantic probability of erroneous 3D points. Based on probability scores, we achieve 3D semantic segmentation on open vocabularies without any supervision or training. In addition, the semantic probability of 3D points can serve as pseudo-labels for 3D distillation, and the geometric features of the 3D scene can be exploited to improve the segmentation performance. Experimental results demonstrate that the proposed method exhibits competitive performance on publicly available benchmark datasets, including ScanNet, Matterport3D, and nuScenes. Lingfeng Shen, Xiaoyao Wei, Gang Pan 0001, Yanlong Cao |
IEEE Trans. Image Process. | 3 |
| 2026 | Generic-to-Personalised Learning for Multimodal Image Synthesis With Bidirectional Variational GANabstractMultimodal image synthesis, which predicts target-modality images from source-modality images, has garnered considerable attention in the field of clinical diagnosis. Both unidirectional and bidirectional multimodal image synthesis methods have been explored in the medical domain, however, unidirectional models heavily rely on paired images, while current bidirectional models typically overlook local image details due to their unsupervised training patterns. In this work, we propose a Bidirectional Variational Generative Adversarial Network (BVGAN) for multimodal image synthesis, which achieves high-quality bidirectional translations between any two modalities using only a limited number paired images. Firstly, BVGAN's generator incorporates a variational structure (VAS) to regularise the latent space for noise reduction. This regularisation imposes smoothness to the latent space, enabling BVGAN to produce high-quality, noise-free images. Secondly, a novel generic-to-personalised (GTP) learning strategy is introduced to train BVGAN and reduce its reliance on a large sets of paired images. GTP initially leverages an unsupervised learning model to capture the global mapping between two modalities using unpaired images from generic patients. It then applies a supervised learning model to refine the mapping for individual patient, enhancing image details. Finally, the GTP learning strategy along with VAS enables BVGAN to achieve state-of-the-art performance on two multi-modality medical datasets: Brain CTMRI and BRATS. Long Chen 0019, Xirui Dong, Jiangrong Shen, Lu Zhang 0053, Qi Xu 0008, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Multim. | 6 |
| 2025 | EvHDR-GS: Event-guided HDR Video Reconstruction with 3D Gaussian SplattingabstractHigh Dynamic Range (HDR) video reconstruction seeks to accurately restore the extensive dynamic range present in real-world scenes and is widely employed in downstream applications. Existing methods typically operate on one or a small number of consecutive frames, which often leads to inconsistent brightness across the video due to their limited perspective on the video sequence. Moreover, supervised learning-based approaches are susceptible to data bias, resulting in reduced effectiveness when confronted with test inputs exhibiting a domain gap relative to the training data. To address these limitations, we present an event-guided HDR video reconstruction method through building 3D Gaussian Splatting (3DGS), to ensure consistent brightness imposed by 3D consistency. We introduce HDR 3D Gaussians capable of simultaneously representing HDR and low-dynamic-range (LDR) colors. Furthermore, we incorporate a learnable HDR-to-LDR transformation optimized by input event streams and LDR frames to eliminate the data bias. Experimental results on both synthetic and real-world datasets demonstrate that the proposed method achieves state-of-the-art performance. Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
AAAI | 7 |
| 2025 | EvHDR-NeRF: Building High Dynamic Range Radiance Fields with Single Exposure Images and EventsabstractWe present EvHDR-NeRF to recover a High Dynamic Range (HDR) radiance field from event streams and a set of Low Dynamic Range (LDR) views with single exposures. Using the EvHDR-NeRF, we can generate both novel HDR views and novel LDR views under different exposures. The key to our method is to model the new relationship between events streams and LDR images, which considers both the Camera Response Function (CRF) and exposure time. Based on this relationship, we categorize events into inter-frame events and intra-exposure. The former is utilized for building HDR radiance field and the latter is used to deblur potentially blurred images. Compared to existing methods, this method can effectively reconstruct the HDR radiance field even when the input images are degraded. Experimental results demonstrate that our method achieves state-of-the-art HDR reconstruction, providing a more adaptable and accurate solution for complex imaging applications. Zhanfeng Liao, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 6 |
| 2025 | EvSTVSR: Event Guided Space-Time Video Super-ResolutionabstractIn the domain of space-time video super-resolution, it is typically challenging to handle complex motions (including large and nonlinear motions) and varying illumination scenes due to the lack of inter-frame information. Leveraging the dense temporal information provided by event signals offers a promising solution. Traditional event-based methods typically rely on multiple images, using motion estimation and compensation, which can introduce errors. Accumulated errors from multiple frames often lead to artifacts and blurriness in the output. To mitigate these issues, we propose EvSTVSR, a method that uses fewer adjacent frames and integrates dense temporal information from events to guide alignment. Additionally, we introduce a coordinate-based feature fusion upsampling module to achieve spatial super-resolution. Experimental results demonstrate that our method not only outperforms existing RGB-based approaches but also excels in handling large motion scenarios. Haojie Yan, Zhan Lu, De Ma, Huajin Tang, Gang Pan 0001 |
AAAI | 7 |
| 2025 | Personalized Sleep Staging Leveraging Source-free Unsupervised Domain AdaptationabstractSleep staging is important for monitoring sleep quality and diagnosing sleep-related disorders. Recently, numerous deep learning-based models have been proposed for automatic sleep staging using polysomnography recordings. Most of them are trained and tested on the same labeled datasets which results in poor generalization to unseen target domains. However, they regard the subjects in the target domains as a whole and overlook the individual discrepancies, which limits the model's generalization ability to new patients (i.e., unseen subjects) and plug-and-play applicability in clinics. To address this, we propose a novel Source-Free Unsupervised Individual Domain Adaptation (SF-UIDA) framework for sleep staging, leveraging sequential cross-view contrasting and pseudo-label based fine-tuning. It is actually a two-step subject-specific adaptation scheme, which enables the source model to effectively adapt to newly appeared unlabeled individual without access to the source data. It meets the practical needs in real-world scenarios, where the personalized customization can be plug-and-play applied to new ones. Our framework is applied to three classic sleep staging models and evaluated on three public sleep datasets, achieving the state-of-the-art performance. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Benyan Luo, Gang Pan 0001 |
AAAI | 8 |
| 2025 | Improving Cross-Task Applicability of Parameter Sharing in Cooperative Multi-Agent Reinforcement LearningabstractParameter sharing is a widely adopted approach in cooperative Multi-Agent Reinforcement Learning (MARL), often achieving strong performance. However, its effectiveness can vary, as the policy’s similarity induced by parameter sharing may hinder performance in certain tasks. In this study, we propose a novel framework, termed Composite Shared Policy (CSP), to enhance the cross-task applicability of parameter sharing. CSP is designed to model multiple diverse policies concurrently, thereby introducing inherent policy diversity without relying on task-specific designs. By increasing the differences among the policies of individual agents, CSP effectively mitigates the policy similarity problem commonly associated with parameter sharing. These characteristics collectively enable CSP to improve the cross-task applicability of parameter sharing. To empirically validate the effectiveness of CSP, we implement it based on QMIX, a classic cooperative MARL method, and conduct experiments across two widely used MARL testbeds. The experimental results demonstrate that CSP significantly enhances the cross-task applicability of parameter sharing. Additionally, we conduct ablation studies to evaluate the contributions of each component within CSP. The results highlight that each component plays a critical role in the overall effectiveness of the framework. The source code is available at https://github.com/Yurui-Li/CSP. Jianyu Zhang 0001, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ECAI | 5 |
| 2025 | VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision MakingabstractRecent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm.We show that this paradigm fundamentally limits performance in multi-input multi-output (MIMO) scenarios, where parallel task execution is required.In MISO architectures, tasks compete for a shared output channel, creating mutual exclusion effects that cause unbalanced optimization and degraded performance.To address this gap, we introduce MIMO-VLA (VLASCD), a unified training framework that enables concurrent multi-task outputs, exemplified by simultaneous dialogue generation and decision-making.Inspired by human cognition, MIMO-VLA eliminates interference between tasks and supports efficient parallel processing.Experiments on the CARLA autonomous driving platform demonstrate that MIMO-VLA substantially outperforms stateof-the-art MISO-based LLMs, reinforcement learning models, and VLAs in MIMO settings, establishing a new direction for multimodal and multitask learning.Our code is available at: Zuojin Tang, De Ma, Gang Pan 0001 |
EMNLP | 5 |
| 2025 | Improving Stability of Parameter Sharing in Cooperative Multi-agent Reinforcement Learning
Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ICANN (1) | 4 |
| 2025 | Bridging the Gap Between Brain and Machine in Interpreting Visual Semantics: Towards Self-Adaptive Brain-to-Text Decoding
Yueming Wang 0001, Gang Pan 0001 |
ICCV | 4 |
| 2025 | E-NeMF: Event-based Neural Motion Field for Novel Space-time View Synthesis of Dynamic Scenes
Haojie Yan, De Ma, Huajin Tang, Gang Pan 0001 |
ICCV | 7 |
| 2025 | Unsupervised Rgb-D Point Cloud Registration for Scenes With Low Overlap and Photometric Inconsistency
Yejun Shou, Lingfeng Shen, Gang Pan 0001, Yanlong Cao |
ICCV | 5 |
| 2025 | Mitigating Reward Over-Optimization in RLHF via Behavior-Supported RegularizationabstractReinforcement learning from human feedback (RLHF) is an effective method for aligning large language models (LLMs) with human values. However, reward over-optimization remains an open challenge leading to discrepancies between the performance of LLMs under the reward model and the true human objectives. A primary contributor to reward over-optimization is the extrapolation error that arises when the reward model evaluates out-of-distribution (OOD) responses. However, current methods still fail to prevent the increasing frequency of OOD response generation during the reinforcement learning (RL) process and are not effective at handling extrapolation errors from OOD responses. In this work, we propose the *Behavior-Supported Policy Optimization* (BSPO) method to mitigate the reward over-optimization issue. Specifically, we define *behavior policy* as the next token distribution of the reward training dataset to model the in-distribution (ID) region of the reward model. Building on this, we introduce the behavior-supported Bellman operator to regularize the value function, penalizing all OOD values without impacting the ID ones. Consequently, BSPO reduces the generation of OOD responses during the RL process, thereby avoiding overestimation caused by the reward model’s extrapolation errors. Theoretically, we prove that BSPO guarantees a monotonic improvement of the supported policy until convergence to the optimal behavior-supported policy. Empirical results from extensive experiments show that BSPO outperforms baselines in preventing reward over-optimization due to OOD evaluation and finding the optimal ID policy. Juntao Dai, Taiye Chen, Yaodong Yang 0001, Gang Pan 0001 |
ICLR | 5 |
| 2025 | Improving the Sparse Structure Learning of Spiking Neural Networks from the View of Compression EfficiencyabstractThe human brain utilizes spikes for information transmission and dynamically reorganizes its network structure to boost energy efficiency and cognitive capabilities throughout its lifespan. Drawing inspiration from this spike-based computation, Spiking Neural Networks (SNNs) have been developed to construct event-driven models that emulate this efficiency. Despite these advances, deep SNNs continue to suffer from over-parameterization during training and inference, a stark contrast to the brain’s ability to self-organize. Furthermore, existing sparse SNNs are challenged by maintaining optimal pruning levels due to a static pruning ratio, resulting in either under or over-pruning.
In this paper, we propose a novel two-stage dynamic structure learning approach for deep SNNs, aimed at maintaining effective sparse training from scratch while optimizing compression efficiency.
The first stage evaluates the compressibility of existing sparse subnetworks within SNNs using the PQ index, which facilitates an adaptive determination of the rewiring ratio for synaptic connections based on data compression insights. In the second stage, this rewiring ratio critically informs the dynamic synaptic connection rewiring process, including both pruning and regrowth. This approach significantly improves the exploration of sparse structures training in deep SNNs, adapting sparsity dynamically from the point view of compression efficiency.
Our experiments demonstrate that this sparse training approach not only aligns with the performance of current deep SNNs models but also significantly improves the efficiency of compressing sparse SNNs. Crucially, it preserves the advantages of initiating training with sparse models and offers a promising solution for implementing Edge AI on neuromorphic hardware. Jiangrong Shen, Qi Xu 0008, Gang Pan 0001, Badong Chen |
ICLR | 3 |
| 2025 | CBraMod: A Criss-Cross Brain Foundation Model for EEG DecodingabstractElectroencephalography (EEG) is a non-invasive technique to measure and record brain electrical activity, widely used in various BCI and healthcare applications. Early EEG decoding methods rely on supervised learning, limited by specific tasks and datasets, hindering model performance and generalizability. With the success of large language models, there is a growing body of studies focusing on EEG foundation models. However, these studies still leave challenges: Firstly, most of existing EEG foundation models employ full EEG modeling strategy. It models the spatial and temporal dependencies between all EEG patches together, but ignores that the spatial and temporal dependencies are heterogeneous due to the unique structural characteristics of EEG signals. Secondly, existing EEG foundation models have limited generalizability on a wide range of downstream BCI tasks due to varying formats of EEG data, making it challenging to adapt to. To address these challenges, we propose a novel foundation model called CBraMod. Specifically, we devise a criss-cross transformer as the backbone to thoroughly leverage the structural characteristics of EEG signals, which can model spatial and temporal dependencies separately through two parallel attention mechanisms. And we utilize an asymmetric conditional positional encoding scheme which can encode positional information of EEG patches and be easily adapted to the EEG with diverse formats. CBraMod is pre-trained on a very large corpus of EEG through patch-based masked EEG reconstruction. We evaluate CBraMod on up to 10 downstream BCI tasks (12 public datasets). CBraMod achieves the state-of-the-art performance across the wide range of tasks, proving its strong capability and generalizability. The source code is publicly available at https://github.com/wjq-learning/CBraMod. Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
ICLR | 8 |
| 2025 | BrainUICL: An Unsupervised Individual Continual Learning Framework for EEG ApplicationsabstractElectroencephalography (EEG) is a non-invasive brain-computer interface technology used for recording brain electrical activity. It plays an important role in human life and has been widely uesd in real life, including sleep staging, emotion recognition, and motor imagery. However, existing EEG-related models cannot be well applied in practice, especially in clinical settings, where new patients with individual discrepancies appear every day. Such EEG-based model trained on fixed datasets cannot generalize well to the continual flow of numerous unseen subjects in real-world scenarios. This limitation can be addressed through continual learning (CL), wherein the CL model can continuously learn and advance over time. Inspired by CL, we introduce a novel Unsupervised Individual Continual Learning paradigm for handling this issue in practice. We propose the BrainUICL framework, which enables the EEG-based model to continuously adapt to the incoming new subjects. Simultaneously, BrainUICL helps the model absorb new knowledge during each adaptation, thereby advancing its generalization ability for all unseen subjects. The effectiveness of the proposed BrainUICL has been evaluated on three different mainstream EEG tasks. The BrainUICL can effectively balance both the plasticity and stability during CL, achieving better plasticity on new individuals and better stability across all the unseen individuals, which holds significance in a practical setting. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
ICLR | 7 |
| 2025 | Hybrid Spiking Vision Transformer for Object Detection with Event CamerasabstractEvent-based object detection has attracted increasing attention for its high temporal resolution, wide dynamic range, and asynchronous address-event representation. Leveraging these advantages, spiking neural networks (SNNs) have emerged as a promising approach, offering low energy consumption and rich spatiotemporal dynamics. To further enhance the performance of event-based object detection, this study proposes a novel hybrid spike vision Transformer (HsVT) model. The HsVT model integrates a spatial feature extraction module to capture local and global features, and a temporal feature extraction module to model time dependencies and long-term patterns in event sequences. This combination enables HsVT to capture spatiotemporal features, improving its capability in handling complex event-based object detection tasks. To support research in this area, we developed the Fall Detection dataset as a benchmark for event-based object detection tasks. The Fall DVS detection dataset protects facial privacy and reduces memory usage thanks to its event-based representation. Experimental results demonstrate that HsVT outperforms existing SNN methods and achieves competitive performance compared to ANN-based models, with fewer parameters and lower energy consumption. Qi Xu 0008, Jiangrong Shen, Biwu Chen, Huajin Tang, Gang Pan 0001 |
ICML | 6 |
| 2025 | MetricEmbedding: Accelerate Metric Nearness by Tropical Inner ProductabstractThe Metric Nearness Problem involves restoring a non-metric matrix to its closest metric-compliant form, addressing issues such as noise, missing values, and data inconsistencies. Ensuring metric properties, particularly the $O(N^3)$ triangle inequality constraints, presents significant computational challenges, especially in large-scale scenarios where traditional methods suffer from high time and space complexity. We propose a novel solution based on the tropical inner product (max-plus operation), which we prove satisfies the triangle inequality for non-negative real matrices. By transforming the problem into a continuous optimization task, our method directly minimizes the distance to the target matrix. This approach not only restores metric properties but also generates metric-preserving embeddings, enabling real-time updates and reducing computational and storage overhead for downstream tasks. Experimental results demonstrate that our method achieves up to 60$\times$ speed improvements over state-of-the-art approaches, and efficiently scales from $1e4 \times 1e4$ to $1e5 \times 1e5$ matrices with significantly lower memory usage. Muyang Cao, Jiajun Yu, Xin Du 0002, Gang Pan 0001, Wei Wang 0011 |
ICML | 4 |
| 2025 | Efficient ANN-SNN Conversion with Error Compensation LearningabstractArtificial neural networks (ANNs) have demonstrated outstanding performance in numerous tasks, but deployment in resource-constrained environments remains a challenge due to their high computational and memory requirements. Spiking neural networks (SNNs) operate through discrete spike events and offer superior energy efficiency, providing a bio-inspired alternative. However, current ANN-to-SNN conversion often results in significant accuracy loss and increased inference time due to conversion errors such as clipping, quantization, and uneven activation. This paper proposes a novel ANN-to-SNN conversion framework based on error compensation learning. We introduce a learnable threshold clipping function, dual-threshold neurons, and an optimized membrane potential initialization strategy to mitigate the conversion error. Together, these techniques address the clipping error through adaptive thresholds, dynamically reduce the quantization error through dual-threshold neurons, and minimize the non-uniformity error by effectively managing the membrane potential. Experimental results on CIFAR-10, CIFAR-100, ImageNet datasets show that our method achieves high-precision and ultra-low latency among existing conversion methods. Using only two time steps, our method significantly reduces the inference time while maintains competitive accuracy of 94.75% on CIFAR-10 dataset under ResNet-18 structure. This research promotes the practical application of SNNs on low-power hardware, making efficient real-time processing possible. Chang Liu 0030, Jiangrong Shen, Xuming Ran, Mingkun Xu, Qi Xu 0008, Yi Xu 0008, Gang Pan 0001 |
ICML | 7 |
| 2025 | Flow Matching for Few-Trial Neural Adaptation with Stable Latent DynamicsabstractThe primary goal of brain-computer interfaces (BCIs) is to establish a direct linkage between neural activities and behavioral actions via neural decoders. Due to the nonstationary property of neural signals, BCIs trained on one day usually obtain degraded performance on other days, hindering the user experience. Existing studies attempted to address this problem by aligning neural signals across different days. However, these neural adaptation methods may exhibit instability and poor performance when only a few trials are available for alignment, limiting their practicality in real-world BCI deployment. To achieve efficient and stable neural adaptation with few trials, we propose Flow-Based Distribution Alignment (FDA), a novel framework that utilizes flow matching to learn flexible neural representations with stable latent dynamics, thereby facilitating source-free domain alignment through likelihood maximization. The latent dynamics of FDA framework is theoretically proven to be stable using Lyapunov exponents, allowing for robust adaptation. Further experiments across multiple motor cortex datasets demonstrate the superior performance of FDA, achieving reliable results with fewer than five trials. Our FDA approach offers a novel and efficient solution for few-trial neural data adaptation, offering significant potential for improving the long-term viability of real-world BCI applications. Puli Wang, Yueming Wang 0001, Gang Pan 0001 |
ICML | 4 |
| 2025 | Training High Performance Spiking Neural Network by Temporal Model CalibrationabstractSpiking Neural Networks (SNNs) are considered promising energy-efficient models due to their dynamic capability to process spatial-temporal spike information. Existing work has demonstrated that SNNs exhibit temporal heterogeneity, which leads to diverse outputs of SNNs at different time steps and has the potential to enhance their performance. Although SNNs obtained by direct training methods achieve state-of-the-art performance, current methods introduce limited temporal heterogeneity through the dynamics of spiking neurons or network structures. They lack the improvement of temporal heterogeneity through the lens of the gradient. In this paper, we first conclude that the diversity of the temporal logit gradients in current methods is limited. This leads to insufficient temporal heterogeneity and results in temporally miscalibrated SNNs with degraded performance. Based on the above analysis, we propose a Temporal Model Calibration (TMC) method, which can be seen as a logit gradient rescaling mechanism across time steps. Experimental results show that our method can improve the temporal logit gradient diversity and generate temporally calibrated SNNs with enhanced performance. In particular, our method achieves state-of-the-art accuracy on ImageNet, DVSCIFAR10, and N-Caltech101. Codes are available at https://github.com/zju-bmi-lab/TMC. Changping Wang, De Ma, Huajin Tang, Gang Pan 0001 |
ICML | 6 |
| 2025 | TS-SNN: Temporal Shift Module for Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) are increasingly recognized for their biological plausibility and energy efficiency, positioning them as strong alternatives to Artificial Neural Networks (ANNs) in neuromorphic computing applications. SNNs inherently process temporal information by leveraging the precise timing of spikes, but balancing temporal feature utilization with low energy consumption remains a challenge. In this work, we introduce Temporal Shift module for Spiking Neural Networks (TS-SNN), which incorporates a novel Temporal Shift (TS) module to integrate past, present, and future spike features within a single timestep via a simple yet effective shift operation. A residual combination method prevents information loss by integrating shifted and original features. The TS module is lightweight, requiring only one additional learnable parameter, and can be seamlessly integrated into existing architectures with minimal additional computational cost. TS-SNN achieves state-of-the-art performance on benchmarks like CIFAR-10 (96.72%), CIFAR-100 (80.28%), and ImageNet (70.61%) with fewer timesteps, while maintaining low energy consumption. This work marks a significant step forward in developing efficient and accurate SNN architectures. Kairong Yu, Tianqing Zhang, Qi Xu 0008, Gang Pan 0001, Hongwei Wang 0001 |
ICML | 4 |
| 2025 | HSRL: A Hierarchical Control System Based on Spiking Deep Reinforcement Learning for Robot NavigationabstractReinforcement Learning (RL) has shown promise in robotic navigation tasks, yet applying it to real-world environments remains challenging due to dynamic complexities and the need for dynamically feasible actions. We propose a hierarchical control framework based on Spiking Deep Reinforcement Learning (SDRL) for robust robot navigation in real environments. Our approach utilizes a two-layer architecture: a high-level decision layer powered by a Spiking GRU network for handling partially observable environments, and a low-level executive layer employing Continuous Attractor Neural Networks (CANNs) to ensure precise and continuous actions. This hierarchical structure allows real-time decisionmaking that respects the physical constraints of the robot. Experimental results show that our method adapts effectively to new environments without fine-tuning and surpasses existing methods in performance. We also explore the implementation on the Darwin3 chip, paving the way for biologically inspired motion control in future robotic applications. Shibo Zhou, Chaohui Lin, Qingao Chai, Rui Yan 0005, De Ma, Gang Pan 0001, Huajin Tang |
ICRA | 7 |
| 2025 | Bidirectional Distillation: A Mixed-Play Framework for Multi-Agent Generalizable Behaviors
Lang Feng 0002, Dong Xing, Li Zhang 0045, De Ma, Gang Pan 0001 |
AAMAS | 6 |
| 2025 | EDyGS: Event Enhanced Dynamic 3D Radiance Fields from Blurry Monocular VideoabstractThe task of generating novel views in dynamic scenes plays a critical role in the 3D vision domain. Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS) have shown great promise in this domain but struggle with motion blur, which often arises in real-world scenarios due to camera or object motion. Existing methods address camera motion blur but fall short in dynamic scenes, where the coupling of camera and object motion complicates multi-view consistency and temporal coherence. In this work, we propose EDyGS, a model designed to reconstruct sharp novel views from event streams and monocular videos of dynamic scenes with motion blur. Our approach introduces a motion-mask 3D Gaussian model that assigns each Gaussian an additional attribute to distinguish between static and dynamic regions. By leveraging this motion mask field, we separate and optimize the static and dynamic regions independently. A progressive learning strategy is adopted, where static regions are reconstructed by jointly optimizing camera poses and learnable 3D Gaussians, while dynamic regions are modeled using an implicit deformation field alongside learnable 3D Gaussians. We conduct both quantitative and qualitative experiments on synthetic and real-world data. Experimental results demonstrate that EDyGS effectively handles blurry inputs in dynamic scenes. Mengxu Lu, De Ma, Huajin Tang, Gang Pan 0001 |
IJCAI | 7 |
| 2025 | An NoC-Based Latency Upper-Bound Model for Reducing Timestep Length in SNN CommunicationabstractSpiking Neural Networks (SNNs) with real-time demands are deployed on Network-on-Chip (NoC)-based neuromorphic processors for specific tasks. While meeting real-time constraints for spiking data streams is crucial, overly long timesteps (e.g., 1ms) result in idle neuron cores and routers, reducing efficiency. This paper proposes a communication performance model using a recursive calculation method to assess the worst-case upper-bound latency. Integrated with the NoC router microarchitecture, the model analyzes the behavior of spiking streams during communication. It effectively balances timestep length to ensure real-time communication while minimizing idle time. Empirical results show that our model outperforms worst-contention delay (WCD) and worst-case traversal time (WCTT) models, achieving latency reductions of 0.55× to 2.53× compared to WCTT and 0.83× to 25.57× compared to WCD across various spiking datasets and virtual channels. Encouragingly, with only a 1% accuracy reduction, latency was reduced by 18%. Ziyang Kang, Lei Wang 0011, De Ma, Gang Pan 0001 |
ISCAS | 5 |
| 2025 | Spiking Neural Networks with Temporal Attention-Guided Adaptive Fusion for imbalanced Multi-modal Learning
Jiangrong Shen, Yulin Xie, Qi Xu 0008, Gang Pan 0001, Huajin Tang, Badong Chen |
ACM Multimedia | 4 |
| 2025 | Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion
Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Shijian Li, Shurong Dong, Gang Pan 0001 |
ACM Multimedia | 9 |
| 2025 | Neural-Driven Image EditingabstractTraditional image editing typically relies on manual prompting, making it labor-intensive and inaccessible to individuals with limited motor control or language abilities. Leveraging recent advances in brain-computer interfaces (BCIs) and generative models, we propose LoongX, a hands-free image editing approach driven by multimodal neurophysiological signals.
LoongX utilizes state-of-the-art diffusion models trained on a comprehensive dataset of 23,928 image editing pairs, each paired with synchronized electroencephalography (EEG), functional near-infrared spectroscopy (fNIRS), photoplethysmography (PPG), and head motion signals that capture user intent.
To effectively address the heterogeneity of these signals, LoongX integrates two key modules. The cross-scale state space (CS3) module encodes informative modality-specific features. The dynamic gated fusion (DGF) module further aggregates these features into a unified latent space, which is then aligned with edit semantics via fine-tuning on a diffusion transformer (DiT).
Additionally, we pre-train the encoders using contrastive learning to align cognitive states with semantic intentions from embedded natural language.
Extensive experiments demonstrate that LoongX achieves performance comparable to text-driven methods (CLIP-I: 0.6605 vs. 0.6558; DINO: 0.4812 vs. 0.4637) and outperforms them when neural signals are combined with speech (CLIP-T: 0.2588 vs. 0.2549). These results highlight the promise of neural-driven generative models in enabling accessible, intuitive image editing and open new directions for cognitive-driven creative technologies. The code and dataset are released on the project website: https://loongx1.github.io. Xiaopeng Peng 0001, Wangbo Zhao, Zilong Ye, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Kai Wang 0036, Xiaojun Chang, Gang Pan 0001, Shurong Dong, Kaipeng Zhang, Yang You 0001 |
NeurIPS | 14 |
| 2025 | SPICED: A Synaptic Homeostasis-Inspired Framework for Unsupervised Continual EEG DecodingabstractHuman brain achieves dynamic stability-plasticity balance through synaptic homeostasis, a self-regulatory mechanism that stabilizes critical memory traces while preserving optimal learning capacities. Inspired by this biological principle, we propose SPICED: a neuromorphic framework that integrates the synaptic homeostasis mechanism for unsupervised continual EEG decoding, particularly addressing practical scenarios where new individuals with inter-individual variability emerge continually. SPICED comprises a novel synaptic network that enables dynamic expansion during continual adaptation through three bio-inspired neural mechanisms: (1) critical memory reactivation, which mimics brain functional specificity, selectively activates task-relevant memories to facilitate adaptation; (2) synaptic consolidation, which strengthens these reactivated critical memory traces and enhances their replay prioritizations for further adaptations and (3) synaptic renormalization, which are periodically triggered to weaken global memory traces to preserve learning capacities. The interplay within synaptic homeostasis dynamically strengthens task-discriminative memory traces and weakens detrimental memories. By integrating these mechanisms with continual learning system, SPICED preferentially replays task-discriminative memory traces that exhibit strong associations with newly emerging individuals, thereby achieving robust adaptations. Meanwhile, SPICED effectively mitigates catastrophic forgetting by suppressing the replay prioritization of detrimental memories during long-term continual learning. Validated on three EEG datasets, SPICED show its effectiveness. More importantly, SPICED bridges biological neural mechanisms and artificial intelligence through synaptic homeostasis, providing insights into the broader applicability of bio-inspired principles. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
NeurIPS | 7 |
| 2025 | FGDC: A fine-grained divide-and-conquer approach for extending NCO to solve large-scale Traveling Salesman Problem
Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Expert Syst. Appl. | 6 |
| 2025 | M-MDD: A multi-task deep learning framework for major depressive disorder diagnosis using EEG
Yilin Wang 0014, Sha Zhao, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
Neurocomputing | 6 |
| 2025 | NeuroSimWorm: A multisensory framework for modeling and simulating neural circuits of Caenorhabditis elegansabstractBiological behaviors emerge from the dynamic interplay among the inner neurodynamic system, embodied mechanical structure, and external environmental inputs. Nonetheless, existing approaches simply consider the static brain model that cannot fully exploit the potential of continuous interaction and feedback from the body and the environment. To address these problems, we introduce NeuroSimWorm , a multisensory closed-loop neural circuit simulation approach of the widely studied organism, Caenorhabditis elegans ( C. elegans ). The full closed-loop simulation platform integrates four key subcomponents including Environment, Neural Computing, Biomechanical Model, and Visualization modules. We initially define multiple sensory environments for chemical, mechanical, and thermal stimuli. Subsequently, we construct four types of neural circuits including Locomotion, Chemosensation , Thermosensation, and Mechanosensation . Fitness functions based on multisensory optimization are proposed, enabling the virtual nematode to achieve complex intelligent behaviors such as autonomous locomotion, foraging, thermosensory regulation, and tactile avoidance. Thus, NeuroSimWorm presents a feasible way to understand the mechanism of biological intelligence by modeling the connectome and simulating behaviors in the physical surroundings. Mengxiao Zhang 0003, Lijun Kang, Gang Pan 0001, Huajin Tang |
Neurocomputing | 6 |
| 2025 | Context gating in spiking neural networks: Achieving lifelong learning through integration of local and global plasticity
Jiangrong Shen, Wenyao Ni, Qi Xu 0008, Gang Pan 0001, Huajin Tang |
Knowl. Based Syst. | 4 |
| 2025 | Temporal spiking generative adversarial networks for heading direction decoding
Jiangrong Shen, Jian K. Liu, Qi Xu 0008, Gang Pan 0001, Xiaodong Chen 0005, Huajin Tang |
Neural Networks | 6 |
| 2025 | Post-training quantization for efficient ANN-SNN conversion
Ruimin Sun, De Ma, Gang Pan 0001 |
Neural Networks | 3 |
| 2025 | EEGMamba: An EEG foundation model with Mamba
Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Shijian Li, Gang Pan 0001 |
Neural Networks | 6 |
| 2025 | Revisiting Supervised Learning-Based Photometric Stereo NetworksabstractDeep learning has significantly propelled the development of photometric stereo by handling the challenges posed by unknown reflectance and global illumination effects. However, how supervised learning-based photometric stereo networks resolve these challenges remains to be elucidated. In this paper, we aim to reveal how existing methods address these challenges by revisiting their deep features, deep feature encoding strategies, and network architectures. Based on the insights gained from our analysis, we propose ESSENCE-Net, which effectively encodes deep shading features with an easy-first-encoding strategy, enhances shading features with shading supervision, and accurately decodes normal with spatial context-aware attention. The experimental results verify that the proposed method outperforms state-of-the-art methods on three benchmark datasets, whether with dense or sparse inputs. Xiaoyao Wei, Zongrui Li 0001, Binjie Ding, Boxin Shi, Xudong Jiang 0001, Gang Pan 0001, Yanlong Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Human-Inspired Computing for Robust and Efficient Audio-Visual Speech RecognitionabstractHumans excel at audiovisual speech recognition (AVSR), motivating the development of human-inspired computing for robust and efficient AVSR models. Spiking neural networks (SNNs), mimicking the brain’s information-processing mechanisms, offer a promising foundation. However, research on SNN-based AVSR remains limited, with most audio-visual methods focusing on object or digit recognition. These methods oversimplify multimodal fusion, neglecting modality-specific characteristics and interactions. Additionally, they often rely on future information, increasing recognition latency and limiting real-time applicability. Inspired by human speech perception, this paper proposes a novel human-inspired SNN named HI-AVSNN for AVSR, incorporating three computing characteristics: spike activity, cueing interaction, and causal processing. For cueing interaction, we introduce a Spike-Driven Visual-Cued Speech Processing (sVCSP) scheme, where visual features hierarchically guide speech processing to enhance critical features. For causal processing, we align the temporal dimensions of SNN with audio-visual inputs and apply temporal masking to ensure only past and current information is used. For spike activity, in addition to SNNs, we incorporate event cameras to capture lip movements as spikes, efficiently encoding visual data like the human retina. Experiments on two event-based AVSR datasets demonstrate our method outperforms existing audio-visual SNN fusion techniques, showcasing the effectiveness, robustness, and efficiency achieved through our human-inspired computing. Qianhui Liu, Yang Wang 0106, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001 |
IEEE Trans. Computers | 5 |
| 2025 | LAC-PS: A Light Direction Selection Policy Under the Accuracy Constraint for Photometric StereoabstractPhotometric stereo (PS) methods recover surface normals from appearance changes under varying light directions, excelling in tasks like 3D surface reconstruction and defect inspection. However, collecting the illumination images is expensive, and current PS methods cannot obtain the light direction set that satisfies the pre-defined accuracy constraint, limiting their adaptability to various applications with varying accuracy requirements. To address this issue, we propose the LAC-PS, a light direction selection policy under the accuracy constraint for photometric stereo, which optimizes the light direction set to meet target reconstruction accuracy. In our method, we develop an accuracy assessment network that estimates reconstruction accuracy without ground truth. With this estimated accuracy, we put forward a reinforcement learning-based method that can utilize policy to sequentially select light directions and obtain the light directions satisfying the desired PS recovery accuracy constraint. Experimental results on real and synthetic datasets demonstrate that our method effectively selects light directions that satisfy accuracy constraints. Wenjia Meng, Huimin Han, Xiankai Lu, Yilong Yin, Gang Pan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MindGPT: Interpreting What You See With Non-Invasive Brain RecordingsabstractDecoding of seen visual contents with non-invasive brain recordings has important scientific and practical values. Efforts have been made to recover the seen images from brain signals. However, most existing approaches cannot faithfully reflect the visual contents due to insufficient image quality or semantic mismatches. Compared with reconstructing pixel-level visual images, speaking is a more efficient and effective way to explain visual information. Here we introduce a non-invasive neural decoder, termed MindGPT, which interprets perceived visual stimuli into natural languages from functional Magnetic Resonance Imaging (fMRI) signals in an end-to-end manner. Specifically, our model builds upon a visually guided neural encoder with a cross-attention mechanism. By the collaborative use of data augmentation techniques, this architecture permits us to guide latent neural representations towards a desired language semantic direction in a self-supervised fashion. Through doing so, we found that the neural representations of the MindGPT are explainable, which can be used to evaluate the contributions of visual properties to language semantics. Our experiments show that the generated word sequences truthfully represented the visual information (with essential details) conveyed in the seen stimuli. The results also suggested that with respect to language decoding tasks, the higher visual cortex (HVC) is more semantically informative than the lower visual cortex (LVC), and using only the HVC can recover most of the semantic information. The source code for the MindGPT model is publicly available at https://github.com/JxuanC/MindGPT. Jiaxuan Chen 0002, Yueming Wang 0001, Gang Pan 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | EvoMoE: Evolutionary Mixture-of-Experts for SSVEP-EEG Classification With User-Independent TrainingabstractThe analysis of EEG data in BCI systems captures unique individual characteristics, presenting diverse patterns that deviate from conventional identical distribution assumptions. Therefore, applying AI models directly to brain data becomes challenging due to the non-identical distribution issue. Meanwhile, as user numbers in BCI systems rise, scalable models are crucial to handle the growing data volume. Moreover, the limited availability of individual data necessitates the use of collective data for training, requiring models with strong generalization capabilities. To address these challenges, we propose Evolutionary Mixture of Experts (EvoMoE), a framework leveraging a set of diverse experts to model data from individuals. Users with similar distributions are grouped together, allowing experts to handle EEG data with different distribution types. The gating network of EvoMoE selects experts that closely match the distribution of the current sample, effectively tackling non-identical distribution issues. When encountering an unrecognized distribution, a new expert is introduced to accommodate the new data pattern, ensuring model adaptability. Evaluations on two 40-category BCI Speller datasets demonstrate significant performance improvements over state-of-the-art methods. On the BETA dataset, our online EvoMoE achieves 13.06% increase in accuracy and a 27.24-point increase in high information transfer rate (ITR) compared to the online UI method. The Bench dataset shows 3.64% increase in accuracy and a 10.42-point increase in ITR. These qualities make it a promising solution for practical BCI implementation, while setting the stage for the development of comprehensive biological big models. Jianyu Zhang 0001, Hui-yuan Tian, Shijian Li, Gang Pan 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Off-OAB: Off-Policy Policy Gradient Method With Optimal Action-Dependent BaselineabstractThe policy-based methods have achieved remarkable success in solving challenging reinforcement learning (RL) problems. Among these methods, the off-policy policy gradient (OPPG) methods are particularly important because they can benefit from off-policy data. However, these methods suffer from the high variance of the OPPG estimator, which results in poor sample efficiency during training. In this article, we propose an off-policy policy gradient method with the optimal action-dependent baseline (Off-OAB) to mitigate this variance issue. Specifically, this baseline maintains the OPPG estimator's unbiasedness while theoretically minimizing its variance. To enhance practical computational efficiency, we design an approximated version of this optimal baseline. Utilizing this approximation, our method (Off-OAB) aims to decrease the OPPG estimator's variance during policy optimization. We evaluate the proposed Off-OAB method on six representative tasks from OpenAI Gym and MuJoCo, where it demonstrably surpasses the state-of-the-art methods on the majority of these tasks. Wenjia Meng, Long Yang 0004, Yilong Yin, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Robust Sensory Information Reconstruction and Classification With Augmented SpikesabstractSensory information recognition is primarily processed through the ventral and dorsal visual pathways in the primate brain visual system, which exhibits layered feature representations bearing a strong resemblance to convolutional neural networks (CNNs), encompassing reconstruction and classification. However, existing studies often treat these pathways as distinct entities, focusing individually on pattern reconstruction or classification tasks, overlooking a key feature of biological neurons, the fundamental units for neural computation of visual sensory information. Addressing these limitations, we introduce a unified framework for sensory information recognition with augmented spikes. By integrating pattern reconstruction and classification within a single framework, our approach not only accurately reconstructs multimodal sensory information but also provides precise classification through definitive labeling. Experimental evaluations conducted on various datasets including video scenes, static images, dynamic auditory scenes, and functional magnetic resonance imaging (fMRI) brain activities demonstrate that our framework delivers state-of-the-art pattern reconstruction quality and classification accuracy. The proposed framework enhances the biological realism of multimodal pattern recognition models, offering insights into how the primate brain visual system effectively accomplishes the reconstruction and classification tasks through the integration of ventral and dorsal pathways. Qi Xu 0008, Sibo Liu, Xuming Ran, Jiangrong Shen, Huajin Tang, Jian K. Liu, Gang Pan 0001, Qiang Zhang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Mapping Large-Scale Spiking Neural Network on Arbitrary Meshed Neuromorphic HardwareabstractNeuromorphic hardware systems—designed as 2D-mesh structures with parallel neurosynaptic cores—have proven highly efficient at executing large-scale spiking neural networks (SNNs). A critical challenge, however, lies in mapping neurons efficiently to these cores. While existing approaches work well with regular, fully functional mesh structures, they falter in real-world scenarios where hardware has irregular shapes or non-functional cores caused by defects or resource fragmentation. To address these limitations, we propose a novel mapping method based on an innovative space-filling curve: the Adaptive Locality-Preserving (ALP) curve. Using a unique divide-and-conquer construction algorithm, the ALP curve ensures adaptability to meshes of any shape while maintaining crucial locality properties—essential for efficient mapping. Our method demonstrates exceptional computational efficiency, making it ideal for large-scale deployments. These distinctive characteristics enable our approach to handle complex scenarios that challenge conventional methods. Experimental results show that our method matches state-of-the-art solutions in regular-shape mapping while achieving significant improvements in irregular scenarios, reducing communication overhead by up to 57.1%. Ouwen Jin, Qinghui Xing, Zhuo Chen 0044, Ming Zhang 0018, De Ma, Ying Li 0001, Xin Du 0002, Shuibing He, Shuiguang Deng, Gang Pan 0001 |
IEEE Trans. Parallel Distributed Syst. | 10 |
| 2025 | HetSub: A Heterogeneous Multi-NoC With Reconfigurable Long-Range Links for Neuromorphic Systems
Youneng Hu, Xiaofei Jin, Ziyang Kang, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Bridging the Semantic Latent Space between Brain and Machine: Similarity Is All You NeedabstractHow our brain encodes complex concepts has been a longstanding mystery in neuroscience. The answer to this problem can lead to new understandings about how the brain retrieves information in large-scale data with high efficiency and robustness. Neuroscience studies suggest the brain represents concepts in a locality-sensitive hashing (LSH) strategy, i.e., similar concepts will be represented by similar responses. This finding has inspired the design of similarity-based algorithms, especially in contrastive learning. Here, we hypothesize that the brain and large neural network models, both using similarity-based learning rules, could contain a similar semantic embedding space. To verify that, this paper proposes a functional Magnetic Resonance Imaging (fMRI) semantic learning network named BrainSem, aimed at seeking a joint semantic latent space that bridges the brain and a Contrastive Language-Image Pre-training (CLIP) model. Given that our perception is inherently cross-modal, we introduce a fuzzy (one-to-many) matching loss function to encourage the models to extract high-level semantic components from neural signals. Our results claimed that using only a small set of fMRI recordings for semantic space alignment, we could obtain shared embedding valid for unseen categories out of the training set, which provided potential evidence for the semantic representation similarity between the brain and large neural networks. In a zero-shot classification task, our BrainSem achieves an 11.6% improvement over the state-of-the-art. Jiaxuan Chen 0007, Yueming Wang 0001, Gang Pan 0001 |
AAAI | 4 |
| 2024 | Spiking NeRF: Representing the Real-World Geometry by a Discontinuous RepresentationabstractA crucial reason for the success of existing NeRF-based methods is to build a neural density field for the geometry representation via multiple perceptron layers (MLPs). MLPs are continuous functions, however, real geometry or density field is frequently discontinuous at the interface between the air and the surface. Such a contrary brings the problem of unfaithful geometry representation. To this end, this paper proposes spiking NeRF, which leverages spiking neurons and a hybrid Artificial Neural Network (ANN)-Spiking Neural Network (SNN) framework to build a discontinuous density field for faithful geometry representation. Specifically, we first demonstrate the reason why continuous density fields will bring inaccuracy. Then, we propose to use the spiking neurons to build a discontinuous density field. We conduct a comprehensive analysis for the problem of existing spiking neuron models and then provide the numerical relationship between the parameter of the spiking neuron and the theoretical accuracy of geometry. Based on this, we propose a bounded spiking neuron to build the discontinuous density field. Our method achieves SOTA performance. The source code and the supplementary material are available at https://github.com/liaozhanfeng/Spiking-NeRF. Zhanfeng Liao, Gang Pan 0001 |
AAAI | 4 |
| 2024 | Generalizable Sleep Staging via Multi-Level Domain AlignmentabstractAutomatic sleep staging is essential for sleep assessment and disorder diagnosis. Most existing methods depend on one specific dataset and are limited to be generalized to other unseen datasets, for which the training data and testing data are from the same dataset. In this paper, we introduce domain generalization into automatic sleep staging and propose the task of generalizable sleep staging which aims to improve the model generalization ability to unseen datasets. Inspired by existing domain generalization methods, we adopt the feature alignment idea and propose a framework called SleepDG to solve it. Considering both of local salient features and sequential features are important for sleep staging, we propose a Multi-level Feature Alignment combining epoch-level and sequence-level feature alignment to learn domain-invariant feature representations. Specifically, we design an Epoch-level Feature Alignment to align the feature distribution of each single sleep epoch among different domains, and a Sequence-level Feature Alignment to minimize the discrepancy of sequential features among different domains. SleepDG is validated on five public datasets, achieving the state-of-the-art performance. Jiquan Wang, Sha Zhao, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
AAAI | 6 |
| 2024 | Mind Artist: Creating Artistic Snapshots with Human ThoughtabstractWe introduce Mind Artist (MindArt), a novel and efficient neural decoding architecture to snap artistic photographs from our mind in a controllable manner. Recently, progress has been made in image reconstruction with non-invasive brain recordings, but it's still difficult to generate realistic images with high semantic fidelity due to the scarcity of data annotations. Unlike previous methods, this work casts the neural decoding into optimal transport (OT) and representation decoupling problems. Specifically, under discrete OT theory, we design a graph matching-guided neural representation learning framework to seek the underlying correspondences between conceptual semantics and neural signals, which yields a natural and meaningful self-supervisory task. Moreover, the proposed MindArt, structured with multiple stand-alone modal branches, enables the seamless incorporation of semantic representation into any visual style information, thus leaving it to have multi-modal reconstruction and training-free semantic editing capabilities. By doing so, the reconstructed images of MindArt have phenomenal realism both in terms of semantics and appearance. We compare our MindArt with leading alternatives, and achieve SOTA performance in different decoding tasks. Importantly, our approach can directly generate a series of stylized “mind snapshots” w/o extra optimizations, which may open up more potential applications. Code is available at https://github.com/JxuanC/MindArt. Jiaxuan Chen 0007, Yueming Wang 0001, Gang Pan 0001 |
CVPR | 4 |
| 2024 | Spin-UP: Spin Light for Natural Light Uncalibrated Photometric StereoabstractNatural Light Uncalibrated Photometric Stereo (NaUPS) relieves the strict environment and light assumptions in classical Uncalibrated Photometric Stereo (UPS) methods. However, due to the intrinsic ill-posedness and high-dimensional ambiguities, addressing NaUPS is still an open question. Existing works impose strong assumptions on the environment lights and objects' material, restricting the effectiveness in more general scenarios. Alternatively, some methods leverage supervised learning with intricate models while lacking interpretability, resulting in a biased estimation. In this work, we propose Spin Light Uncalibrated Photometric Stereo (Spin-UP), an unsupervised method to tackle NaUPS in various environment lights and objects. The proposed method uses a novel setup that captures the object's images on a rotatable platform, which mitigates NaUPS's ill-posedness by reducing unknowns and provides reliable priors to alleviate NaUPS's ambiguities. Leveraging neural inverse rendering and the proposed training strategies, Spin-UP recovers surface normals, environment light, and isotropic reflectance under complex natural light with low computational cost. Experiments have shown that Spin-UP outperforms other supervised / unsupervised NaUPS meth-ods and achieves state-of-the-art performance on synthetic and real-world datasets. Codes and data are available at https://github.com/LMozart/CVPR2024-SpinUP. Zongrui Li 0001, Zhan Lu, Haojie Yan, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001 |
CVPR | 5 |
| 2024 | Adaptive deep spiking neural network with global-local learning via balanced excitatory and inhibitory mechanismabstractThe training method of Spiking Neural Networks (SNNs) is an essential problem, and how to integrate local and global learning is a worthy research interest. However, the current integration methods do not consider the network conditions suitable for local and global learning, and thus fail to balance their advantages. In this paper, we propose an Excitation-Inhibition Mechanism-assisted Hybrid Learning(EIHL) algorithm that adjusts the network connectivity by using the excitation-inhibition mechanism and then switches between local and global learning according to the network connectivity. The experimental results on CIFAR10/100 and DVS-CIFAR10 demonstrate that the EIHL not only has better accuracy performance than other methods but also has excellent sparsity advantage. Especially, the Spiking VGG11 is trained by EIHL, STBP, and STDP on DVS_CIFAR10, respectively. The accuracy of the Spiking VGG11 model on EIHL is 62.45%, which is 4.35% higher than STBP and 11.40% higher than STDP, and the sparsity is 18.74%, which is 18.74% higher than the other two methods. Moreover, the excitation-inhibition mechanism used in our method also offers a new perspective on the field of SNN learning. Qi Xu 0008, Xuming Ran, Jiangrong Shen, Pan Lv, Qiang Zhang 0008, Gang Pan 0001 |
ICLR | 7 |
| 2024 | Safe Reinforcement Learning using Finite-Horizon Gradient-based EstimationabstractA key aspect of Safe Reinforcement Learning (Safe RL) involves estimating the constraint condition for the next policy, which is crucial for guiding the optimization of safe policy updates. However, the existing Advantage-based Estimation (ABE) method relies on the infinite-horizon discounted advantage function. This dependence leads to catastrophic errors in finite-horizon scenarios with non-discounted constraints, resulting in safety-violation updates. In response, we propose the first estimation method for finite-horizon non-discounted constraints in deep Safe RL, termed Gradient-based Estimation (GBE), which relies on the analytic gradient derived along trajectories. Our theoretical and empirical analyses demonstrate that GBE can effectively estimate constraint changes over a finite horizon. Constructing a surrogate optimization problem with GBE, we developed a novel Safe RL algorithm called Constrained Gradient-based Policy Optimization (CGPO). CGPO identifies feasible optimal policies by iteratively resolving sub-problems within trust regions. Our empirical results reveal that CGPO, unlike baseline algorithms, successfully estimates the constraint functions of subsequent policies, thereby ensuring the efficiency and feasibility of each update. Juntao Dai, Yaodong Yang 0001, Gang Pan 0001 |
ICML | 4 |
| 2024 | Resisting Stochastic Risks in Diffusion Planners with the Trajectory Aggregation TreeabstractDiffusion planners have shown promise in handling long-horizon and sparse-reward tasks due to the non-autoregressive plan generation. However, their inherent stochastic risk of generating infeasible trajectories presents significant challenges to their reliability and stability. We introduce a novel approach, the Trajectory Aggregation Tree (TAT), to address this issue in diffusion planners. Compared to prior methods that rely solely on raw trajectory predictions, TAT aggregates information from both historical and current trajectories, forming a dynamic tree-like structure. Each trajectory is conceptualized as a branch and individual states as nodes. As the structure evolves with the integration of new trajectories, unreliable states are marginalized, and the most impactful nodes are prioritized for decision-making. TAT can be deployed without modifying the original training and sampling pipelines of diffusion planners, making it a training-free, ready-to-deploy solution. We provide both theoretical analysis and empirical evidence to support TAT’s effectiveness. Our results highlight its remarkable ability to resist the risk from unreliable trajectories, guarantee the performance boosting of diffusion planners in 100% of tasks, and exhibit an appreciable tolerance margin for sample quality, thereby enabling planning with a more than $3\times$ acceleration. Lang Feng 0002, Pengjie Gu, Bo An 0001, Gang Pan 0001 |
ICML | 4 |
| 2024 | Towards efficient deep spiking neural networks construction with spiking activity based pruningabstractThe emergence of deep and large-scale spiking neural networks (SNNs) exhibiting high performance across diverse complex datasets has led to a need for compressing network models due to the presence of a significant number of redundant structural units, aiming to more effectively leverage their low-power consumption and biological interpretability advantages. Currently, most model compression techniques for SNNs are based on unstructured pruning of individual connections, which requires specific hardware support. Hence, we propose a structured pruning approach based on the activity levels of convolutional kernels named Spiking Channel Activity-based (SCA) network pruning framework. Inspired by synaptic plasticity mechanisms, our method dynamically adjusts the network’s structure by pruning and regenerating convolutional kernels during training, enhancing the model’s adaptation to the current target task. While maintaining model performance, this approach refines the network architecture, ultimately reducing computational load and accelerating the inference process. This indicates that structured dynamic sparse learning methods can better facilitate the application of deep SNNs in low-power and high-efficiency scenarios. Qi Xu 0008, Jiangrong Shen, Hongming Xu 0002, Long Chen 0019, Gang Pan 0001 |
ICML | 6 |
| 2024 | A Framework for Image Synthesis Using Supervised Contrastive Learning
Jianyu Zhang 0001, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ICPR (6) | 5 |
| 2024 | LitE-SNN: Designing Lightweight and Efficient Spiking Neural Network through Spatial-Temporal Compressive Network Search and Joint Optimization
Qianhui Liu, Malu Zhang, Gang Pan 0001, Haizhou Li 0001 |
IJCAI | 4 |
| 2024 | Event-ID: Intrinsic Decomposition Using an Event CameraabstractReconstructing 3D scenes from multi-view images is challenging, especially under extreme scenarios. We propose Event-ID, an event-based intrinsic decomposition framework that leverages events and images for stable decomposition under extreme scenarios. Our method is based on two observations: event cameras maintain good imaging quality under blurry or poorly exposed scenarios, and event signals from different viewpoints exhibit similarity in diffuse regions while varying in specular regions. We establish an event-based reflectance model and introduce an event-based warping method to extract specular clues. Our two-stage framework constructs a radiance field and decomposes the scene into normal, material, and lighting. Experimental results demonstrate superior performance compared to state-of-the-art methods. Our project can be found at https://zehaoc.github.io/EventID.github.io/ Zhan Lu, De Ma, Huajin Tang, Xudong Jiang 0001, Gang Pan 0001 |
ACM Multimedia | 7 |
| 2024 | RSNN: Recurrent Spiking Neural Networks for Dynamic Spatial-Temporal Information Processing
Qi Xu 0008, Xuanye Fang, Jiangrong Shen, De Ma, Yi Xu 0008, Gang Pan 0001 |
ACM Multimedia | 7 |
| 2024 | Reversing Structural Pattern Learning with Biologically Inspired Knowledge Distillation for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have superb characteristics in sensory information recognition tasks due to their biological plausibility. However, the performance of some current spiking-based models is limited by their structures which means either fully connected or too-deep structures bring too much redundancy. This redundancy from both connection and neurons is one of the key factors hindering the practical application of SNNs. Although Some pruning methods were proposed to tackle this problem, they normally ignored the fact the neural topology in the human brain could be adjusted dynamically. Inspired by this, this paper proposed an evolutionary-based structure construction method for constructing more reasonable SNNs. By integrating the knowledge distillation and connection pruning method, the synaptic connections in SNNs can be optimized dynamically to reach an optimal state. As a result, the structure of SNNs could not only absorb knowledge from the teacher model but also search for deep but sparse network topology. Experimental results on CIFAR100, Tiny-imagenet and DVS-Gesture show that the proposed structure learning method can get pretty well performance while reducing the connection redundancy. The proposed method explores a novel dynamical way for structure learning from scratch in SNNs which could build a bridge to close the gap between deep learning and bio-inspired neural dynamics. Qi Xu 0008, Xuanye Fang, Jiangrong Shen, Qiang Zhang 0008, Gang Pan 0001 |
ACM Multimedia | 6 |
| 2024 | Rethinking the Membrane Dynamics and Optimization Objectives of Spiking Neural NetworksabstractDespite spiking neural networks (SNNs) have demonstrated notable energy efficiency across various fields, the limited firing patterns of spiking neurons within fixed time steps restrict the expression of information, which impedes further improvement of SNN performance. In addition, current implementations of SNNs typically consider the firing rate or average membrane potential of the last layer as the output, lacking exploration of other possibilities. In this paper, we identify that the limited spike patterns of spiking neurons stem from the initial membrane potential (IMP), which is set to 0. By adjusting the IMP, the spiking neurons can generate additional firing patterns and pattern mappings. Furthermore, we find that in static tasks, the accuracy of SNNs at each time step increases as the membrane potential evolves from zero. This observation inspires us to propose a learnable IMP, which can accelerate the evolution of membrane potential and enables higher performance within a limited number of time steps. Additionally, we introduce the last time step (LTS) approach to accelerate convergence in static tasks, and we propose a label smooth temporal efficient training (TET) loss to mitigate the conflicts between optimization objective and regularization term in the vanilla TET. Our methods improve the accuracy by 4.05\% on ImageNet compared to baseline and achieve state-of-the-art performance of 87.80\% on CIFAR10-DVS and 87.86\% on N-Caltech101. Hangchi Shen, Gang Pan 0001 |
NeurIPS | 4 |
| 2024 | FEEL-SNN: Robust Spiking Neural Networks with Frequency Encoding and Evolutionary Leak FactorabstractCurrently, researchers think that the inherent robustness of spiking neural networks (SNNs) stems from their biologically plausible spiking neurons, and are dedicated to developing more bio-inspired models to defend attacks. However, most work relies solely on experimental analysis and lacks theoretical support, and the direct-encoding method and fixed membrane potential leak factor they used in spiking neurons are simplified simulations of those in the biological nervous system, which makes it difficult to ensure generalizability across all datasets and networks. Contrarily, the biological nervous system can stay reliable even in a highly complex noise environment, one of the reasons is selective visual attention and non-fixed membrane potential leaks in biological neurons. This biological finding has inspired us to design a highly robust SNN model that closely mimics the biological nervous system. In our study, we first present a unified theoretical framework for SNN robustness constraint, which suggests that improving the encoding method and evolution of the membrane potential leak factor in spiking neurons can improve SNN robustness. Subsequently, we propose a robust SNN (FEEL-SNN) with Frequency Encoding (FE) and Evolutionary Leak factor (EL) to defend against different noises, mimicking the selective visual attention mechanism and non-fixed leak observed in biological systems. Experimental results confirm the efficacy of both our FE, EL, and FEEL methods, either in isolation or in conjunction with established robust enhancement algorithms, for enhancing the robustness of SNNs. Mengting Xu, De Ma, Huajin Tang, Gang Pan 0001 |
NeurIPS | 5 |
| 2024 | State-sensitive convolutional sparse coding for potential biomarker identification in brain signals
Puli Wang, Gang Pan 0001 |
Sci. China Inf. Sci. | 3 |
| 2024 | SpikingMiniLM: energy-efficient spiking transformer for natural language understanding
Jiangrong Shen, Zeke Wang, Qinghai Guo, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Sci. China Inf. Sci. | 6 |
| 2024 | Multi-depth branch network for efficient image super-resolution
Hui-yuan Tian, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Image Vis. Comput. | 5 |
| 2024 | Efficient spiking neural network design via neural architecture search
Qianhui Liu, Malu Zhang, Lang Feng 0002, De Ma, Haizhou Li 0001, Gang Pan 0001 |
Neural Networks | 7 |
| 2024 | Trainable Spiking-YOLO for low-latency and high-performance object detection
Mengwen Yuan, Huixiang Liu, Gang Pan 0001, Huajin Tang |
Neural Networks | 5 |
| 2024 | Enhancing SNN-based spatio-temporal learning: A benchmark dataset and Cross-Modality Attention model
Shibo Zhou, Mengwen Yuan, Runhao Jiang, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Neural Networks | 6 |
| 2024 | SPLiT: Single Portrait Lighting Estimation via a Tetrad of Face IntrinsicsabstractThis paper proposes a novel pipeline to estimate a non-parametric environment map with high dynamic range from a single human face image. Lighting-independent and -dependent intrinsic images of the face are first estimated separately in a cascaded network. The influence of face geometry on the two lighting-dependent intrinsics, diffuse shading and specular reflection, are further eliminated by distributing the intrinsics pixel-wise onto spherical representations using the surface normal as indices. This results in two representations simulating images of a diffuse sphere and a glossy sphere under the input scene lighting. Taking into account the distinctive nature of light sources and ambient terms, we further introduce a two-stage lighting estimator to predict both accurate and realistic lighting from these two representations. Our model is trained supervisedly on a large-scale and high-quality synthetic face image dataset. We demonstrate that our method allows accurate and detailed lighting estimation and intrinsic decomposition, outperforming state-of-the-art methods both qualitatively and quantitatively on real face images. Yean Cheng, Yongjie Zhu, Si Li 0001, Gang Pan 0001, Boxin Shi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Brain-Inspired Computing: A Systematic Survey and Future TrendsabstractBrain-inspired computing (BIC) is an emerging research field that aims to build fundamental theories, models, hardware architectures, and application systems toward more general artificial intelligence (AI) by learning from the information processing mechanisms or structures/functions of biological nervous systems. It is regarded as one of the most promising research directions for future intelligent computing in the post-Moore era. In the past few years, various new schemes in this field have sprung up to explore more general AI. These works are quite divergent in the aspects of modeling/algorithm, software tool, hardware platform, and benchmark data since BIC is an interdisciplinary field that consists of many different domains, including computational neuroscience, AI, computer science, statistical physics, material science, and microelectronics. This situation greatly impedes researchers from obtaining a clear picture and getting started in the right way. Hence, there is an urgent requirement to do a comprehensive survey in this field to help correctly recognize and analyze such bewildering methodologies. What are the key issues to enhance the development of BIC? What roles do the current mainstream technologies play in the general framework of BIC? Which techniques are truly useful in real-world applications? These questions largely remain open. To address the above issues, in this survey, we first clarify the biggest challenge of BIC: how can AI models benefit from the recent advancements in computational neuroscience? With this challenge in mind, we will focus on discussing the concept of BIC and summarize four components of BIC infrastructure development: 1) modeling/algorithm; 2) hardware platform; 3) software tool; and 4) benchmark data. For each component, we will summarize its recent progress, main challenges to resolve, and future trends. Based on these studies, we present a general framework for the real-world applications of BIC systems, which is promising to benefit both AI and brain science. Finally, we claim that it is extremely important to build a research ecology to promote prosperity continuously in this field. Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 4 |
| 2024 | Corrections to "Brain-Inspired Computing: A Systematic Survey and Future Trends"abstractPresents corrections to the paper, (Corrections to “Brain-Inspired Computing: A Systematic Survey and Future Trends”). Guoqi Li 0002, Lei Deng 0003, Huajin Tang, Gang Pan 0001, Yonghong Tian 0001, Kaushik Roy 0001, Wolfgang Maass 0001 |
Proc. IEEE | 4 |
| 2024 | CareSleepNet: A Hybrid Deep Learning Network for Automatic Sleep StagingabstractSleep staging is essential for sleep assessment and plays an important role in disease diagnosis, which refers to the classification of sleep epochs into different sleep stages. Polysomnography (PSG), consisting of many different physiological signals, e.g. electroencephalogram (EEG) and electrooculogram (EOG), is a gold standard for sleep staging. Although existing studies have achieved high performance on automatic sleep staging from PSG, there are still some limitations: 1) they focus on local features but ignore global features within each sleep epoch, and 2) they ignore cross-modality context relationship between EEG and EOG. In this paper, we propose CareSleepNet, a novel hybrid deep learning network for automatic sleep staging from PSG recordings. Specifically, we first design a multi-scale Convolutional-Transformer Epoch Encoder to encode both local salient wave features and global features within each sleep epoch. Then, we devise a Cross-Modality Context Encoder based on co-attention mechanism to model cross-modality context relationship between different modalities. Next, we use a Transformer-based Sequence Encoder to capture the sequential relationship among sleep epochs. Finally, the learned feature representations are fed into an epoch-level classifier to determine the sleep stages. We collected a private sleep dataset, SSND, and use two public datasets, Sleep-EDF-153 and ISRUC to evaluate the performance of CareSleepNet. The experiment results show that our CareSleepNet achieves the state-of-the-art performance on the three datasets. Moreover, we conduct ablation studies and attention visualizations to prove the effectiveness of each module and to analyze the influence of each modality. Jiquan Wang, Sha Zhao, Haiteng Jiang, Yangxuan Zhou, Zhenghe Yu, Shijian Li, Gang Pan 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2024 | PU-Detector: A PU Learning-based Framework for Real Money Trading Detection in MMORPGabstractMassive multiplayer online role-playing games (MMORPG) have been becoming one of the most popular and exciting online games. In recent years, a cheating phenomenon called real money trading (RMT) has arisen and damaged the fantasy world in many ways. RMT is the sale of in-game items, currency, or even characters to earn real money, breaking the balance of the game economy ecosystem and damaging the game experience. Therefore, some studies have emerged to address the problem of RMT detection. However, they cannot well handle the label uncertainty problem in practice, where there are only labeled RMT samples (positive samples) and unlabeled samples, which could either be RMT samples or normal transactions (negative samples). Meanwhile, the trading relationship between RMTers is modeled in a simple way, leading to some normal transactions being falsely classified as RMT. In this article, we propose PU-Detector, a novel framework based on PU learning (learning from positive and unlabeled data) for RMT detection, considering the fact that there are only labeled RMT samples and other unlabeled transactions. We first automatically estimate the likelihood of one transaction being RMT by developing an improved PU learning method and proposing an assessment rule. Sequentially, we use the estimated likelihood as edge weight to construct a trading graph to learn trader representation. Then, with the trader representations and basic trading features, we detect RMT samples by the improved PU learning method. PU-Detector is evaluated on a large-scale real world dataset consisting of 33,809,956 transaction logs generated by 43,217 unique players. Compared with other approaches, it achieves the state-of-the-art performance and demonstrates its advantages in detecting underlying RMT samples. Yilin Wang 0014, Sha Zhao, Runze Wu 0001, Yuhong Xu, Jianrong Tao, Tangjie Lv, Shijian Li, Zhipeng Hu, Gang Pan 0001 |
ACM Trans. Knowl. Discov. Data | 10 |
| 2024 | A Human-Machine Joint Learning Framework to Boost Endogenous BCI TrainingabstractBrain-computer interfaces (BCIs) provide a direct pathway from the brain to external devices and have demonstrated great potential for assistive and rehabilitation technologies. Endogenous BCIs based on electroencephalogram (EEG) signals, such as motor imagery (MI) BCIs, can provide some level of control. However, mastering spontaneous BCI control requires the users to generate discriminative and stable brain signal patterns by imagery, which is challenging and is usually achieved over a very long training time (weeks/months). Here, we propose a human-machine joint learning framework to boost the learning process in endogenous BCIs, by guiding the user to generate brain signals toward an optimal distribution estimated by the decoder, given the historical brain signals of the user. To this end, we first model the human-machine joint learning process in a uniform formulation. Then a human-machine joint learning framework is proposed: 1) for the human side, we model the learning process in a sequential trial-and-error scenario and propose a novel "copy/new" feedback paradigm to help shape the signal generation of the subject toward the optimal distribution and 2) for the machine side, we propose a novel adaptive learning algorithm to learn an optimal signal distribution along with the subject's learning process. Specifically, the decoder reweighs the brain signals generated by the subject to focus more on "good" samples to cope with the learning process of the subject. Online and psuedo-online BCI experiments with 18 healthy subjects demonstrated the advantages of the proposed joint learning process over coadaptive approaches in both learning efficiency and effectiveness. Lin Yao 0002, Yueming Wang 0001, Dario Farina, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Hierarchical Spiking-Based Model for Efficient Image Classification With Enhanced Feature Extraction and EncodingabstractThanks to their event-driven nature, spiking neural networks (SNNs) are surmised to be great computation-efficient models. The spiking neurons encode beneficial temporal facts and possess excessive anti-noise properties. However, the high-quality encoding of spatio-temporal complexity and also its training optimization of SNNs are restricted by means of the contemporary problem, this article proposes a novel hierarchical event-driven visual device to explore how information transmits and signifies in the retina the usage of biologically manageable mechanisms. This cognitive model is an augmented spiking-based framework consisting of the function learning capacity of convolutional neural networks (CNNs) with the cognition capability of SNNs. Furthermore, this visual device is modeled in a biological realism way with unsupervised learning rules and advanced spike firing rate encoding methods. We train and test them on some image datasets (Modified National Institute of Standards and Technology (MNIST), Canadian Institute for Advanced Research (CIFAR)10, and its noisy versions) to show that our mannequin can process greater vital data than present cognitive models. This article also proposes a novel quantization approach to make the proposed spiking-based model more efficient for neuromorphic hardware implementation. The outcomes show this joint CNN-SNN model can reap excessive focus accuracy and get more effective generalization ability. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | LSM-Based Hotspot Prediction and Hotspot-Aware Routing in NoC-Based Neuromorphic ProcessorabstractThe traffic patterns of spiking neural networks (SNNs) exhibit high variability and stochastic, leading to the emergence of elevated traffic hotspots on the network-on-chip (NoC)-based neuromorphic processors. Predicting the occurrence of hotspots remains one of the most challenging issues in NoC design. This article presents the first attempt toward traffic hotspot prediction by utilizing liquid state machine (HP-LSM). The predictor extracts essential information reflecting the current state of the NoC to predict potential routing hotspots in the subsequent time step. Furthermore, we designed the hardware architecture for HP-LSM, which incorporates leaky-integrate-and-fire (LIF) neurons with configurable biological parameters. Meanwhile, we introduce a novel hotspot-aware path-based multicast (HaPM) routing algorithm that utilizes advanced knowledge acquired from HP-LSM to guide packet routing throughout the network, aiming to improve the performance of NoC. Results indicate that the HP-LSM can forecast hotspot formation with an accuracy up to 89.36% and 90.19% for two spiking-based datasets, respectively. The hardware experiment results demonstrate a 92.03% reduction in the average execution time of zero skipping compared with nonzero skipping. Moreover, the HP-LSM exhibits a reduction of up to 79.30% in the number of neurons compared with other related SNN predictor models. The experiments reveal a reduction of 73.67% and 53.42% in the average length of the multicast path when compared with dual-path (DP) or multipath (MP) multicast routing. The HaPM demonstrates improved performance in terms of average latency and throughput compared with DP, MP, and path-based multicast (PbM) multicast routing. Ziyang Kang, Xun Xiao, Lei Wang 0011, De Ma, Gang Pan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | Augmented Proximal Policy Optimization for Safe Reinforcement LearningabstractSafe reinforcement learning considers practical scenarios that maximize the return while satisfying safety constraints. Current algorithms, which suffer from training oscillations or approximation errors, still struggle to update the policy efficiently with precise constraint satisfaction. In this article, we propose Augmented Proximal Policy Optimization (APPO), which augments the Lagrangian function of the primal constrained problem via attaching a quadratic deviation term. The constructed multiplier-penalty function dampens cost oscillation for stable convergence while being equivalent to the primal constrained problem to precisely control safety costs. APPO alternately updates the policy and the Lagrangian multiplier via solving the constructed augmented primal-dual problem, which can be easily implemented by any first-order optimizer. We apply our APPO methods in diverse safety-constrained tasks, setting a new state of the art compared with a comprehensive list of safe RL baselines. Extensive experiments verify the merits of our method in easy implementation, stable convergence, and precise cost control. Juntao Dai, Jiaming Ji, Long Yang 0004, Gang Pan 0001 |
AAAI | 5 |
| 2023 | Extracting Semantic-Dynamic Features for Long-Term Stable Brain Computer InterfaceabstractBrain-computer Interface (BCI) builds a neural signal to the motor command pathway, which is a prerequisite for the realization of neural prosthetics. However, a long-term stable BCI suffers from the neural data drift across days while retraining the BCI decoder is expensive and restricts its application scenarios. Recent solutions of neural signal recalibration treat the continuous neural signals as discrete, which is less effective in temporal feature extraction. Inspired by the observation from biologists that low-dimensional dynamics could describe high-dimensional neural signals, we model the underlying neural dynamics and propose a semantic-dynamic feature that represents the semantics and dynamics in a shared feature space facilitating the BCI recalibration. Besides, we present the joint distribution alignment instead of the common used marginal alignment strategy, dealing with the various complex changes in neural data distribution. Our recalibration approach achieves state-of-the-art performance on the real neural data of two monkeys in both classification and regression tasks. Our approach is also evaluated on a simulated dataset, which indicates its robustness in dealing with various common causes of neural signal instability. Gang Pan 0001 |
AAAI | 4 |
| 2023 | Off-Policy Proximal Policy OptimizationabstractProximal Policy Optimization (PPO) is an important reinforcement learning method, which has achieved great success in sequential decision-making problems. However, PPO faces the issue of sample inefficiency, which is due to the PPO cannot make use of off-policy data. In this paper, we propose an Off-Policy Proximal Policy Optimization method (Off-Policy PPO) that improves the sample efficiency of PPO by utilizing off-policy data. Specifically, we first propose a clipped surrogate objective function that can utilize off-policy data and avoid excessively large policy updates. Next, we theoretically clarify the stability of the optimization process of the proposed surrogate objective by demonstrating the degree of policy update distance is consistent with that in the PPO. We then describe the implementation details of the proposed Off-Policy PPO which iteratively updates policies by optimizing the proposed clipped surrogate objective. Finally, the experimental results on representative continuous control tasks validate that our method outperforms the state-of-the-art methods on most tasks. Wenjia Meng, Gang Pan 0001, Yilong Yin |
AAAI | 3 |
| 2023 | ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural NetworksabstractSpiking neural networks (SNNs) have manifested remarkable advantages in power consumption and event-driven property during the inference process. To take full advantage of low power consumption and improve the efficiency of these models further, the pruning methods have been explored to find sparse SNNs without redundancy connections after training. However, parameter redundancy still hinders the efficiency of SNNs during training. In the human brain, the rewiring process of neural networks is highly dynamic, while synaptic connections maintain relatively sparse during brain development. Inspired by this, here we propose an efficient evolutionary structure learning (ESL) framework for SNNs, named ESL-SNNs, to implement the sparse SNN training from scratch. The pruning and regeneration of synaptic connections in SNNs evolve dynamically during learning, yet keep the structural sparsity at a certain level. As a result, the ESL-SNNs can search for optimal sparse connectivity by exploring all possible parameters across time. Our experiments show that the proposed ESL-SNNs framework is able to learn SNNs with sparse structures effectively while reducing the limited accuracy. The ESL-SNNs achieve merely 0.28% accuracy loss with 10% connection density on the DVS-Cifar10 dataset. Our work presents a brand-new approach for sparse training of SNNs from scratch with biologically plausible evolutionary mechanisms, closing the gap in the expressibility between sparse training and dense training. Hence, it has great potential for SNN lightweight training and inference with low power consumption and small memory usage. Jiangrong Shen, Qi Xu 0008, Jian K. Liu, Yueming Wang 0001, Gang Pan 0001, Huajin Tang |
AAAI | 5 |
| 2023 | Loan Fraud Users Detection in Online Lending Leveraging Multiple Data ViewsabstractIn recent years, online lending platforms have been becoming attractive for micro-financing and popular in financial industries. However, such online lending platforms face a high risk of failure due to the lack of expertise on borrowers' creditworthness. Thus, risk forecasting is important to avoid economic loss. Detecting loan fraud users in advance is at the heart of risk forecasting. The purpose of fraud user (borrower) detection is to predict whether one user will fail to make required payments in the future. Detecting fraud users depend on historical loan records. However, a large proportion of users lack such information, especially for new users. In this paper, we attempt to detect loan fraud users from cross domain heterogeneous data views, including user attributes, installed app lists, app installation behaviors, and app-in logs, which compensate for the lack of historical loan records. However, it is difficult to effectively fuse the multiple heterogeneous data views. Moreover, some samples miss one or even more data views, increasing the difficulty in fusion. To address the challenges, we propose a novel end-to-end deep multiview learning approach, which encodes heterogeneous data views into homogeneous ones, generates the missing views based on the learned relationship among all the views, and then fuses all the views together to a comprehensive view for identifying fraud users. Our model is evaluated on a real-world large-scale dataset consisting of 401,978 loan records of 228,117 users from January 1, 2019, to September 30, 2019, achieving the state-of-the-art performance. Sha Zhao, Yongrui Huang, Shijian Li, Gang Pan 0001 |
AAAI | 7 |
| 2023 | Mapping Very Large Scale Spiking Neuron Network to Neuromorphic HardwareabstractNeuromorphic hardware is a multi-core computer system specifically designed to run Spiking Neuron Network (SNN) applications. As the scale of neuromorphic hardware increases, it becomes very challenging to efficiently map a large SNN to hardware. In this paper, we proposed an efficient approach to map very large scale SNN applications to neuromorphic hardware, aiming to reduce energy consumption, spike latency, and on-chip network communication congestion. The approach consists of two steps. Firstly, it solves the initial placement using the Hilbert curve, a space-filling curve with unique properties that are particularly suitable for mapping SNNs. Secondly, the Force Directed (FD) algorithm is developed to optimize the initial placement. The FD algorithm formulates the connections of clusters as tension forces, thus converts the local optimization of placement as a force analysis problem. The proposed approach is evaluated with the scale of 4 billion neurons, which is more than 200 times larger than previous research. The results show that our approach achieves state-of-the-art performance, significantly exceeding existing approaches. Ouwen Jin, Qinghui Xing, Ying Li 0001, Shuiguang Deng, Shuibing He, Gang Pan 0001 |
ASPLOS (3) | 6 |
| 2023 | DANI-Net: Uncalibrated Photometric Stereo by Differentiable Shadow Handling, Anisotropic Reflectance Modeling, and Neural Inverse RenderingabstractUncalibrated photometric stereo (UPS) is challenging due to the inherent ambiguity brought by the unknown light. Although the ambiguity is alleviated on non-Lambertian objects, the problem is still difficult to solve for more general objects with complex shapes introducing irregular shadows and general materials with complex reflectance like anisotropic reflectance. To exploit cues from shadow and reflectance to solve UPS and improve performance on general materials, we propose DANI-Net, an inverse rendering framework with differentiable shadow handling and anisotropic reflectance modeling. Unlike most previous methods that use non-differentiable shadow maps and assume isotropic material, our network benefits from cues of shadow and anisotropic reflectance through two differentiable paths. Experiments on multiple real-world datasets demonstrate our superior and robust performance. Zongrui Li 0001, Boxin Shi, Gang Pan 0001, Xudong Jiang 0001 |
CVPR | 4 |
| 2023 | Constructing Deep Spiking Neural Networks from Artificial Neural Networks with Knowledge DistillationabstractSpiking neural networks (SNNs) are well-known as brain-inspired models with high computing efficiency, due to a key component that they utilize spikes as information units, close to the biological neural systems. Although spiking based models are energy efficient by taking advantage of discrete spike signals, their performance is limited by current network structures and their training methods. As discrete signals, typical SNNs cannot apply the gradient descent rules directly into parameter adjustment as artificial neural networks (ANNs). Aiming at this limitation, here we propose a novel method of constructing deep SNN models with knowledge distillation (KD) that uses ANN as the teacher model and SNN as the student model. Through the ANN-SNN joint training algorithm, the student SNN model can learn rich feature information from the teacher ANN model through the KD method, yet it avoids training SNN from scratch when communicating with non-differentiable spikes. Our method can not only build a more efficient deep spiking structure feasibly and reasonably but use few time steps to train the whole model compared to direct training or ANN to SNN methods. More importantly, it has a superb ability of noise immunity for various types of artificial noises and natural signals. The proposed novel method provides efficient ways to improve the performance of SNN through constructing deeper structures in a high-throughput fashion, with potential usage for light and efficient brain-inspired computing of practical scenarios. Qi Xu 0008, Jiangrong Shen, Jian K. Liu, Huajin Tang, Gang Pan 0001 |
CVPR | 6 |
| 2023 | Rethinking Visual Reconstruction: Experience-Based Content Completion Guided by Visual CuesabstractDecoding seen images from brain activities has been an absorbing field. However, the reconstructed images still suffer from low quality with existing studies. This can be because our visual system is not like a camera that ''remembers'' every pixel. Instead, only part of the information can be perceived with our selective attention, and the brain ''guesses'' the rest to form what we think we see. Most existing approaches ignored the brain completion mechanism. In this work, we propose to reconstruct seen images with both the visual perception and the brain completion process, and design a simple, yet effective visual decoding framework to achieve this goal. Specifically, we first construct a shared discrete representation space for both brain signals and images. Then, a novel self-supervised token-to-token inpainting network is designed to implement visual content completion by building context and prior knowledge about the visual objects from the discrete latent space. Our approach improved the quality of visual reconstruction significantly and achieved state-of-the-art. Jiaxuan Chen 0007, Gang Pan 0001 |
ICML | 3 |
| 2023 | Controlling Type Confounding in Ad Hoc Teamwork with Instance-wise Teammate Feedback RectificationabstractAd hoc teamwork requires an agent to cooperate with unknown teammates without prior coordination. Many works propose to abstract teammate instances into high-level representation of types and then pre-train the best response for each type. However, most of them do not consider the distribution of teammate instances within a type. This could expose the agent to the hidden risk of type confounding. In the worst case, the best response for an abstract teammate type could be the worst response for all specific instances of that type. This work addresses the issue from the lens of causal inference. We first theoretically demonstrate that this phenomenon is due to the spurious correlation brought by uncontrolled teammate distribution. Then, we propose our solution, CTCAT, which disentangles such correlation through an instance-wise teammate feedback rectification. This operation reweights the interaction of teammate instances within a shared type to reduce the influence of type confounding. The effect of CTCAT is evaluated in multiple domains, including classic ad hoc teamwork tasks and real-world scenarios. Results show that CTCAT is robust to the influence of type confounding, a practical issue that directly hazards the robustness of our trained agents but was unnoticed in previous works. Dong Xing, Pengjie Gu, Xinrun Wang, Shanqi Liu, Longtao Zheng, Bo An 0001, Gang Pan 0001 |
ICML | 8 |
| 2023 | Spiking Reinforcement Learning with Memory Ability for Mapless NavigationabstractOur study focuses on mapless navigation in robotics, which involves navigating without an established obstacle map of the environment. Spiking Neural Networks (SNNs) have recently been applied to this task using Deep Reinforcement Learning (DRL), but face challenges in dynamic and partially observable environments, as well as inaccuracies in transmitted data. To overcome these issues, we propose a Multi-Critic DDPG with Spiking Memory (MC-DDPGSM) framework. Our approach introduces a spiking Gate Recurrent Unit layer (Spiking-GRU) to provide memory function and evaluates the state-action value with multi-critic networks. The experimental results demonstrate that our method achieves better performance (success rate, navigation distance, navigation time spent, and power consumption) in complex navigation tasks compared to the state-of-the-art approaches. Furthermore, our model can be transferred to unseen environments without the need for fine-tuning. Mengwen Yuan, Chaofei Hong, Gang Pan 0001, Huajin Tang |
IROS | 5 |
| 2023 | Alleviating the Semantic Gap for Generalized fMRI-to-Image ReconstructionabstractAlthough existing fMRI-to-image reconstruction methods could predict high-quality images, they do not explicitly consider the semantic gap between training and testing data, resulting in reconstruction with unstable and uncertain semantics. This paper addresses the problem of generalized fMRI-to-image reconstruction by explicitly alleviates the semantic gap. Specifically, we leverage the pre-trained CLIP model to map the training data to a compact feature representation, which essentially extends the sparse semantics of training data to dense ones, thus alleviating the semantic gap of the instances nearby known concepts (i.e., inside the training super-classes). Inspired by the robust low-level representation in fMRI data, which could help alleviate the semantic gap for instances that far from the known concepts (i.e., outside the training super-classes), we leverage structural information as a general cue to guide image reconstruction. Further, we quantify the semantic uncertainty based on probability density estimation and achieve Generalized fMRI-to-image reconstruction by adaptively integrating Expanded Semantics and Structural information (GESS) within a diffusion process. Experimental results demonstrate that the proposed GESS model outperforms state-of-the-art methods, and we propose a generalized scenario split strategy to evaluate the advantage of GESS in closing the semantic gap. Gang Pan 0001 |
NeurIPS | 3 |
| 2023 | Enhancing Adaptive History Reserving by Spiking Convolutional Block Attention Module in Recurrent Neural NetworksabstractSpiking neural networks (SNNs) serve as one type of efficient model to process spatio-temporal patterns in time series, such as the Address-Event Representation data collected from Dynamic Vision Sensor (DVS). Although convolutional SNNs have achieved remarkable performance on these AER datasets, benefiting from the predominant spatial feature extraction ability of convolutional structure, they ignore temporal features related to sequential time points. In this paper, we develop a recurrent spiking neural network (RSNN) model embedded with an advanced spiking convolutional block attention module (SCBAM) component to combine both spatial and temporal features of spatio-temporal patterns. It invokes the history information in spatial and temporal channels adaptively through SCBAM, which brings the advantages of efficient memory calling and history redundancy elimination. The performance of our model was evaluated in DVS128-Gesture dataset and other time-series datasets. The experimental results show that the proposed SRNN-SCBAM model makes better use of the history information in spatial and temporal dimensions with less memory space, and achieves higher accuracy compared to other models. Qi Xu 0008, Yuyuan Gao, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001 |
NeurIPS | 7 |
| 2023 | Pyramid-VAE-GAN: Transferring hierarchical latent variables for image inpaintingabstractSignificant progress has been made in image inpainting methods in recent years. However, they are incapable of producing inpainting results with reasonable structures, rich detail, and sharpness at the same time. In this paper, we propose the Pyramid-VAE-GAN network for image inpainting to address this limitation. Our network is built on a variational autoencoder (VAE) backbone that encodes high-level latent variables to represent complicated high-dimensional prior distributions of images. The prior assists in reconstructing reasonable structures when inpainting. We also adopt a pyramid structure in our model to maintain rich detail in low-level latent variables. To avoid the usual incompatibility of requiring both reasonable structures and rich detail, we propose a novel cross-layer latent variable transfer module. This transfers information about long-range structures contained in high-level latent variables to low-level latent variables representing more detailed information. We further use adversarial training to select the most reasonable results and to improve the sharpness of the images. Extensive experimental results on multiple datasets demonstrate the superiority of our method. Our code is available at https://github.com/thy960112/Pyramid-VAE-GAN . Hui-yuan Tian, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Comput. Vis. Media | 5 |
| 2023 | RM-FSP: Regret minimization optimizes neural fictitious self-play
Li Zhang 0045, Shijian Li, Xili Chen, Gang Pan 0001 |
Neurocomputing | 5 |
| 2023 | Fast-SNN: Fast Spiking Neural Network by Converting Quantized ANNabstractSpiking neural networks (SNNs) have shown advantages in computation and energy efficiency over traditional artificial neural networks (ANNs) thanks to their event-driven representations. SNNs also replace weight multiplications in ANNs with additions, which are more energy-efficient and less computationally intensive. However, it remains a challenge to train deep SNNs due to the discrete spiking function. A popular approach to circumvent this challenge is ANN-to-SNN conversion. However, due to the quantization error and accumulating error, it often requires lots of time steps (high inference latency) to achieve high performance, which negates SNN's advantages. To this end, this paper proposes Fast-SNN that achieves high performance with low latency. We demonstrate the equivalent mapping between temporal quantization in SNNs and spatial quantization in ANNs, based on which the minimization of the quantization error is transferred to quantized ANN training. With the minimization of the quantization error, we show that the sequential error is the primary cause of the accumulating error, which is addressed by introducing a signed IF neuron model and a layer-wise fine-tuning mechanism. Our method achieves state-of-the-art performance and low latency on various computer vision tasks, including image classification, object detection, and semantic segmentation. Codes are available at: https://github.com/yangfan-hu/Fast-SNN. Yangfan Hu, Xudong Jiang 0001, Gang Pan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Spatio-temporal analysis of urban crime leveraging multisource crowdsensed data
Binbin Zhou 0005, Longbiao Chen, Sha Zhao, Fangxun Zhou, Shijian Li, Gang Pan 0001 |
Pers. Ubiquitous Comput. | 6 |
| 2023 | DILI: A Distribution-Driven Learned IndexabstractTargeting in-memory one-dimensional search keys, we propose a novel DIstribution-driven Learned Index tree ( DILI ), where a concise and computation-efficient linear regression model is used for each node. An internal node's key range is equally divided by its child nodes such that a key search enjoys perfect model prediction accuracy to find the relevant leaf node. A leaf node uses machine learning models to generate searchable data layout and thus accurately predicts the data record position for a key. To construct DILI, we first build a bottom-up tree with linear regression models according to global and local key distributions. Using the bottom-up tree, we build DILI in a top-down manner, individualizing the fanouts for internal nodes according to local distributions. DILI strikes a good balance between the number of leaf nodes and the height of the tree, two critical factors of key search time. Moreover, we design flexible algorithms for DILI to efficiently insert and delete keys and automatically adjust the tree structure when necessary. Extensive experimental results show that DILI outperforms the state-of-the-art alternatives on different kinds of workloads. Pengfei Li 0005, Hua Lu 0001, Bolin Ding, Long Yang 0004, Gang Pan 0001 |
Proc. VLDB Endow. | 6 |
| 2023 | Explainable AI for Cheating Detection and Churn Prediction in Online GamesabstractOnline gaming is a multibillion dollar industry that entertains a large, global population. Empowering online games with AI has made a great success, however, ignores the explainability of black-box model makes AI less responsible and hinders its further development. In this article, we introduce and discuss the audience and the concept of XAI (eXplainable AI) in online games. We propose a GXAI workflow, which combines the strong expressiveness of multiview data sources and the clear transparency of multiview black-box models. We present four specific classifiers and explainers in the character portrait view, the behavior sequence view, the client image view, and the social graph view. Experiments conducted on real-world datasets for game cheating detection and player churn prediction show the accuracy of classification and the rationality of explanation. We also discover and present numerous interesting and valuable findings from the individual, local, and global explanations. We implement and deploy three practical applications, including evidence and reason generation, model debugging and testing, and model compression and comparison in NetEase Games and have received quite positive reviews from user studies. More future work is in progress since this is the first work that introduces XAI in online games. Jianrong Tao, Runze Wu 0001, Tangjie Lyu, Changjie Fan, Zhipeng Hu, Sha Zhao, Gang Pan 0001 |
IEEE Trans. Games | 10 |
| 2023 | Unsupervised Domain Adaptation for Crime Risk Prediction Across CitiesabstractCrime risk prediction is crucial for city safety and residents’ life quality. However, without labeled data, it is challenging to predict crime risk in cities. Due to municipal regulations and maintenance costs, it is not trivial for many cities to collect high-quality labeled crime data. In particular, some cities have lots of labeled data while others may have few. It has been possible to develop a crime prediction model for a city without labeled crime data by learning knowledge from a city with abundant data. Nevertheless, the inconsistency of relevant context data between cities exacerbates the difficulty of this prediction task. To this end, this article proposes an effective unsupervised domain adaptation model (UDAC) for crime risk prediction across cities while addressing the contexts’ inconsistency issue. More specifically, we first identify several similar source city grids for each target city grid. Based on these source city grids, we then construct auxiliary contexts for the target city, to make contexts consistent between the two cities. A dense convolutional network with unsupervised domain adaptation is designed to learn high-level representations for accurate crime risk prediction and simultaneously learn domain-invariant features for domain adaptation. The effectiveness of our model is verified through extensive experiments using three real-world datasets. Binbin Zhou 0005, Longbiao Chen, Sha Zhao, Shijian Li, Zengwei Zheng, Gang Pan 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2023 | Spiking Deep Residual NetworksabstractSpiking neural networks (SNNs) have received significant attention for their biological plausibility. SNNs theoretically have at least the same computational power as traditional artificial neural networks (ANNs). They possess the potential of achieving energy-efficient machine intelligence while keeping comparable performance to ANNs. However, it is still a big challenge to train a very deep SNN. In this brief, we propose an efficient approach to build deep SNNs. Residual network (ResNet) is considered a state-of-the-art and fundamental model among convolutional neural networks (CNNs). We employ the idea of converting a trained ResNet to a network of spiking neurons named spiking ResNet (S-ResNet). We propose a residual conversion model that appropriately scales continuous-valued activations in ANNs to match the firing rates in SNNs and a compensation mechanism to reduce the error caused by discretization. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on CIFAR-10, CIFAR-100, and ImageNet 2012 with low latency. This work is the first time to build an asynchronous SNN deeper than 100 layers, with comparable performance to its original ANN. Yangfan Hu, Huajin Tang, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | A Thompson Sampling Algorithm With Logarithmic Regret for Unimodal Gaussian BanditabstractIn this article, we propose a Thompson sampling algorithm with Gaussian prior for unimodal bandit under Gaussian reward setting, where the expected reward is unimodal over the partially ordered arms. To exploit the unimodal structure better, at each step, instead of exploration from the entire decision space, the proposed algorithm makes decisions according to posterior distribution only in the arm's neighborhood with the highest empirical mean estimate. We theoretically prove that the asymptotic regret of our algorithm reaches O(logT) , i.e., it shares the same regret order with asymptotic optimal algorithms, which is comparable to extensive existing state-of-the-art unimodal multiarm bandit (U-MAB) algorithms. Finally, we use extensive experiments to demonstrate the effectiveness of the proposed algorithm on both synthetic datasets and real-world applications. Long Yang 0004, Zhao Li 0007, Zehong Hu, Shasha Ruan, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Policy Optimization with Stochastic Mirror DescentabstractImproving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes VRMPO algorithm: a sample efficient policy gradient method with stochastic mirror descent. In VRMPO, a novel variance-reduced policy gradient estimator is presented to improve sample efficiency. We prove that the proposed VRMPO needs only O(ε−3) sample trajectories to achieve an ε-approximate first-order stationary point, which matches the best sample complexity for policy optimization. Extensive empirical results demonstrate that VRMP outperforms the state-of-the-art policy gradient methods in various settings. Long Yang 0004, Yu Zhang 0009, Gang Zheng 0005, Pengfei Li 0005, Jianhang Huang, Gang Pan 0001 |
AAAI | 7 |
| 2022 | Event-Based Multimodal Spiking Neural Network with Attention MechanismabstractHuman brain can effectively integrate visual and auditory information. Dynamic Vision Sensor (DVS) and Dynamic Audio Sensor (DAS) are event-based sensors imitating the mechanism of human retina and cochlea. Since the sensors record the visual and auditory input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on unimodality, however, audiovisual multimodal SNNs are still limited. In this paper, we propose an end-to-end event-based multimodal spiking neural network. The network consists of visual and auditory unimodal subnetworks and a novel attention-based cross-modal subnetwork for fusion. The attention mechanism measures the significance of each modality and allocates the weights to two modalities. We evaluate our proposed multimodal network on an event-based audiovisual joint dataset (MNIST-DVS and N-TIDIGITS datasets). Experimental results show the performance improvement of this multimodal network and the effectiveness of our proposed attention mechanism. Qianhui Liu, Dong Xing, Lang Feng 0002, Huajin Tang, Gang Pan 0001 |
ICASSP | 5 |
| 2022 | T-Detector: A Trajectory based Pre-trained Model for Game Bot Detection in MMORPGsabstractGame bots are programmed to automatically play games and illegally obtain profit, seriously affecting game experience of honest players and breaking the balance of game ecosystem. Therefore, bot detection needs to be addressed urgently, especially for MMORPGs, one of the most rapidly expanding genres of games. There have been some studies for bot detection, but the features they used are dependent on specific games and the methods cannot be generalized to other games. In this paper, we propose a trajectory based pre-trained model for game bot detection from game character trajectories and mouse trajectories, named T-Detector, which is independent to specific games and can be generalized to others. More specifically, we propose a pretrain method of LocationTime2Vec to learn representations of trajectories from huge unlabeled samples, which deeply embed spatial and temporal information hidden in trajectories. Moreover, we extract universal features based on behavioral differences in movement trajectories between human players and bots. We design an Angle Pretrain to extract features of turning angle, and propose an attention pooling module to extract features of moving speed and distance. Such features are not dependent on any specific game, enabling T-Detector to be generalized to many MMORPGs. Evaluated by two large-scale real-world datasets of 143,938 samples from two MMORPGs, T-Detector achieves the state-of-the-art performance in bot detection, and demonstrates powerful generalization ability. Sha Zhao, Junwei Fang, Runze Wu 0001, Jianrong Tao, Shijian Li, Gang Pan 0001 |
ICDE | 7 |
| 2022 | Multi-Level Firing with Spiking DS-ResNet: Enabling Better and Deeper Directly-Trained Spiking Neural NetworksabstractSpiking neural networks (SNNs) are bio-inspired neural networks with asynchronous discrete and sparse characteristics, which have increasingly manifested their superiority in low energy consumption. Recent research is devoted to utilizing spatio-temporal information to directly train SNNs by backpropagation. However, the binary and non-differentiable properties of spike activities force directly trained SNNs to suffer from serious gradient vanishing and network degradation, which greatly limits the performance of directly trained SNNs and prevents them from going deeper. In this paper, we propose a multi-level firing (MLF) method based on the existing spatio-temporal back propagation (STBP) method, and spiking dormant-suppressed residual network (spiking DS-ResNet). MLF enables more efficient gradient propagation and the incremental expression ability of the neurons. Spiking DS-ResNet can efficiently perform identity mapping of discrete spikes, as well as provide a more suitable connection for gradient propagation in deep SNNs. With the proposed method, our model achieves superior performances on a non-neuromorphic dataset and two neuromorphic datasets with much fewer trainable parameters and demonstrates the great ability to combat the gradient vanishing and degradation problem in deep SNNs. Lang Feng 0002, Qianhui Liu, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 5 |
| 2022 | TinyLight: Adaptive Traffic Signal Control on Devices with Extremely Limited ResourcesabstractRecent advances in deep reinforcement learning (DRL) have largely promoted the performance of adaptive traffic signal control (ATSC). Nevertheless, regarding the implementation, most works are cumbersome in terms of storage and computation. This hinders their deployment on scenarios where resources are limited. In this work, we propose TinyLight, the first DRL-based ATSC model that is designed for devices with extremely limited resources. TinyLight first constructs a super-graph to associate a rich set of candidate features with a group of light-weighted network blocks. Then, to diminish the model's resource consumption, we ablate edges in the super-graph automatically with a novel entropy-minimized objective function. This enables TinyLight to work on a standalone microcontroller with merely 2KB RAM and 32KB ROM. We evaluate TinyLight on multiple road networks with real-world traffic demands. Experiments show that even with extremely limited resources, TinyLight still achieves competitive performance. The source code and appendix of this work can be found at https://bit.ly/38hH8t8. Dong Xing, Qianhui Liu, Gang Pan 0001 |
IJCAI | 4 |
| 2022 | Rapid Earthquake Magnitude Estimation Using Deep LearningabstractEarthquake magnitude estimation is one of the critical parts of earthquake early warning systems. It uses the first few seconds of a waveform recorded by an earthquake detection station, which is required to be rapid and accurate. In this paper, we propose a novel framework to estimate magnitude integrated with deep learning, consisting of feature stage and regression stage. In the feature stage, we extract temporal & spatial features by deep learning methods, and combine them with hand-crafted features embedded expert domain knowledge. Then, each earthquake can be represented by a hybrid feature. Therefore, magnitude estimation can be modeled as a regression problem to solve. Our framework is evaluated on 5,503 earthquake records collected in Sichuan province, China. It is found that, learning the temporal & spatial features by deep neural networks is critical for magnitude estimation. The results demonstrate the state-of-the-art performance, compared with other approaches. Sha Zhao, Yizhi Xu, Zhiling Luo, Jin Dong Song, Shijian Li, Gang Pan 0001 |
IJCNN | 7 |
| 2022 | Online Neural Sequence Detection with Hierarchical Dirichlet Point ProcessabstractNeural sequence detection plays a vital role in neuroscience research. Recent impressive works utilize convolutive nonnegative matrix factorization and Neyman-Scott process to solve this problem. However, they still face two limitations. Firstly, they accommodate the entire dataset into memory and perform iterative updates of multiple passes, which can be inefficient when the dataset is large or grows frequently. Secondly, they rely on the prior knowledge of the number of sequence types, which can be impractical with data when the future situation is unknown. To tackle these limitations, we propose a hierarchical Dirichlet point process model for efficient neural sequence detection. Instead of computing the entire data, our model can sequentially detect sequences in an online unsupervised manner with Particle filters. Besides, the Dirichlet prior enables our model to automatically introduce new sequence types on the fly as needed, thus avoiding specifying the number of types in advance. We manifest these advantages on synthetic data and neural recordings from songbird higher vocal center and rodent hippocampus. Gang Pan 0001 |
NeurIPS | 3 |
| 2022 | Constrained Update Projection Approach to Safe Policy OptimizationabstractSafe reinforcement learning (RL) studies problems where an intelligent agent has to not only maximize reward but also avoid exploring unsafe areas. In this study, we propose CUP, a novel policy optimization method based on Constrained Update Projection framework that enjoys rigorous safety guarantee. Central to our CUP development is the newly proposed surrogate functions along with the performance bound. Compared to previous safe reinforcement learning meth- ods, CUP enjoys the benefits of 1) CUP generalizes the surrogate functions to generalized advantage estimator (GAE), leading to strong empirical performance. 2) CUP unifies performance bounds, providing a better understanding and in- terpretability for some existing algorithms; 3) CUP provides a non-convex im- plementation via only first-order optimizers, which does not require any strong approximation on the convexity of the objectives. To validate our CUP method, we compared CUP against a comprehensive list of safe RL baselines on a wide range of tasks. Experiments show the effectiveness of CUP both in terms of reward and safety constraint satisfaction. We have opened the source code of CUP at https://github.com/zmsn-2077/CUP-safe-rl. Long Yang 0004, Jiaming Ji, Juntao Dai, Linrui Zhang, Binbin Zhou 0005, Pengfei Li 0005, Yaodong Yang 0001, Gang Pan 0001 |
NeurIPS | 8 |
| 2022 | Tracking Functional Changes in Nonstationary Signals with Evolutionary Ensemble Bayesian Model for Robust Neural DecodingabstractNeural signals are typical nonstationary data where the functional mapping between neural activities and the intentions (such as the velocity of movements) can occasionally change. Existing studies mostly use a fixed neural decoder, thus suffering from an unstable performance given neural functional changes. We propose a novel evolutionary ensemble framework (EvoEnsemble) to dynamically cope with changes in neural signals by evolving the decoder model accordingly. EvoEnsemble integrates evolutionary computation algorithms in a Bayesian framework where the fitness of models can be sequentially computed with their likelihoods according to the incoming data at each time slot, which enables online tracking of time-varying functions. Two strategies of evolve-at-changes and history-model-archive are designed to further improve efficiency and stability. Experiments with simulations and neural signals demonstrate that EvoEnsemble can track the changes in functions effectively thus improving the accuracy and robustness of neural decoding. The improvement is most significant in neural signals with functional changes. Xinyun Zhu, Gang Pan 0001, Yueming Wang 0001 |
NeurIPS | 3 |
| 2022 | Answering medical questions in Chinese using automatically mined knowledge and deep neural networks: an end-to-end solutionabstractBACKGROUND: Medical information has rapidly increased on the internet and has become one of the main targets of search engine use. However, medical information on the internet is subject to the problems of quality and accessibility, so ordinary users are unable to obtain answers to their medical questions conveniently. As a solution, researchers build medical question answering (QA) systems. However, research on medical QA in the Chinese language lags behind work on English-based systems. This lag is mainly due to the difficulty of constructing a high-quality knowledge base and the underutilization of medical corpora in the Chinese language. RESULTS: This study developed an end-to-end solution to implement a medical QA system for the Chinese language with low cost and time. First, we created a high-quality medical knowledge graph from hospital data (electronic health/medical records) in a nearly automatic manner that trained a supervised model based on data labeled using bootstrapping techniques. Then, we designed a QA system based on a memory-based neural network and attention mechanism. Finally, we trained the system to generate answers from the knowledge base and a QA corpus on the internet. CONCLUSIONS: Bootstrapping and deep neural network techniques can construct a knowledge graph from electronic health/medical records with satisfactory precision and coverage. Our proposed context bridge mechanisms perform training with a variety of language features. Our QA system can achieve state-of-the-art quality in answering medical questions with constrained topics. As we evaluated, complex Chinese language processing techniques, such as segmentation and parsing, were not necessary for practice and complex architectures were not necessary to build the QA system. Lastly, we created an application using our method for internet QA usage. Li Zhang 0045, Shijian Li, Tianyi Liao, Gang Pan 0001 |
BMC Bioinform. | 5 |
| 2022 | DeepOffense: a recurrent network based approach for crime prediction
Fangxun Zhou, Binbin Zhou 0005, Sha Zhao, Gang Pan 0001 |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2022 | Dynamic road crime risk prediction with urban open data
Binbin Zhou 0005, Longbiao Chen, Fangxun Zhou, Shijian Li, Sha Zhao, Gang Pan 0001 |
Frontiers Comput. Sci. | 6 |
| 2022 | Subdomain contraction in deep networks for robust representation learning
Zhentao Pan, Gang Pan 0001, Yueming Wang 0001 |
Neurocomputing | 3 |
| 2022 | Training Deep Convolutional Spiking Neural Networks With Spike Probabilistic Global PoolingabstractRecent work on spiking neural networks (SNNs) has focused on achieving deep architectures. They commonly use backpropagation (BP) to train SNNs directly, which allows SNNs to go deeper and achieve higher performance. However, the BP training procedure is computing intensive and complicated by many trainable parameters. Inspired by global pooling in convolutional neural networks (CNNs), we present the spike probabilistic global pooling (SPGP) method based on a probability function for training deep convolutional SNNs. It aims to remove the difficulty of too many trainable parameters brought by multiple layers in the training process, which can reduce the risk of overfitting and get better performance for deep SNNs (DSNNs). We use the discrete leaky-integrate-fire model and the spatiotemporal BP algorithm for training DSNNs directly. As a result, our model trained with the SPGP method achieves competitive performance compared to the existing DSNNs on image and neuromorphic data sets while minimizing the number of trainable parameters. In addition, the proposed SPGP method shows its effectiveness in performance improvement, convergence, and generalization ability. Shuang Lian, Qianhui Liu, Rui Yan 0005, Gang Pan 0001, Huajin Tang |
Neural Comput. | 4 |
| 2022 | Understanding Smartphone Users From Installed App Lists Using Boolean Matrix FactorizationabstractSmartphones are changing humans' lifestyles. Mobile applications (apps) on smartphones serve as entries for users to access a wide range of services in our daily lives. The apps installed on one's smartphone convey lots of personal information, such as demographics, interests, and needs. This provides a new lens to understand smartphone users. However, it is difficult to compactly characterize a user with his/her installed app list. In this article, a user representation framework is proposed, where we model the underlying relations between apps and users with Boolean matrix factorization (BMF). It builds a compact user subspace by discovering basic components from installed app lists. Each basic component encapsulates a semantic interpretation of a series of special-purpose apps, which is a reflection of user needs and interests. Each user is represented by a linear combination of the semantic basic components. With this user representation framework, we use supervised and unsupervised learning methods to understand users, including mining user attributes, discovering user groups, and labeling semantic tags to users. Extensive experiments were conducted on three data subsets from a large-scale real-world dataset for evaluation, each consisting of installed app lists from over 10 000 users. The results demonstrated the effectiveness of our user representation framework. Sha Zhao, Gang Pan 0001, Jianrong Tao, Zhiling Luo, Shijian Li, Zhaohui Wu 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Jointly Optimizing Expressional and Residual Models for 3D Facial Expression RemovalabstractThis article proposes a facial expression removal method to recover a 3D neutral face from a single 3D expressional or non-neutral face. We treat a 3D non-neutral face as the sum of its neutral one and the residual. This can be satisfied if the correspondence between 3D vertices of expressional faces and those of neutral faces is established. We propose a non-rigid deformation method to establish the correspondence between 3D faces. Then, according to algebra inequality, the minimization of a neutral face model can be replaced by the minimization of its upper bound, i.e., the errors of an expressional face model and a residual model. Thus, we co-optimize the representation errors of the latter two models and build the relationship between the representation coefficients of the two models. Given an expressional face as the input, its corresponding neutral face can be inferred by the associative representation parameters in these two models. In the testing stage, we use an iterative joint fitting scheme to obtain a more accurate recovery. Extensive experiments are conducted to evaluate our method. The results show that our method obtains considerably better performance than existing methods in terms of average root mean square errors and recognition rates, and also better visual effects. Yueming Wang 0001, Zhenfang Hu, Zhaohui Wu 0001, Gang Pan 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2022 | An Off-Policy Trust Region Policy Optimization Method With Monotonic Improvement Guarantee for Deep Reinforcement LearningabstractIn deep reinforcement learning, off-policy data help reduce on-policy interaction with the environment, and the trust region policy optimization (TRPO) method is efficient to stabilize the policy optimization procedure. In this article, we propose an off-policy TRPO method, off-policy TRPO, which exploits both on- and off-policy data and guarantees the monotonic improvement of policies. A surrogate objective function is developed to use both on- and off-policy data and keep the monotonic improvement of policies. We then optimize this surrogate objective function by approximately solving a constrained optimization problem under arbitrary parameterization and finite samples. We conduct experiments on representative continuous control tasks from OpenAI Gym and MuJoCo. The results show that the proposed off-policy TRPO achieves better performance in the majority of continuous control tasks compared with other trust region policy-based methods using off-policy data. Wenjia Meng, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Robust Transcoding Sensory Information With Neural SpikesabstractNeural coding, including encoding and decoding, is one of the key problems in neuroscience for understanding how the brain uses neural signals to relate sensory perception and motor behaviors with neural systems. However, most of the existed studies only aim at dealing with the continuous signal of neural systems, while lacking a unique feature of biological neurons, termed spike, which is the fundamental information unit for neural computation as well as a building block for brain-machine interface. Aiming at these limitations, we propose a transcoding framework to encode multi-modal sensory information into neural spikes and then reconstruct stimuli from spikes. Sensory information can be compressed into 10% in terms of neural spikes, yet re-extract 100% of information by reconstruction. Our framework can not only feasibly and accurately reconstruct dynamical visual and auditory scenes, but also rebuild the stimulus patterns from functional magnetic resonance imaging (fMRI) brain activities. More importantly, it has a superb ability of noise immunity for various types of artificial noises and background signals. The proposed framework provides efficient ways to perform multimodal feature representation and reconstruction in a high-throughput fashion, with potential usage for efficient neuromorphic computing in a noisy environment. Qi Xu 0008, Jiangrong Shen, Xuming Ran, Huajin Tang, Gang Pan 0001, Jian K. Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Sample Complexity of Policy Gradient Finding Second-Order Stationary PointsabstractThe policy-based reinforcement learning (RL) can be considered as maximization of its objective. However, due to the inherent non-concavity of its objective, the policy gradient method to a first-order stationary point (FOSP) cannot guar- antee a maximal point. A FOSP can be a minimal or even a saddle point, which is undesirable for RL. It has be found that if all the saddle points are strict, all the second-order station- ary points (SOSP) are exactly equivalent to local maxima. Instead of FOSP, we consider SOSP as the convergence criteria to characterize the sample complexity of policy gradient. Our result shows that policy gradient converges to an (ε, √εχ)-SOSP with probability at least 1 − O(δ) after the total cost of O(ε−9/2)sinificantly improves the state of the art cost O(ε−9).Our analysis is based on the key idea that decomposes the parameter space Rp into three non-intersected regions: non-stationary point region, saddle point region, and local optimal region, then making a local improvement of the objective of RL in each region. This technique can be potentially generalized to extensive policy gradient methods. For the complete proof, please refer to https://arxiv.org/pdf/2012.01491.pdf. Long Yang 0004, Gang Pan 0001 |
AAAI | 3 |
| 2021 | On Convergence of Gradient Expected Sarsa(λ)abstractWe study the convergence of Expected Sarsa(λ) with function approximation. We show that with off-line es- timate (multi-step bootstrapping) to ExpectedSarsa(λ) is unstable for off-policy learning. Furthermore, based on convex-concave saddle-point framework, we propose a con- vergent Gradient Expected Sarsa(λ) (GES(λ)) algorithm. The theoretical analysis shows that the proposed GES(λ) converges to the optimal solution at a linear convergence rate under true gradient setting. Furthermore, we develop a Lyapunov function technique to investigate how the step- size influences finite-time performance of GES(λ). Addition- ally, such a technique of Lyapunov function can be poten- tially generalized to other gradient temporal difference algo- rithms. Finally, our experiments verify the effectiveness of our GES(λ). For the details of proof, please refer to https: //arxiv.org/pdf/2012.07199.pdf. Long Yang 0004, Gang Zheng 0005, Yu Zhang 0009, Pengfei Li 0005, Gang Pan 0001 |
AAAI | 6 |
| 2021 | Indoor Lighting Estimation Using an Event CameraabstractImage-based methods for indoor lighting estimation suffer from the problem of intensity-distance ambiguity. This paper introduces a novel setup to help alleviate the ambiguity based on the event camera. We further demonstrate that estimating the distance of a light source becomes a well-posed problem under this setup, based on which an optimization-based method and a learning-based method are proposed. Our experimental results validate that our approaches not only achieve superior performance for indoor lighting estimation (especially for the close light) but also significantly alleviate the intensity-distance ambiguity. Peisong Niu, Huajin Tang, Gang Pan 0001 |
CVPR | 5 |
| 2021 | Event-based Action Recognition Using Motion Information and Spiking Neural NetworksabstractEvent-based cameras have attracted increasing attention due to their advantages of biologically inspired paradigm and low power consumption. Since event-based cameras record the visual input as asynchronous discrete events, they are inherently suitable to cooperate with the spiking neural network (SNN). Existing works of SNNs for processing events mainly focus on the task of object recognition. However, events from the event-based camera are triggered by dynamic changes, which makes it an ideal choice to capture actions in the visual scene. Inspired by the dorsal stream in visual cortex, we propose a hierarchical SNN architecture for event-based action recognition using motion information. Motion features are extracted and utilized from events to local and finally to global perception for action recognition. To the best of the authors’ knowledge, it is the first attempt of SNN to apply motion information to event-based action recognition. We evaluate our proposed SNN on three event-based action recognition datasets, including our newly published DailyAction-DVS dataset comprising 12 actions collected under diverse recording conditions. Extensive experimental results show the effectiveness of motion information and our proposed SNN architecture for event-based action recognition. Qianhui Liu, Dong Xing, Huajin Tang, De Ma, Gang Pan 0001 |
IJCAI | 5 |
| 2021 | Learning with Generated Teammates to Achieve Type-Free Ad-Hoc TeamworkabstractIn ad-hoc teamwork, an agent is required to cooperate with unknown teammates without prior coordination. To swiftly adapt to an unknown teammate, most works adopt a type-based approach, which pre-trains the agent with a set of pre-prepared teammate types, then associates the unknown teammate with a particular type. Typically, these types are collected manually. This hampers previous works by both the availability and diversity of types they manage to obtain. To eliminate these limitations, this work addresses to achieve ad-hoc teamwork in a type-free approach. Specifically, we propose the model of Entropy-regularized Deep Recurrent Q-Network (EDRQN) to generate teammates automatically, meanwhile utilize them to pre-train our agent. These teammates are obtained from scratch and are designed to perform the task with various behaviors, therefore their availability and diversity are both ensured. We evaluate our model on several benchmark domains of ad-hoc teamwork. The result shows that even if our model has no access to any pre-prepared teammate types, it still achieves significant performance. Dong Xing, Qianhui Liu, Gang Pan 0001 |
IJCAI | 4 |
| 2021 | A Monte Carlo Neural Fictitious Self-Play approach to approximate Nash Equilibrium in imperfect-information dynamic games
Li Zhang 0045, Wei Wang 0011, Ziliang Han, Shijian Li, Gang Pan 0001 |
Frontiers Comput. Sci. | 7 |
| 2021 | Player Behavior Modeling for Enhancing Role-Playing Game EngagementabstractRole-playing games (RPGs) are one of the most exciting and most rapidly expanding genres of online games. Virtual characters that are not controlled by players, have become an integral part, which helps to advance narratives of RPGs. Believable characters can enhance game engagement and further improve player retention. However, game players easily find that most characters' behaviors are limited and improbable, resulting in a less meaningful game experience. In this work, we propose a framework to model game behaviors to learn behavior patterns of human players. Based on the learned behavior patterns, it generates human-like action sequences that can be used for the design of believable virtual characters in RPGs, so as to enhance game engagement. Specifically, considering the influence of game context in behavior patterns, we integrate game context (players' levels and game classes) with actions together to model behaviors. We propose a long-term memory cell on actions and game context to learn the hidden representations. We also introduce an attention mechanism to measure the contribution of the actions previously performed to the next action. Given only one action, our model can generate action sequences by predicting the succeeding action based on the previously generated actions. The model was evaluated on a real-world data set of over 22 000 players and more than 51 million action logs of an RPG game in 21 days. The results demonstrate the state-of-the-art performance. Sha Zhao, Yizhi Xu, Zhiling Luo, Jianrong Tao, Shijian Li, Changjie Fan, Gang Pan 0001 |
IEEE Trans. Comput. Soc. Syst. | 7 |
| 2021 | HisRect: Features from Historical Visits and Recent Tweet for Co-Location JudgementabstractEnabled by smartphones, social media users are increasingly going mobile. This trend fosters various location based services on social media platforms (e.g., Twitter). Many services like friends notification and community detection benefit from co-location judgement, i.e., to decide whether two Twitter users are co-located in some point-of-interest (POI). This problem is challenging due to the limited information in tweets and the lack of explicit geo-tags in tweets that can be used as labeled data. Our approach to this problem is based on a novel concept of HisRect features extracted from users' historical visits and recent tweets: The former has impacts on where a user visits in general, whereas the latter gives more hints about where a user is currently. In practice, labeled data is scarce. Therefore, we design a semi-supervised learning (SSL) framework that leverages unlabeled data to extract HisRect features. Moreover, we employ an embedding neural network layer to process HisRect features of two users, which decides co-location based on the embedding difference between the two features. Our model is extensively evaluated on two large sets of real Twitter data from more than one million users. The experimental results demonstrate that our HisRect features and SSL framework are highly effective at deciding co-locations. In terms of multiple metrics, our approach clearly outperforms alternative approaches using state-of-the-art techniques. Pengfei Li 0005, Hua Lu 0001, Shijian Li, Gang Pan 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2020 | Effective AER Object Classification Using Segmented Probability-Maximization Learning in Spiking Neural NetworksabstractAddress event representation (AER) cameras have recently attracted more attention due to the advantages of high temporal resolution and low power consumption, compared with traditional frame-based cameras. Since AER cameras record the visual input as asynchronous discrete events, they are inherently suitable to coordinate with the spiking neural network (SNN), which is biologically plausible and energy-efficient on neuromorphic hardware. However, using SNN to perform the AER object classification is still challenging, due to the lack of effective learning algorithms for this new representation. To tackle this issue, we propose an AER object classification model using a novel segmented probability-maximization (SPA) learning algorithm. Technically, 1) the SPA learning algorithm iteratively maximizes the probability of the classes that samples belong to, in order to improve the reliability of neuron responses and effectiveness of learning; 2) a peak detection (PD) mechanism is introduced in SPA to locate informative time points segment by segment, based on which information within the whole event stream can be fully utilized by the learning. Extensive experimental results show that, compared to state-of-the-art methods, not only our model is more effective, but also it requires less information to reach a certain level of accuracy. Qianhui Liu, Haibo Ruan, Dong Xing, Huajin Tang, Gang Pan 0001 |
AAAI | 5 |
| 2020 | HisRect: Features from Historical Visits and Recent Tweet for Co-Location JudgementabstractThis study explores the problem of co-location judgement, i.e., to decide whether two Twitter users are co-located at some point-of-interest (POI). We extract novel features, named HisRect, from users' historical visits and recent tweets: The former has impact on where a user visits in general, whereas the latter gives more hints about where a user is currently. To alleviate the issue of data scarcity, a semi-supervised learning (SSL) framework is designed to extract HisRect features. Moreover, we use an embedding neural network layer to decide co-location based on the difference between two users' His-Rect features. Extensive experiments on real Twitter data suggest that our HisRect features and SSL framework are highly effective at deciding co-locations. Pengfei Li 0005, Hua Lu 0001, Shijian Li, Gang Pan 0001 |
ICDE | 5 |
| 2020 | Maximum Entropy Reinforcement Learning with Evolution StrategiesabstractEvolution strategies (ES) have recently raised attention in solving challenging tasks with low computation costs and high scalability. However, it is well-known that evolution strategies reinforcement learning (RL) methods suffer from low stability. Without careful consideration, ES methods are sensitive to local optima and are unstable in learning. Therefore, there is an urgent need for improving the stability of ES methods in solving RL problems. In this paper, we propose a simple yet efficient ES method to stabilize the learning. Specifically, we propose a framework to incorporate the maximum entropy reinforcement learning with evolution strategies and derive an efficient entropy calculation method for linear policies. We further present a practical algorithm called maximum entropy evolution policy search based on the proposed framework, which is efficient and stable for policy search in continuous control. Our algorithm shows high stability across different random seeds and can obtain comparable results in performance against some existing derivative-free RL methods on several of the well-known benchmark MuJoCo robotic control tasks. Longxiang Shi, Shijian Li, Longbing Cao, Long Yang 0004, Gang Pan 0001 |
IJCNN | 6 |
| 2020 | Reconstructing Perceptive Images from Brain Activity by Shape-Semantic GANabstractReconstructing seeing images from fMRI recordings is an absorbing research area in neuroscience and provides a potential brain-reading technology. The challenge lies in that visual encoding in brain is highly complex and not fully revealed. Inspired by the theory that visual features are hierarchically represented in cortex, we propose to break the complex visual signals into multi-level components and decode each component separately. Specifically, we decode shape and semantic representations from the lower and higher visual cortex respectively, and merge the shape and semantic information to images by a generative adversarial network (Shape-Semantic GAN). This 'divide and conquer' strategy captures visual information more accurately. Experiments demonstrate that Shape-Semantic GAN improves the reconstruction similarity and image quality, and achieves the state-of-the-art image reconstruction performance. Gang Pan 0001 |
NeurIPS | 3 |
| 2020 | LISA: A Learned Index Structure for Spatial DataabstractIn spatial query processing, the popular index R-tree may incur large storage consumption and high IO cost. Inspired by the recent learned index [17] that replaces B-tree with machine learning models, we study an analogy problem for spatial data. We propose a novel Learned Index structure for Spatial dAta (LISA for short). Its core idea is to use machine learning models, through several steps, to generate searchable data layout in disk pages for an arbitrary spatial dataset. In particular, LISA consists of a mapping function that maps spatial keys (points) into 1-dimensional mapped values, a learned shard prediction function that partitions the mapped space into shards, and a series of local models that organize shards into pages. Based on LISA, a range query algorithm is designed, followed by a lattice regression model that enables us to convert a KNN query to range queries. Algorithms are also designed for LISA to handle data updates. Extensive experiments demonstrate that LISA clearly outperforms R-tree and other alternatives in terms of storage consumption and IO cost for queries. Moreover, LISA can handle data insertions and deletions efficiently. Pengfei Li 0005, Hua Lu 0001, Long Yang 0004, Gang Pan 0001 |
SIGMOD Conference | 5 |
| 2020 | Perception-enhancement based task learning and action scheduling for robotic limb in CPS environment
Shijian Li, Minhao Shi, Runhe Huang, Gang Pan 0001 |
Future Gener. Comput. Syst. | 5 |
| 2020 | Binless Kernel Machine: Modeling Spike Train Transformation for Cognitive Neural ProsthesesabstractModeling spike train transformation among brain regions helps in designing a cognitive neural prosthesis that restores lost cognitive functions. Various methods analyze the nonlinear dynamic spike train transformation between two cortical areas with low computational eficiency. The application of a real-time neural prosthesis requires computational eficiency, performance stability, and better interpretation of the neural firing patterns that modulate target spike generation. We propose the binless kernel machine in the point-process framework to describe nonlinear dynamic spike train transformations. Our approach embeds the binless kernel to eficiently capture the feedforward dynamics of spike trains and maps the input spike timings into reproducing kernel Hilbert space (RKHS). An inhomogeneous Bernoulli process is designed to combine with a kernel logistic regression that operates on the binless kernel to generate an output spike train as a point process. Weights of the proposed model are estimated by maximizing the log likelihood of output spike trains in RKHS, which allows a global-optimal solution. To reduce computational complexity, we design a streaming-based clustering algorithm to extract typical and important spike train features. The cluster centers and their weights enable the visualization of the important input spike train patterns that motivate or inhibit output neuron firing. We test the proposed model on both synthetic data and real spike train data recorded from the dorsal premotor cortex and the primary motor cortex of a monkey performing a center-out task. Performances are evaluated by discrete-time rescaling Kolmogorov-Smirnov tests. Our model outperforms the existing methods with higher stability regardless of weight initialization and demonstrates higher eficiency in analyzing neural patterns from spike timing with less historical input (50%). Meanwhile, the typical spike train patterns selected according to weights are validated to encode output spike from the spike train of single-input neuron and the interaction of two input neurons. Cunle Qian, Xuyun Sun, Yueming Wang 0001, Xiaoxiang Zheng, Yiwen Wang 0002, Gang Pan 0001 |
Neural Comput. | 6 |
| 2020 | Deep CovDenseSNN: A hierarchical event-driven dynamic framework with spiking neurons in noisy environment
Qi Xu 0008, Jianxin Peng, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
Neural Networks | 5 |
| 2020 | Gender Profiling From a Single Snapshot of Apps Installed on a Smartphone: An Empirical StudyabstractThe integration of the fifth generation (5G) networks and artificial intelligence (AI) benefits to create a more holistic and better connected ecosystem for industries. User profiling has become an important issue for industries to improve company profit. In the 5G era, smartphone applications have become an indispensable part in our everyday lives. Users determine what apps to install based on their personal needs, interests, and tastes, which is likely shaped by their genders-the behavioral, cultural, or psychological traits typically associated with their sex. It is possible to profile users' gender based simply on a single snapshot of apps installed on their smartphones. With this inference based on easy to access data, we can make smartphone systems more user-friendly, and provide better personalized products and services. In this article, we explore such possibilities through an empirical study on a large-scale dataset of installed app lists from 15 000 Android users. More specifically, we investigate the following research questions: 1) What differences between females and males can be explored from installed app lists? 2) Can user gender be reliably inferred from a snapshot of apps installed? Which snapshot feature(s) are the most predictive? What is the best combination of features for building the gender prediction model? 3) What are the limitations of a gender prediction model based solely on a snapshot of apps installed on a smartphone? We find significant gender differences in app type, function, and icon design. We then extract the corresponding features from a snapshot of apps installed to infer the gender of each user. We assess the gender predictive ability of individual features and combinations of different features. We achieve an accuracy of 76.62% and area under the curve of 84.23% with the best set of features, outperforming the existing work by around 5% and 10%, respectively. Finally, we perform an error analysis on misclassified users and discussed the implications and limitations of this article. Sha Zhao, Yizhi Xu, Xiaojuan Ma, Ziwen Jiang, Zhiling Luo, Shijian Li, Laurence T. Yang, Anind K. Dey, Gang Pan 0001 |
IEEE Trans. Ind. Informatics | 9 |
| 2020 | Forecasting Price Trend of Bulk Commodities Leveraging Cross-domain Open Data FusionabstractForecasting price trend of bulk commodities is important in international trade, not only for markets participants to schedule production and marketing plans but also for government administrators to adjust policies. Previous studies cannot support accurate fine-grained short-term prediction, since they mainly focus on coarse-grained long-term prediction using historical data. Recently, cross-domain open data provides possibilities to conduct fine-grained price forecasting, since they can be leveraged to extract various direct and indirect factors of the price. In this article, we predict the price trend over upcoming days, by leveraging cross-domain open data fusion. More specifically, we formulate the price trend into three classes (rise, slight-change, and fall), and then we predict the specific class in which the price trend of the future day lies. We take three factors into consideration: (1) supply factor considering sources providing bulk commodities,<?brk?> (2) demand factor focusing on vessel transportation with reflection of short time needs, and (3) expectation factor encompassing indirect features (e.g., air quality) with latent influences. A hybrid classification framework is proposed for the price trend forecasting. Evaluation conducted on nine real-world cross-domain open datasets shows that our framework can forecast the price trend accurately, outperforming multiple state-of-the-art baselines. Binbin Zhou 0005, Sha Zhao, Longbiao Chen, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2020 | Unsupervised AER Object Recognition Based on Multiscale Spatio-Temporal Features and Spiking NeuronsabstractThis article proposes an unsupervised address event representation (AER) object recognition approach. The proposed approach consists of a novel multiscale spatio-temporal feature (MuST) representation of input AER events and a spiking neural network (SNN) using spike-timing-dependent plasticity (STDP) for object recognition with MuST. MuST extracts the features contained in both the spatial and temporal information of AER event flow, and forms an informative and compact feature spike representation. We show not only how MuST exploits spikes to convey information more effectively, but also how it benefits the recognition using SNN. The recognition process is performed in an unsupervised manner, which does not need to specify the desired status of every single neuron of SNN, and thus can be flexibly applied in real-world recognition tasks. The experiments are performed on five AER datasets including a new one named GESTURE-DVS. Extensive experimental results show the effectiveness and advantages of the proposed approach. Qianhui Liu, Gang Pan 0001, Haibo Ruan, Dong Xing, Qi Xu 0008, Huajin Tang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-NetworkabstractThe deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. The DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample trajectories. In this brief, we propose a general framework to combine the DQN and most of the return-based reinforcement learning algorithms, named R-DQN. We show that the performance of the traditional DQN can be significantly improved by introducing return-based algorithms. In order to further improve the R-DQN, we design a strategy with two measurements to qualitatively measure the policy discrepancy. We conduct experiments on several representative tasks from the OpenAI Gym and Atari games. The state-of-the-art performance achieved by our method with this proposed strategy validates its effectiveness. Wenjia Meng, Long Yang 0004, Pengfei Li 0005, Gang Pan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Summary study of data-driven photometric stereo methodsabstractA photometric stereo method aims to recover the surface normal of a 3D object observed under varying light directions. It is an ill-defined problem because the general reflectance properties of the surface are unknown. This paper reviews existing data-driven methods, with a focus on their technical insights into the photometric stereo problem. We divide these methods into two categories, per-pixel and all-pixel, according to how they process an image. We discuss the differences and relationships between these methods from the perspective of inputs, networks, and data, which are key factors in designing a deep learning approach. We demonstrate the performance of the models using a popular benchmark dataset. Data-driven photometric stereo methods have shown that they possess a superior performance advantage over traditional methods. However, these methods suffer from various limitations, such as limited generalization capability. Finally, this study suggests directions for future research. Boxin Shi, Gang Pan 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2019 | LSTM with Uniqueness Attention for Human Activity Recognition
Zengwei Zheng, Lifei Shi, Lin Sun 0006, Gang Pan 0001 |
ICANN (3) | 5 |
| 2019 | Location Inference for Non-Geotagged Tweets in User Timelines [Extended Abstract]abstractThis study explores the problem of inferring locations for individual tweets. We scrutinize Twitter user timelines in a novel fashion. First of all, we split each user's tweet timeline temporally into a number of clusters, each tending to imply a distinct location. Subsequently, we adapt machine learning models to our setting and design classifiers that classify each tweet cluster into one of the pre-defined location classes at the city level. Extensive experiments on a large set of real Twitter data suggest that our models are effective at inferring locations for non-geotagged tweets and outperform the state-of-the-art approaches significantly in terms of inference accuracy. Pengfei Li 0005, Hua Lu 0001, Nattiya Kanhabua, Sha Zhao, Gang Pan 0001 |
ICDE | 5 |
| 2019 | AppUsage2Vec: Modeling Smartphone App Usage for PredictionabstractApp usage prediction, i.e. which apps will be used next, is very useful for smartphone system optimization, such as operating system resource management, battery energy consumption optimization, and user experience improvement as well. However, it is still challenging to achieve usage prediction of high accuracy. In this paper, we propose a novel framework for app usage prediction, called AppUsage2Vec, inspired by Doc2Vec. It models app usage records by considering the contribution of different apps, user personalized characteristics, and temporal context. We measure the contribution of each app to the target app by introducing an app-attention mechanism. The user personalized characteristics in app usage are learned by a module of dual-DNN. Furthermore, we encode the top-k supervised information in loss function for training the model to predict the app most likely to be used next. The AppUsage2Vec was evaluated on a dataset of 10,360 users and 46,434,380 records in three months. The results demonstrate the state-of-the-art performance. Sha Zhao, Zhiling Luo, Ziwen Jiang, Shijian Li, Jianwei Yin, Gang Pan 0001 |
ICDE | 8 |
| 2019 | STCA: Spatio-Temporal Credit Assignment with Delayed Feedback in Deep Spiking Neural NetworksabstractThe temporal credit assignment problem, which aims to discover the predictive features hidden in distracting background streams with delayed feedback, remains a core challenge in biological and machine learning. To address this issue, we propose a novel spatio-temporal credit assignment algorithm called STCA for training deep spiking neural networks (DSNNs). We present a new spatiotemporal error backpropagation policy by defining a temporal based loss function, which is able to credit the network losses to spatial and temporal domains simultaneously. Experimental results on MNIST dataset and a music dataset (MedleyDB) demonstrate that STCA can achieve comparable performance with other state-of-the-art algorithms with simpler architectures. Furthermore, STCA successfully discovers predictive sensory features and shows the highest performance in the unsegmented sensory event detection tasks. Pengjie Gu, Rong Xiao 0001, Gang Pan 0001, Huajin Tang |
IJCAI | 3 |
| 2019 | Multi-layer Temporal Network Analysis Reveals Increasing Temporal Reachability and Spreadability in the First Two Years of Life
Zhen Zhou 0004, Han Zhang 0002, Li-Ming Hsu, Weili Lin, Gang Pan 0001, Dinggang Shen |
MICCAI (3) | 5 |
| 2019 | Dynamic Ensemble Modeling Approach to Nonstationary Neural Decoding in Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-art neural signal decoders such as Kalman filter assume fixed relationship between neural activities and motor movements, thus will fail if this assumption is not satisfied. We propose a dynamic ensemble modeling (DyEnsemble) approach that is capable of adapting to changes in neural signals by employing a proper combination of decoding functions. The DyEnsemble method firstly learns a set of diverse candidate models. Then, it dynamically selects and combines these models online according to Bayesian updating mechanism. Our method can mitigate the effect of noises and cope with different task behaviors by automatic model switching, thus gives more accurate predictions. Experiments with neural data demonstrate that the DyEnsemble method outperforms Kalman filters remarkably, and its advantage is more obvious with noisy signals. Yueming Wang 0001, Gang Pan 0001 |
NeurIPS | 4 |
| 2019 | Exploiting Ratings, Reviews and Relationships for Item Recommendations in Topic Based Social NetworksabstractMany e-commerce platforms today allow users to give their rating scores and reviews on items as well as to establish social relationships with other users. As a result, such platforms accumulate heterogeneous data including numeric scores, short textual reviews, and social relationships. However, many recommender systems only consider historical user feedbacks in modeling user preferences. More specifically, most existing recommendation approaches only use rating scores but ignore reviews and social relationships in the user-generated data. In this paper, we propose TSNPF-a latent factor model to effectively capture user preferences and item features. Employing Poisson factorization, TSNPF fully exploits the wealth of information in rating scores, review text and social relationships altogether. It extracts topics of items and users from the review text and makes use of similarities between user pairs with social relationships, which results in a comprehensive understanding of user preferences. Experimental results on real-world datasets demonstrate that our TSNPF approach is highly effective at recommending items to users. Pengfei Li 0005, Hua Lu 0001, Gang Zheng 0005, Long Yang 0004, Gang Pan 0001 |
WWW | 6 |
| 2019 | Investigating smartphone user differences in their application usage behaviors: an empirical study
Sha Zhao, Yizhi Xu, Xiaojuan Ma, Zhiling Luo, Shijian Li, Anind K. Dey, Gang Pan 0001 |
CCF Trans. Pervasive Comput. Interact. | 8 |
| 2019 | Activity-dependent neuron model for noise resistance
Yueming Wang 0001, Gang Pan 0001 |
Neurocomputing | 6 |
| 2019 | Overfitting remedy by sparsifying regularization on fully-connected layers of CNNs
Qi Xu 0008, Ming Zhang 0018, Zonghua Gu 0001, Gang Pan 0001 |
Neurocomputing | 4 |
| 2019 | State Distribution-Aware Sampling for Deep Q-Learning
Weichao Li 0003, Fuxian Huang, Xi Li 0001, Gang Pan 0001, Fei Wu 0001 |
Neural Process. Lett. | 4 |
| 2019 | User profiling from their use of smartphone applications: A survey
Sha Zhao, Shijian Li, Julian Ramos 0001, Zhiling Luo, Ziwen Jiang, Anind K. Dey, Gang Pan 0001 |
Pervasive Mob. Comput. | 7 |
| 2019 | Deep Attention Network for Egocentric Action RecognitionabstractRecognizing a camera wearer's actions from videos captured by an egocentric camera is a challenging task. In this paper, we employ a two-stream deep neural network composed of an appearance-based stream and a motion-based stream to recognize egocentric actions. Based on the insight that human action and gaze behavior are highly coordinated in object manipulation tasks, we propose a spatial attention network to predict human gaze in the form of attention map. The attention map helps each of the two streams to focus on the most relevant spatial region of the video frames to predict actions. To better model the temporal structure of the videos, a temporal network is proposed. The temporal network incorporates bi-directional long short-term memory to model the long-range dependencies to recognize egocentric actions. The experimental results demonstrate that our method is able to predict attention maps that are consistent with human attention and achieve competitive action recognition performance with the state-of-the-art methods on the GTEA Gaze and GTEA Gaze+ datasets. Minlong Lu, Ze-Nian Li, Yueming Wang 0001, Gang Pan 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Numerical Reflectance Compensation for Non-Lambertian Photometric StereoabstractThe surface normal estimation from photometric stereo becomes less reliable when the surface reflectance deviates from the Lambertian assumption. The non-Lambertian effect can be explicitly addressed by physics modeling to the reflectance function, at the cost of introducing highly nonlinear optimization. This paper proposes a numerical compensation scheme that attempts to minimize the angular error to address the non-Lambertian photometric stereo problem. Due to the multifaceted influence in the modeling of non-Lambertian reflectance in photometric stereo, directly minimizing the angular errors of surface normal is a highly complex problem. We introduce an alternating strategy, in which the estimated reflectance can be temporarily regarded as a known variable, to simplify the formulation of angular error. To reduce the impact of inaccurately estimated reflectance in this simplification, we propose a numerical compensation scheme whose compensation weight is formulated to reflect the reliability of estimated reflectance. Finally, the solution for the proposed numerical compensation scheme is efficiently computed by using cosine difference to approximate the angular difference. The experimental results show that our method can significantly improve the performance of the state-of-the-art methods on both synthetic data and real data with small additive costs. Moreover, our method initialized by results from the baseline method (least-square-based) achieves the state-of-the-art performance on both synthetic data and real data with significantly smaller overall computation, i.e., about eight times faster compared with the state-of-the-art methods. Ajay Kumar 0001, Boxin Shi, Gang Pan 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Location Inference for Non-Geotagged Tweets in User TimelinesabstractSocial media like Twitter have become globally popular in the past decade. Thanks to the high penetration of smartphones, social media users are increasingly going mobile. This trend has contributed to foster various location based services deployed on social media, the success of which heavily depends on the availability and accuracy of users' location information. However, only a very small fraction of tweets in Twitter are geo-tagged. Therefore, it is necessary to infer locations for tweets in order to attain the purpose of those location based services. In this paper, we tackle this problem by scrutinizing Twitter user timelines in a novel fashion. First of all, we split each user's tweet timeline temporally into a number of clusters, each tending to imply a distinct location. Subsequently, we adapt two machine learning models to our setting and design classifiers that classify each tweet cluster into one of the pre-defined location classes at the city level. The Bayes based model focuses on the information gain of words with location implications in the user-generated contents. The convolutional LSTM model treats user-generated contents and their associated locations as sequences and employs bidirectional LSTM and convolution operation to make location inferences. The two models are evaluated on a large set of real Twitter data. The experimental results suggest that our models are effective at inferring locations for non-geotagged tweets and the models outperform the state-of-the-art and alternative approaches significantly in terms of inference accuracy. Pengfei Li 0005, Hua Lu 0001, Nattiya Kanhabua, Sha Zhao, Gang Pan 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Editorial: Booming of Neural Networks and Learning SystemsabstractAs you open this January issue of the IEEE Transactions on Neural Networks and Learning Systems (TNNLS), I hope everyone enjoyed a great holiday season and is excited for the new year of 2019. I am very delighted and honored to report several key metrics of IEEE TNNLS to the community. Akira Hirose 0001, Alessio Micheli, Artur S. d'Avila Garcez, Choon Ki Ahn, Gang Pan 0001, Hamid Reza Karimi, Jianbing Shen, José de Jesús Rubio, Lei Zhang 0005, Lingjia Liu 0001, Lorenzo Livi, Nishchal K. Verma, Pedro Antonio Gutiérrez, Qi Tian 0001, Qinglai Wei, Seiichi Ozawa, Stuart Harvey Rubin, Weineng Chen, Xi Li 0001, Xiaofeng Liao 0001, Youmin Zhang 0001, Zhen Ni, Haibo He |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Epileptic State Segmentation with Temporal-Constrained ClusteringabstractAutomatic seizure identification plays an important role in epilepsy evaluation. Most existing methods regard seizure identification as a classification problem and rely on labelled training set. However, labelling seizure onset is very expensive and seizure data for each individual is especially limited, classifier-based methods are usually impractical in use. Clustering methods could learn useful information from unlabelled data, while they may lead to unstable results given epileptic signals with high noises. In this paper, we propose to use Gaussian temporal-constrained k-medoids method for seizure state segmentation. Using temporal information, the noises could be effectively suppressed and robust clustering performance is achieved. Besides, a new criterion called signed total variation (STV) which describes temporal integrity and consistency is proposed for temporal-constrained clustering evaluation. Experimental results show that, compared with the existing methods, the k-medoids method with Gaussian temporal constraint achieves the best results on both F1-score and STV. Kang Lin, Shaozhe Feng, Qi Lian, Gang Pan 0001, Yueming Wang 0001 |
ICASSP | 5 |
| 2018 | Knowledge-Guided Agent-Tactic-Aware Learning for StarCraft MicromanagementabstractAs an important and challenging problem in artificial intelligence (AI) game playing, StarCraft micromanagement involves a dynamically adversarial game playing process with complex multi-agent control within a large action space. In this paper, we propose a novel knowledge-guided agent-tactic-aware learning scheme, that is, opponent-guided tactic learning (OGTL), to cope with this micromanagement problem. In principle, the proposed scheme takes a two-stage cascaded learning strategy which is capable of not only transferring the human tactic knowledge from the human-made opponent agents to our AI agents but also improving the adversarial ability. With the power of reinforcement learning, such a knowledge-guided agent-tactic-aware scheme has the ability to guide the AI agents to achieve high winning-rate performances while accelerating the policy exploration process in a tactic-interpretable fashion. Experimental results demonstrate the effectiveness of the proposed scheme against the state-of-the-art approaches in several benchmark combat scenarios. Yue Hu 0008, Xi Li 0001, Gang Pan 0001, Mingliang Xu 0001 |
IJCAI | 4 |
| 2018 | Jointly Learning Network Connections and Link Weights in Spiking Neural NetworksabstractSpiking neural networks (SNNs) are considered to be biologically plausible and power-efficient on neuromorphic hardware. However, unlike the brain mechanisms, most existing SNN algorithms have fixed network topologies and connection relationships. This paper proposes a method to jointly learn network connections and link weights simultaneously. The connection structures are optimized by the spike-timing-dependent plasticity (STDP) rule with timing information, and the link weights are optimized by a supervised algorithm. The connection structures and the weights are learned alternately until a termination condition is satisfied. Experiments are carried out using four benchmark datasets. Our approach outperforms classical learning methods such as STDP, Tempotron, SpikeProp, and a state-of-the-art supervised algorithm. In addition, the learned structures effectively reduce the number of connections by about 24%, thus facilitate the computational efficiency of the network. Jiangrong Shen, Yueming Wang 0001, Huajin Tang, Hang Yu 0010, Zhaohui Wu 0001, Gang Pan 0001 |
IJCAI | 7 |
| 2018 | CSNN: An Augmented Spiking based Framework with Perceptron-InceptionabstractSpiking Neural Networks (SNNs) represent and transmit information in spikes, which is considered more biologically realistic and computationally powerful than the traditional Artificial Neural Networks. The spiking neurons encode useful temporal information and possess highly anti-noise property. The feature extraction ability of typical SNNs is limited by shallow structures. This paper focuses on improving the feature extraction ability of SNNs in virtue of powerful feature extraction ability of Convolutional Neural Networks (CNNs). CNNs can extract abstract features resorting to the structure of the convolutional feature maps. We propose a CNN-SNN (CSNN) model to combine feature learning ability of CNNs with cognition ability of SNNs. The CSNN model learns the encoded spatial temporal representations of images in an event-driven way. We evaluate the CSNN model on the handwritten digits images dataset MNIST and its variational databases. In the presented experimental results, the proposed CSNN model is evaluated regarding learning capabilities, encoding mechanisms, robustness to noisy stimuli and its classification performance. The results show that CSNN behaves well compared to other cognitive models with significantly fewer neurons and training samples. Our work brings more biological realism into modern image classification models, with the hope that these models can inform how the brain performs this high-level vision task. Qi Xu 0008, Hang Yu 0010, Jiangrong Shen, Huajin Tang, Gang Pan 0001 |
IJCAI | 6 |
| 2018 | A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement LearningabstractRecently, a new multi-step temporal learning algorithm Q(σ) unifies n-step Tree-Backup (when σ = 0) and n-step Sarsa (when σ = 1) by introducing a sampling parameter σ. However, similar to other multi-step temporal-difference learning algorithms, Q(σ) needs much memory consumption and computation time. Eligibility trace is an important mechanism to transform the off-line updates into efficient on-line ones which consume less memory and computation time. In this paper, we combine the original Q(σ) with eligibility traces and propose a new algorithm, called Qπ(σ,λ), where λ is trace-decay parameter. This new algorithm unifies Sarsa(λ) (when σ = 1) and Qπ (λ) (when σ = 0). Furthermore, we give an upper error bound of Qπ(σ,λ) policy evaluation algorithm. We prove that Qπ (σ, λ) control algorithm converges to the optimal value function exponentially. We also empirically compare it with conventional temporal-difference learning methods. Results show that, with an intermediate value of σ, Qπ(σ,λ) creates a mixture of the existing algorithms which learn the optimal value significantly faster than the extreme end (σ = 0, or 1). Long Yang 0004, Minhao Shi, Wenjia Meng, Gang Pan 0001 |
IJCAI | 5 |
| 2018 | A Supervised Multi-Spike Learning Algorithm for Spiking Neural NetworksabstractThe formulation of efficient supervised learning algorithms for Spiking Neural Network (SNN) is difficult and remains challenging. This paper presents a supervised multispike learning algorithm, which is used to train neurons to output spike train with a target firing rate. The proposed algorithm simplifies the expression of the membrane potential by assuming a special condition of the threshold, thus allows the application of a gradient descent to optimize the synaptic weights. Additionally, in the presented experimental results, the proposed algorithm is evaluated regarding its initial setups, its classification performance for rate-based and timing-based patterns and its capability to sound recognition. The results also demonstrate that the proposed algorithm can achieve a competitive accuracy in temporal pattern classification and sound recognition. Huajin Tang, Gang Pan 0001 |
IJCNN | 3 |
| 2018 | Nonlinear Modeling of Neural Interaction for Spike Prediction Using the Staged Point-Process ModelabstractNeurons communicate nonlinearly through spike activities. Generalized linear models (GLMs) describe spike activities with a cascade of a linear combination across inputs, a static nonlinear function, and an inhomogeneous Bernoulli or Poisson process, or Cox process if a self-history term is considered. This structure considers the output nonlinearity in spike generation but excludes the nonlinear interaction among input neurons. Recent studies extend GLMs by modeling the interaction among input neurons with a quadratic function, which considers the interaction between every pair of input spikes. However, quadratic effects may not fully capture the nonlinear nature of input interaction. We therefore propose a staged point-process model to describe the nonlinear interaction among inputs using a few hidden units, which follows the idea of artificial neural networks. The output firing probability conditioned on inputs is formed as a cascade of two linear-nonlinear (a linear combination plus a static nonlinear function) stages and an inhomogeneous Bernoulli process. Parameters of this model are estimated by maximizing the log likelihood on output spike trains. Unlike the iterative reweighted least squares algorithm used in GLMs, where the performance is guaranteed by the concave condition, we propose a modified Levenberg-Marquardt (L-M) algorithm, which directly calculates the Hessian matrix of the log likelihood, for the nonlinear optimization in our model. The proposed model is tested on both synthetic data and real spike train data recorded from the dorsal premotor cortex and primary motor cortex of a monkey performing a center-out task. Performances are evaluated by discrete-time rescaled Kolmogorov-Smirnov tests, where our model statistically outperforms a GLM and its quadratic extension, with a higher goodness-of-fit in the prediction results. In addition, the staged point-process model describes nonlinear interaction among input neurons with fewer parameters than quadratic models, and the modified L-M algorithm also demonstrates fast convergence. Cunle Qian, Xuyun Sun, Shaomin Zhang, Dong Xing, Hongbao Li, Xiaoxiang Zheng, Gang Pan 0001, Yiwen Wang 0002 |
Neural Comput. | 7 |
| 2017 | Finding Influential Local Users with Similar Interest from Geo-Tagged Social Media DataabstractGeo-tagged social media data provides abundant resources for people in need of local information. In this paper, we study how to find the top-k influential local users from geo-tagged social media data who have interests similar to a query. Such local users can be of particular importance for a variety of activities from events organizing to online advertising. We formulate the problem as Top-k Influential Similar Local Query (TkISL) and provide a complete set of techniques for solving it. To effectively manage the social media users, we design three hybrid user profiling techniques, an indexing tree, and an upper bound query-user similarity that enables efficient pruning in query processing. To process TkISL queries, we propose a baseline method and a more efficient improved method. The former directly uses the indexing tree and the upper bound for pruning, whereas the latter speeds up the query processing by enhancing the tree and pruning. Finally, we conduct extensive experimental studies to evaluate our proposals on real geo-tagged tweet corpora. The experimental results demonstrate the efficiency and effectiveness of our proposals. Jinling Jiang, Hua Lu 0001, Pengfei Li 0005, Gang Pan 0001, Xike Xie |
MDM | 4 |
| 2017 | Understanding bike trip patterns leveraging bike sharing system open data
Longbiao Chen, Xiaojuan Ma, Thi Mai Trang Nguyen, Gang Pan 0001, Jérémie Jakubowicz |
Frontiers Comput. Sci. | 4 |
| 2017 | Darwin: A neuromorphic hardware co-processor based on spiking neural networks
De Ma, Juncheng Shen, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
J. Syst. Archit. | 9 |
| 2017 | Ubiquitous Intelligence and computing for enabling a smarter world
Diego López-de-Ipiña, Liming Chen 0001, Nathalie Mitton, Gang Pan 0001 |
Pers. Ubiquitous Comput. | 4 |
| 2017 | Fine-Grained Urban Event Detection and Characterization Based on Tensor CofactorizationabstractUnderstanding the irregular crowd movement and social activities caused by urban events such as city festivals and concerts can benefit event management and city planning. Although various urban data can be exploited to detect such irregularities, the crowd mobility data (e.g., bike trip records) are usually in a mixed state with several basic patterns (e.g., eating, working, and recreation), making it difficult to separate concurrent events happening in the same region. The social activity data (e.g., social network check-ins) are usually oversparse, hindering the fine-grained characterization of urban events. In this paper, we propose a tensor cofactorization-based data fusion framework for fine-grained urban event detection and characterization leveraging crowd mobility data and social activity data. First, we adopt a nonnegative tensor cofactorization approach to decompose the crowd mobility tensor into several basic patterns, with the help of the auxiliary social activity tensor. We then use a multivariate-outlier-detection-based method to identify irregularities from the decomposed basic patterns and aggregate them to detect and characterize the associated urban events. We evaluate the performance of our framework using real-world bike trip data and check-in data from New York City and Washington, DC, respectively. Results show that by fusing the two types of urban data, our method achieves fine-grained urban event detection and characterization in both cities and consistently outperforms the baselines. Longbiao Chen, Jérémie Jakubowicz, Dingqi Yang, Daqing Zhang 0001, Gang Pan 0001 |
IEEE Trans. Hum. Mach. Syst. | 5 |
| 2016 | Dynamic cluster-based over-demand prediction in bike sharing systemsabstractBike sharing is booming globally as a green transportation mode, but the occurrence of over-demand stations that have no bikes or docks available greatly affects user experiences. Directly predicting individual over-demand stations to carry out preventive measures is difficult, since the bike usage pattern of a station is highly dynamic and context dependent. In addition, the fact that bike usage pattern is affected not only by common contextual factors (e.g., time and weather) but also by opportunistic contextual factors (e.g., social and traffic events) poses a great challenge. To address these issues, we propose a dynamic cluster-based framework for over-demand prediction. Depending on the context, we construct a weighted correlation network to model the relationship among bike stations, and dynamically group neighboring stations with similar bike usage patterns into clusters. We then adopt Monte Carlo simulation to predict the over-demand probability of each cluster. Evaluation results using real-world data from New York City and Washington, D.C. show that our framework accurately predicts over-demand clusters and outperforms the baseline methods significantly. Longbiao Chen, Daqing Zhang 0001, Leye Wang, Dingqi Yang, Xiaojuan Ma, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001, Thi Mai Trang Nguyen, Jérémie Jakubowicz |
UbiComp | 8 |
| 2016 | Discovering different kinds of smartphone users through their application usage behaviorsabstractUnderstanding smartphone users is fundamental for creating better smartphones, and improving the smartphone usage experience and generating generalizable and reproducible research. However, smartphone manufacturers and most of the mobile computing research community make a simplifying assumption that all smartphone users are similar or, at best, constitute a small number of user types, based on their behaviors. Manufacturers design phones for the broadest audience and hope they work for all users. Researchers mostly analyze data from smartphone-based user studies and report results without accounting for the many different groups of people that make up the user base of smartphones. In this work, we challenge these elementary characterizations of smartphone users and show evidence of the existence of a much more diverse set of users. We analyzed one month of application usage from 106,762 Android users and discovered 382 distinct types of users based on their application usage behaviors, using our own two-step clustering and feature ranking selection approach. Our results have profound implications on the reproducibility and reliability of mobile computing studies, design and development of applications, determination of which apps should be pre-installed on a smartphone and, in general, on the smartphone usage experience for different types of users. Sha Zhao, Julian Ramos 0001, Jianrong Tao, Ziwen Jiang, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001, Anind K. Dey |
UbiComp | 7 |
| 2016 | Darwin: a neuromorphic hardware co-processor based on Spiking Neural Networks
Juncheng Shen, De Ma, Zonghua Gu 0001, Ming Zhang 0018, Xiaoqiang Xu, Qi Xu 0008, Yangjing Shen, Gang Pan 0001 |
Sci. China Inf. Sci. | 9 |
| 2016 | Robust discriminative non-negative matrix factorization
Ruiqing Zhang, Zhenfang Hu, Gang Pan 0001, Yueming Wang 0001 |
Neurocomputing | 3 |
| 2016 | A 3D Feature Descriptor Recovered from a Single 2D Palmprint ImageabstractDesign and development of efficient and accurate feature descriptors is critical for the success of many computer vision applications. This paper proposes a new feature descriptor, referred to as DoN, for the 2D palmprint matching. The descriptor is extracted for each point on the palmprint. It is based on the ordinal measure which partially describes the difference of the neighboring points' normal vectors. DoN has at least two advantages: 1) it describes the 3D information, which is expected to be highly stable under commonly occurring illumination variations during contactless imaging; 2) the size of DoN for each point is only one bit, which is computationally simple to extract, easy to match, and efficient to storage. We show that such 3D information can be extracted from a single 2D palmprint image. The analysis for the effectiveness of ordinal measure for palmprint matching is also provided. Four publicly available 2D palmprint databases are used to evaluate the effectiveness of DoN, both for identification and the verification. Our method on all these databases achieves the state-of-the-art performance. Ajay Kumar 0001, Gang Pan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2016 | Suspecting Less and Doing Better: New Insights on Palmprint Identification for Faster and More Accurate MatchingabstractThis paper introduces a generalized palmprint identification framework to unify several state-of-art 2D and 3D palmprint methods. Through this framework, we argue that the methods employing one-to-one matching strategy and binary representation for feature are more effective for palmprint identification. The analysis for the first argument is based on a statistical matching model and is supported by outperforming results on several publicly available 2D palmprpint databases. These two arguments are further evaluated for 3D palmprint matching and used to introduce a new method for encoding 3D palmprint feature. The proposed 3D feature is binary and more efficiently computed. It encodes the 3D shape of palmprint to either convex or concave. The experimental results on two publicly available, from contactless and contact-base 3D palmprint database of 177 and 200 subjects, respectively, outperform the state-of-the-art methods. This paper also provides our palmprint matching algorithm(s) in public domain, unlike the previous work in this area, which will help to further advance research efforts in this area. Ajay Kumar 0001, Gang Pan 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Container Port Performance Measurement and Comparison Leveraging Ship GPS Traces and Maritime Open DataabstractContainer ports are generally measured and compared using performance indicators such as container throughput and facility productivity. Being able to measure the performance of container ports quantitatively is of great importance for researchers to design models for port operation and container logistics. Instead of relying on the manually collected statistical information from different port authorities and shipping companies, we propose to leverage the pervasive ship GPS traces and maritime open data to derive port performance indicators, including ship traffic, container throughput, berth utilization, and terminal productivity. These performance indicators are found to be directly related to the number of container ships arriving at the terminals and the number of containers handled at each ship. Therefore, we propose a framework that takes the ships' container-handling events at terminals as the basis for port performance measurement. With the inferred port performance indicators, we further compare the strengths and weaknesses of different container ports at the terminal level, port level, and region level, which can potentially benefit terminal productivity improvement, liner schedule optimization, and regional economic development planning. In order to evaluate the proposed framework, we conduct extensive studies on large-scale real-world GPS traces of container ships collected from major container ports worldwide through the year, as well as various maritime open data sources concerning ships and ports. Evaluation results confirm that the proposed framework not only can accurately estimate various port performance indicators but also effectively produces port comparison results such as port performance ranking and port region comparison. Longbiao Chen, Daqing Zhang 0001, Xiaojuan Ma, Leye Wang, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2016 | Weakly Supervised Metric Learning for Traffic Sign Recognition in a LIDAR-Equipped VehicleabstractWe address the problem of traffic sign recognition in a light detection and ranging (LIDAR)-equipped vehicle. With the help of 3-D LIDAR points, the 2-D multiview sign images will be easily detected from the captured images of street signs. After detection, the sign recognition problem is formulated as a multiview object recognition task. We develop a metric-learning-based template matching approach for this task and learn a distance metric between the captured images and the corresponding sign templates. For each sign, recognition is done via soft voting by the recognition results of its corresponding multiview images. We propose a latent structural support vector machine (SVM)-based weakly supervised metric learning (WSMLR) method to learn the metric and a reliability classifier. The reliability classifier is used to determine each image's reliability, which serves as each image's weight in both the learning and soft voting procedure. We evaluate the proposed method for multiview traffic sign recognition on a multiview traffic sign data set with 112 categories and observe very encouraging results compared with other state-of-the-art methods. In addition, the method can be customized to solve the single-view sign recognition. The performance of our method for single-view sign recognition is tested on two public data sets, showing that our method is comparable with other competitive ones. Baoyuan Wang, Zhaohui Wu 0001, Jingdong Wang 0001, Gang Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2016 | Sparse Principal Component Analysis via Rotation and TruncationabstractSparse principal component analysis (sparse PCA) aims at finding a sparse basis to improve the interpretability over the dense basis of PCA, while still covering the data subspace as much as possible. In contrast to most existing work that addresses the problem by adding sparsity penalties on various objectives of PCA, we propose a new method, sparse PCA via rotation and truncation (SPCArt), which finds a rotation matrix and a sparse basis such that the sparse basis approximates the basis of PCA after the rotation. The algorithm of SPCArt consists of three alternating steps: 1) rotating the PCA basis; 2) truncating small entries; and 3) updating the rotation matrix. Its performance bounds are also given. The SPCArt is efficient, with each iteration scaling linearly with the data dimension. Parameter choice is simple, due to explicit physical explanations. We give a unified view to several existing sparse PCA methods and discuss the connections with SPCArt. Some ideas from SPCArt are extended to GPower, a popular sparse PCA algorithm, to address its limitations. Experimental results demonstrate that SPCArt achieves the state-of-the-art performance, along with a good tradeoff among various criteria, including sparsity, explained variance, orthogonality, balance of sparsity among loadings, and computational speed. Zhenfang Hu, Gang Pan 0001, Yueming Wang 0001, Zhaohui Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Improving object detection with deep convolutional networks via Bayesian optimization and structured predictionabstractObject detection systems based on the deep convolutional neural network (CNN) have recently made ground-breaking advances on several object detection benchmarks. While the features learned by these high-capacity neural networks are discriminative for categorization, inaccurate localization is still a major source of error for detection. Building upon high-capacity CNN architectures, we address the localization problem by 1) using a search algorithm based on Bayesian optimization that sequentially proposes candidate regions for an object bounding box, and 2) training the CNN with a structured loss that explicitly penalizes the localization inaccuracy. In experiments, we demonstrate that each of the proposed methods improves the detection performance over the baseline method on PASCAL VOC 2007 and 2012 datasets. Furthermore, two methods are complementary and significantly outperform the previous state-of-the-art when combined. Yuting Zhang 0001, Kihyuk Sohn, Ruben Villegas, Gang Pan 0001, Honglak Lee |
CVPR | 4 |
| 2015 | Bike sharing station placement leveraging heterogeneous urban open dataabstractBike sharing systems have been deployed in many cities to promote green transportation and a healthy lifestyle. One of the key factors for maximizing the utility of such systems is placing bike stations at locations that can best meet users' trip demand. Traditionally, urban planners rely on dedicated surveys to understand the local bike trip demand, which is costly in time and labor, especially when they need to compare many possible places. In this paper, we formulate the bike station placement issue as a bike trip demand prediction problem. We propose a semi-supervised feature selection method to extract customized features from the highly variant, heterogeneous urban open data to predict bike trip demand. Evaluation using real-world open data from Washington, D.C. and Hangzhou shows that our method can be applied to different cities to effectively recommend places with higher potential bike trip demand for placing future bike stations. Longbiao Chen, Daqing Zhang 0001, Gang Pan 0001, Xiaojuan Ma, Dingqi Yang, Kostadin Kushlev, Wangsheng Zhang, Shijian Li |
UbiComp | 3 |
| 2015 | Accelerometer-Based Gait Recognition by Sparse Representation of Signature Points With ClustersabstractGait, as a promising biometric for recognizing human identities, can be nonintrusively captured as a series of acceleration signals using wearable or portable smart devices. It can be used for access control. Most existing methods on accelerometer-based gait recognition require explicit step-cycle detection, suffering from cycle detection failures and intercycle phase misalignment. We propose a novel algorithm that avoids both the above two problems. It makes use of a type of salient points termed signature points (SPs), and has three components: 1) a multiscale SP extraction method, including the localization and SP descriptors; 2) a sparse representation scheme for encoding newly emerged SPs with known ones in terms of their descriptors, where the phase propinquity of the SPs in a cluster is leveraged to ensure the physical meaningfulness of the codes; and 3) a classifier for the sparse-code collections associated with the SPs of a series. Experimental results on our publicly available dataset of 175 subjects showed that our algorithm outperformed existing methods, even if the step cycles were perfectly detected for them. When the accelerometers at five different body locations were used together, it achieved the rank-1 accuracy of 95.8% for identification, and the equal error rate of 2.2% for verification. Yuting Zhang 0001, Gang Pan 0001, Kui Jia, Minlong Lu, Yueming Wang 0001, Zhaohui Wu 0001 |
IEEE Trans. Cybern. | 2 |
| 2015 | City-Scale Social Event Detection and Evaluation with Taxi TracesabstractA social event is an occurrence that involves lots of people and is accompanied by an obvious rise in human flow. Analysis of social events has real-world importance because events bring about impacts on many aspects of city life. Traditionally, detection and impact measurement of social events rely on social investigation, which involves considerable human effort. Recently, by analyzing messages in social networks, researchers can also detect and evaluate country-scale events. Nevertheless, the analysis of city-scale events has not been explored. In this article, we use human flow dynamics, which reflect the social activeness of a region, to detect social events and measure their impacts. We first extract human flow dynamics from taxi traces. Second, we propose a method that can not only discover the happening time and venue of events from abnormal social activeness, but also measure the scale of events through changes in such activeness. Third, we extract traffic congestion information from traces and use its change during social events to measure their impact. The results of experiments validate the effectiveness of both the event detection and impact measurement methods. Wangsheng Zhang, Guande Qi, Gang Pan 0001, Hua Lu 0001, Shijian Li, Zhaohui Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2015 | TripPlanner: Personalized Trip Planning Leveraging Heterogeneous Crowdsourced Digital FootprintsabstractPlanning an itinerary before traveling to a city is one of the most important travel preparation activities. In this paper, we propose a novel framework called TripPlanner, leveraging a combination of location-based social network (i.e., LBSN) and taxi GPS digital footprints to achieve personalized, interactive, and traffic-aware trip planning. First, we construct a dynamic point-of-interest network model by extracting relevant information from crowdsourced LBSN and taxi GPS traces. Then, we propose a two-phase approach for personalized trip planning. In the route search phase, TripPlanner works interactively with users to generate candidate routes with specified venues. In the route augmentation phase, TripPlanner applies heuristic algorithms to add user's preferred venues iteratively to the candidate routes, with the objective of maximizing the route score while satisfying both the venue visiting time and total travel time constraints. To validate the efficiency and effectiveness of the proposed approach, extensive empirical studies were performed on two real-world data sets from the city of San Francisco, which contain more than 391 900 passenger delivery trips generated by 536 taxis in a month and 110 214 check-ins left by 15 680 Foursquare users in six months. Chao Chen 0004, Daqing Zhang 0001, Bin Guo 0001, Xiaojuan Ma, Gang Pan 0001, Zhaohui Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Understanding Taxi Service Strategies From Taxi GPS TracesabstractTaxi service strategies, as the crowd intelligence of massive taxi drivers, are hidden in their historical time-stamped GPS traces. Mining GPS traces to understand the service strategies of skilled taxi drivers can benefit the drivers themselves, passengers, and city planners in a number of ways. This paper intends to uncover the efficient and inefficient taxi service strategies based on a large-scale GPS historical database of approximately 7600 taxis over one year in a city in China. First, we separate the GPS traces of individual taxi drivers and link them with the revenue generated. Second, we investigate the taxi service strategies from three perspectives, namely, passenger-searching strategies, passenger-delivery strategies, and service-region preference. Finally, we represent the taxi service strategies with a feature matrix and evaluate the correlation between service strategies and revenue, informing which strategies are efficient or inefficient. We predict the revenue of taxi drivers based on their strategies and achieve a prediction residual as less as 2.35 RMB/h,1which demonstrates that the extracted taxi service strategies with our proposed approach well characterize the driving behavior and performance of taxi drivers. Daqing Zhang 0001, Lin Sun 0009, Bin Li 0015, Chao Chen 0004, Gang Pan 0001, Shijian Li, Zhaohui Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2014 | Container throughput estimation leveraging ship GPS traces and open dataabstractTraditionally, the port container throughput, a crucial measurement of regional economic development, was manually collected by port authorities. This requires a large amount of human effort and often delays publication of this important figure. In this paper, by leveraging ubiquitous positioning techniques and open data, we propose a two-phase approach to estimation of port container throughput in real-time. First, we obtain the number of container ships arriving at berth by analyzing the ships' GPS traces. Then we estimate the throughput of each ship, in terms of number of containers transshipped, by considering the ship's berthing time, capacity, length, breadth, and crane operation performance, as extracted from different data sources. Evaluation results using real-world datasets from Hong Kong and Singapore show that the proposed approach not only estimates the container throughput quite accurately, but also outperforms the baseline method significantly. Longbiao Chen, Daqing Zhang 0001, Gang Pan 0001, Leye Wang, Xiaojuan Ma, Chao Chen 0004, Shijian Li |
UbiComp | 3 |
| 2014 | High-fidelity compression of extracellular recordings from motor cortexabstractIn invasive brain-machine interfaces (BMI), the recorded high-quality neural signals produce a large data volume. This calls for effective compression. In this paper, we focus on extracellular recording of motor cortex. First the characteristics of the signals are studied, one of which is that peaks of DCT coefficients at high frequency may correspond to spike firing patterns. Based on these characteristics, we propose a high-fidelity compression framework for these signals. The DCT coefficients of the signal are divided into two parts according to amplitude, rather than frequency. The Low-Amplitude-Component (LAC) is encoded by a phase called Symbol Encoding, which helps to reduce overall distortion. The High-Amplitude-Component (HAC), containing major information and spikes, is encoded by another phase called Hybrid Encoding. It combines the Huffman encoding and a novel Zero-Length-Encoding. Experiments show that the algorithm achieves a compression ratio of 18% without obvious distortion. Moreover, spikes are reserved more than 92%, outperforming existing work. Our algorithm enables low-cost storage devices to store long-time neural signals. Rachel Zhang, Gang Pan 0001, Yueming Wang 0001, Zhenfang Hu |
IJCNN | 2 |
| 2014 | Decoding motor cortical activities of Monkey: A datasetabstractMotor brain-machine interface (BMI) has great potentials in neural motor prostheses and has received increasing attention during the past decades in the neural engineering field. It requires an approach to decode neural activities that represents desired movements. Much of the progress in decoding algorithms has been driven by the availability of neural data, e.g. spike trains, in some research groups having animal laboratories and capable of performing surgery and building BMI systems. However, researchers in the neural signal processing field often face a dilemma of lacking neural data. To continue the innovation in decoding algorithms, this paper introduces a public neural dataset, the ZJU Neural Decoding Dataset (ZJUNDD). We give the detailed paradigm of the BMI system on monkey, including the experimental setup and the collection of 96-channel motor cortical activities. The dataset contains spike rates of neurons obtained by a consistent spike sorting method. To improve the data quality and reduce outliers, the spike data are carefully selected according to the quality of hand movements of the monkey. A standard protocol is provided for the assessment of decoding algorithms on the dataset, including the partition of training and testing sets, and the evaluation metrics. We also build an online evaluation system in order to enable a fair comparison between decoding approaches. Further, we benchmark several existing algorithms, which provides a basic performance of the methods. To the best of our knowledge, this is the first public dataset of spike trains for the decoding research of motor cortical activities. Luoqing Zhou, Yueming Wang 0001, Gang Pan 0001, Yiwen Wang 0002, Xiaoxiang Zheng, Zhaohui Wu 0001 |
IJCNN | 4 |
| 2014 | L1-norm latent SVM for compact features in object detection
Gang Pan 0001, Yueming Wang 0001, Yuting Zhang 0001, Zhaohui Wu 0001 |
Neurocomputing | 2 |
| 2014 | Counting moving people in crowds using motion statistics of feature-points
Mahdi Hashemzadeh, Gang Pan 0001 |
Multim. Tools Appl. | 2 |
| 2014 | Facial expression recognition based on meta probability codes
Nacer Farajzadeh, Gang Pan 0001, Zhaohui Wu 0001 |
Pattern Anal. Appl. | 2 |
| 2014 | Collaborative Policy AdministrationabstractPolicy-based management is a very effective method to protect sensitive information. However, the overclaim of privileges is widespread in emerging applications, including mobile applications and social network services, because the applications' users involved in policy administration have little knowledge of policy-based management. The overclaim can be leveraged by malicious applications, then lead to serious privacy leakages and financial loss. To resolve this issue, this paper proposes a novel policy administration mechanism, referred to as collaborative policy administration (CPA for short), to simplify the policy administration. In CPA, a policy administrator can refer to other similar policies to set up their own policies to protect privacy and other sensitive information. This paper formally defines CPA and proposes its enforcement framework. Furthermore, to obtain similar policies more effectively, which is the key step of CPA, a text mining-based similarity measure method is presented. We evaluate CPA with the data of Android applications and demonstrate that the text mining-based similarity measure method is more effective in obtaining similar policies than the previous category-based method. Weili Han, Zheran Fang, Laurence T. Yang, Gang Pan 0001, Zhaohui Wu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Generating fluent tubes in video synopsisabstractVideo synopsis is one of the effective techniques to build a short video representation preserving the essential activities for a long video. Existing methods usually have the problem that a continuous activity (tube) from a single moving object is separated to a few small pieces. In this paper, two schemes are proposed to generate fluent tubes for video synopsis. The Gaussian mixture model and a texture method are combined to detect more compact foreground with shadow removed. The foreground constitutes a set of initial trajectories. A particle filter tracker is used to concatenate two trajectories if they belong to the same foreground activity, which generates more fluent tubes for video synopsis. Experimental results on 4 videos show that our method produces better accuracies and visual effects in video synopsis. Minlong Lu, Yueming Wang 0001, Gang Pan 0001 |
ICASSP | 3 |
| 2013 | Online Community Detection for Large Complex Networks
Wangsheng Zhang, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li |
IJCAI | 2 |
| 2013 | Efficient computation of histograms on densely overlapped polygonal regions
Yuting Zhang 0001, Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001 |
Neurocomputing | 3 |
| 2013 | Combining velocity and Location-Specific Spatial Clues in Trajectories for Counting Crowded Moving ObjectsabstractTrajectory-clustering-based methods have shown a good performance in counting moving objects in densely crowded scenes. However, they still fall into trouble in complex scenes, such as with the close proximity of moving objects, freely moving parts of objects, and different object size in different locations of the scene. This paper proposes a new method combining velocity and location-specific spatial clues in trajectories to deal with these problems. We first extract the velocities of a trajectory over its life-time. To alleviate confusion around the boundary regions between close objects, extracted velocity information is utilized to eliminate unreal-world feature points on objects' boundaries. Then, a function is introduced to measure the similarity of the trajectories integrating both of the spatial and the velocity clues. This function is employed in the Mean-Shift clustering procedure to reduce the effect of freely moving parts of the objects. To address the problem of various object sizes in different regions of the scene, we suggest a technique to learn the location-specific size distribution of objects in different locations of a scene. The experimental results show that our proposed method achieves a good performance. Compared with other trajectory-clustering-based methods, it decreases the counting error rate by about 10%. Mahdi Hashemzadeh, Gang Pan 0001, Yueming Wang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Establishing Point Correspondence of 3D Faces Via Sparse Facial Deformable ModelabstractEstablishing a dense vertex-to-vertex anthropometric correspondence between 3D faces is an important and fundamental problem in 3D face research, which can contribute to most applications of 3D faces. This paper proposes a sparse facial deformable model to automatically achieve this task. For an input 3D face, the basic idea is to generate a new 3D face that has the same mesh topology as a reference face and the highly similar shape to the input face, and whose vertices correspond to those of the reference face in an anthropometric sense. Two constraints: 1) the shape constraint and 2) correspondence constraint are modeled in our method to satisfy the three requirements. The shape constraint is solved by a novel face deformation approach in which a normal-ray scheme is integrated to the closest-vertex scheme to keep high-curvature shapes in deformation. The correspondence constraint is based on an assumption that if the vertices on 3D faces are corresponded, their shape signals lie on a manifold and each face signal can be represented sparsely by a few typical items in a dictionary. The dictionary can be well learnt and contains the distribution information of the corresponded vertices. The correspondence information can be conveyed to the sparse representation of the generated 3D face. Thus, a patch-based sparse representation is proposed as the correspondence constraint. By solving the correspondence constraint iteratively, the vertices of the generated face can be adjusted to correspondence positions gradually. At the early iteration steps, smaller sparsity thresholds are set that yield larger representation errors but better globally corresponded vertices. At the later steps, relatively larger sparsity thresholds are used to encode local shapes. By this method, the vertices in the new face approach the right positions progressively until the final global correspondence is reached. Our method is automatic, and the manual work is needed only in training procedure. The experimental results on a large-scale publicly available 3D face data set, BU-3DFE, demonstrate that our method achieves better performance than existing methods. Gang Pan 0001, Yueming Wang 0001, Zhenfang Hu, Xiaoxiang Zheng, Zhaohui Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2013 | Land-Use Classification Using Taxi GPS TracesabstractDetailed land use, which is difficult to obtain, is an integral part of urban planning. Currently, GPS traces of vehicles are becoming readily available. It conveys human mobility and activity information, which can be closely related to the land use of a region. This paper discusses the potential use of taxi traces for urban land-use classification, particularly for recognizing the social function of urban land by using one year's trace data from 4000 taxis. First, we found that pick-up/set-down dynamics, extracted from taxi traces, exhibited clear patterns corresponding to the land-use classes of these regions. Second, with six features designed to characterize the pick-up/set-down pattern, land-use classes of regions could be recognized. Classification results using the best combination of features achieved a recognition accuracy of 95%. Third, the classification results also highlighted regions that changed land-use class from one to another, and such land-use class transition dynamics of regions revealed unusual real-world social events. Moreover, the pick-up/set-down dynamics could further reflect to what extent each region is used as a certain class. Gang Pan 0001, Guande Qi, Zhaohui Wu 0001, Daqing Zhang 0001, Shijian Li |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2012 | FlyingBuddy2: a brain-controlled assistant for the handicappedabstractThe motor impaired people have much limit in moving. The devices augmenting their mobility will be much helpful for improving their living experiences. This poster develops a brain-controlled assistive system, called FlyingBuddy2, to aid the handicapped in mobility. It uses the brain EEG signals to directly control a quadrotor. Signals from an EEG headset are transmitted wirelessly to a computer, then the decoded brain signals are converted to trigger the quadrotor to move in 3D space. Three applications are developed: thinking to play games, thinking to see, and thinking to take pictures. Yipeng Yu, Weidong Hua, Shijian Li, Yueming Wang 0001, Gang Pan 0001 |
UbiComp | 7 |
| 2012 | Mining the semantics of origin-destination flows using taxi tracesabstractOrigin-destination(OD) flows reflect both human activity and urban dynamic in a city. However, our understanding about their patterns remains limited. In this paper, we study the GPS traces of taxis in a city with several millions people, China and find that there are significant patterns under the OD flows constructed from taxis' random motion. Our spatiotemporal analysis shows that those patterns have close relationship with the semantics of OD flows, hence we can mine the semantics of OD flows from raw GPS trace data. The approach we proposed offers a novel way to explore the human mobility and location characteristic. Wangsheng Zhang, Shijian Li, Gang Pan 0001 |
UbiComp | 3 |
| 2012 | SmartShadow-K: an practical knowledge network for joint context inference in everyday lifeabstractSmart environments require to percept conditions of people. Current context-aware systems mainly model limited user situations, which constrains their coverage and effect in real world usage. This paper proposes an encyclopedic knowledge network to enable practical context inference in our daily life by: 1) expressing essential semantics of contextual concepts and relations into a well-informed relational network, and 2) exploiting relational semantics to infer various contexts simultaneously. The performance of the approach is validated in real challenging problems and compared with inference of human being. Li Zhang 0045, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li, Cho-Li Wang |
UbiComp | 2 |
| 2012 | Prediction of urban human mobility using large-scale taxi traces and its applications
Gang Pan 0001, Zhaohui Wu 0001, Guande Qi, Shijian Li, Daqing Zhang 0001, Wangsheng Zhang, Zonghui Wang |
Frontiers Comput. Sci. China | 2 |
| 2011 | Tilt & touch: mobile phone for 3D interactionabstractMobile phones are becoming de facto pervasive devices for people's daily use. This demonstration illustrates a new interaction, Tilt & Touch, to enable a smart phone to be a 3D controller. It exploits capacitive touchscreen and built-in MEMS motion sensors. When people want to navigate in a virtual reality environment on a large display, they can tilt the phone for viewpoint transforming, touch the phone screen for avatar moving, and pinch screen for viewing camera zooming. The virtual objects in the virtual reality environment can be rotated accordingly by tilting the phone. Yuan Du, Haoyi Ren, Gang Pan 0001, Shijian Li |
UbiComp | 3 |
| 2011 | FlyingBuddy: augment human mobility and perceptibilityabstractTechnologies keep evolving to strengthen and further people's abilities in many aspects. For instance, vehicles expand the range of human moving while mobile phones boost the range of human communication. In this video, we develop a novel mini unmanned aerial vehicle (mini-UAV) named FlyingBuddy to augment human mobility and perceptibility. This prototype is made up of off the shelf components AR. Drone and iPhones with customized software. With help of the built-in magnetometer, GPS, and cameras, as well as Bluetooth, Wi-Fi and 3G connectivity, FlyingBuddy is capable of both manual controlled and self-piloted flying. It provides four typical services: flying to buy, flying to see, flying to report accident, and flying to take pictures. Haoyi Ren, Weidong Hua, Gang Pan 0001, Shijian Li, Zhaohui Wu 0001 |
UbiComp | 4 |
| 2011 | Multiclass Classification Based on Meta Probability CodesabstractThis paper proposes a new approach to improve multiclass classification performance by employing Stacked Generalization structure and One-Against-One decomposition strategy. The proposed approach encodes the outputs of all pairwise classifiers by implicitly embedding two-class discriminative information in a probabilistic manner. The encoded outputs, called Meta Probability Codes (MPCs), are interpreted as the projections of the original features. It is observed that MPC, compared to the original features, has more appropriate features for clustering. Based on MPC, we introduce a cluster-based multiclass classification algorithm, called MPC-Clustering. The MPC-Clustering algorithm uses the proposed approach to project an original feature space to MPC, and then it employs a clustering scheme to cluster MPCs. Subsequently, it trains individual multiclass classifiers on the produced clusters to complete the procedure of multiclass classifier induction. The performance of the proposed algorithm is extensively evaluated on 20 datasets from the UCI machine learning database repository. The results imply that MPC-Clustering is quite efficient with an improvement of 2.4% overall classification rate compared to the state-of-the-art multiclass classifiers. Nacer Farajzadeh, Gang Pan 0001, Zhaohui Wu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2011 | A deformation model to reduce the effect of expressions in 3D face recognition
Yueming Wang 0001, Gang Pan 0001, Jianzhuang Liu |
Vis. Comput. | 2 |
| 2010 | Removal of 3D facial expressions: A learning-based approachabstractThis paper focuses on the task of recovering the neutral 3D face of a person when given his/her 3D face model with facial expression. We propose a learning-based expression removal framework to tackle this task. Our basic idea is to model expression residue from samples, and then use the inferred expression residue from the input expressional face model to recover the neutral one. A two-step non-rigid alignment method is introduced to make all the face models topologically share a common structure. Then we construct two spaces, normal space and expression residue space, for modeling expression. Therefore, the expression removal problem can be formalized as the inference of expression residue from normal spaces. The neutral face model can be generated in a Poisson-based framework by the inferred expression residue. The experimental results on BU-3DFED database demonstrate the effectiveness of our approach. Gang Pan 0001, Zhaohui Wu 0001, Yuting Zhang 0001 |
CVPR | 1 |
| 2010 | Semantic Device Bus for Internet of ThingsabstractThe vision of the Internet of things is very appealing, and gains more and more attention. Since mobile devices in the Internet become more complex and heterogeneous, device collaboration will be full of technical challenges. How to integrate different systems and heterogeneous devices is a big problem. In order to overcome the problem, it is very important to describe and match the heterogeneous device services with semantics. We use web service interface specification to wrap device, and OWL to describe services in semantic. In this paper we introduce a semantic device bus for Internet of things, which will allow service creation of different devices, management of device services, semantic description and matching of device services, and complex collaboration of device services. The bus provides a fundamental platform for large-scale applications of the Internet of things. Shijian Li, Li Zhang 0045, Gang Pan 0001 |
EUC | 4 |
| 2010 | Modeling Files with Context Streams
Qunjie Qiu, Gang Pan 0001, Shijian Li |
UIC | 2 |
| 2010 | GeeAir: a universal multimodal remote control device for home appliances
Gang Pan 0001, Daqing Zhang 0001, Zhaohui Wu 0001, Yingchun Yang, Shijian Li |
Pers. Ubiquitous Comput. | 1 |
| 2010 | Infrastructure and Reliability Analysis of Electric Networks for E-TextilesabstractElectronic textiles (e-textiles), known as computational fabrics, offer an emerging platform for constructing ambient intelligent applications. Computational nodes in e-textiles are driven by batteries. Unlike wireless sensor networks, not each computational node in e-textiles has its own battery. Instead, many computational nodes in e-textiles share a battery. Existing e-textiles use one fixed battery to drive a fixed set of computation nodes (or power consuming electronic components). The fixed battery-component connection may result in electronic components stopping functioning and/or energy waste in batteries when link connection problems occur. In this paper, we propose a new infrastructure of the power networks for e-textiles: flexible power network (FPN). Under the FPN infrastructure, a power consuming node (PCN) is not just connected to one single fixed battery. Instead, it is connected to multiple batteries and can obtain power energy from one of the available battery nodes (BNs) with the help of a battery selector. The electrical features of battery selectors and overcurrent protectors that protect the batteries from wasting the charge when short-circuit faults occur are illustrated. Moreover, by modeling the number of fault occurrence at conductive wires and nodes stochastically, an evaluation algorithm is proposed to analyze the reliability of FPN and to compare the metrics of different design schemes under the perspective of both the BNs and the PCNs. Experimental results show that our FPN is more dependable than some common e-textile electric networks published before with the occurrence of short- and/or open-circuit faults. Nenggan Zheng, Zhaohui Wu 0001, Man Lin, Laurence T. Yang, Gang Pan 0001 |
IEEE Trans. Syst. Man Cybern. Part C | 5 |
| 2009 | SmartShadow: Modeling A User-centric Mobile Virtual SpaceabstractThis paper attempts to model pervasive computing environments as a user-centric ldquoSmartShadowrdquo using the BDP (belief-desire-plan) user model, which maps pervasive computing environments into a dynamic virtual user space. SmartShadow will follow the user to provide him with pervasive services, just like his shadow in the physical world. In the BDP model, desires of a user are inferred from his belief set, and plans are made to satisfy each desire. Pervasive service is introduced to describe computing resources in the cyberspace, which can be organized by the user's BDP to accomplish his desires. The composition process maps pervasive services into a user's SmartShadow. The model is logically natural and simple, and can flexibly model dynamics of pervasive computing spaces. In addition, we implement a simulation system to verify and evaluate the SmartShadow model. Li Zhang 0045, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li, Cho-Li Wang |
PerCom | 2 |
| 2009 | Gesture Recognition with a 3-D Accelerometer
Gang Pan 0001, Daqing Zhang 0001, Guande Qi, Shijian Li |
UIC | 2 |
| 2009 | A Smart Car Control Model for Brake Comfort Based on Car FollowingabstractThis paper demonstrates a novel car-following model focused on passenger comfort, for example, a rapid deceleration will make passengers uncomfortable. The brake comfort model of car following was set up according to the relationship between vehicle deceleration and passenger comfort levels. The model calculates the controlled car's acceleration by measuring the distance between the controlled car and its preceding car, as well as the velocity of the controlled car. By controlling the car's acceleration, the model is able to keep riders feeling comfortable. The friction coefficient between the car and the road surface is also considered. Experiments show that the model is highly compatible with real cases. Zhaohui Wu 0001, Gang Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2008 | Hallucinating 3D facial shapesabstractThis paper focuses on hallucinating a facial shape from a low-resolution 3D facial shape. Firstly, we give a constrained conformal embedding of 3D shape in R2, which establishes an isomorphic mapping between curved facial surface and 2D planar domain. With such conformal embedding, two planar representations of 3D shapes are proposed:Gaussiancurvatureimage(GCI) for a facial surface, andsurfacedisplacementimage(SDI) for a pair of facial surfaces. The conformal planar representation reduces the data complexity from 3D irregular curved surface to 2D regular grid while preserving the necessary information for hallucination. Then, hallucinating a low resolution facial shape is formalized as inference of SDI from GCIs by modeling the relationship between GCI and SDI by RBF regression. The experiments on USF HumanID 3D face database demonstrate the effectiveness of the approach. Our method can be easily extended to hallucinate those category-specific 3D surfaces sharing with similar geometric structures. Gang Pan 0001, Zhaohui Wu 0001 |
CVPR | 1 |
| 2008 | 3D Face Recognition by Local Shape Difference Boosting
Yueming Wang 0001, Xiaoou Tang, Jianzhuang Liu, Gang Pan 0001, Rong Xiao 0003 |
ECCV (1) | 4 |
| 2007 | 3D Face Recognition in the Presence of Expression: A Guidance-based Constraint Deformation ApproachabstractThree-dimensional human face recognition in the presence of expression is a big challenge, since the shape distortion caused by facial expression greatly weakens the rigid matching. This paper proposes a guidance-based constraint deformation(GCD) model to cope with the shape distortion by expression. The basic idea is that, the face model with non-neutral expression is deformed toward its neutral one under certain constraint so that the distortion is reduced while inter-class discriminative information is preserved. The GCD model exploits the neutral 3D face shape to guide the deformation, meanwhile applies a rigid constraint on it. Both steps are smoothly unified in the Poisson equation framework. The GCD approach only needs one neutral model for each person in the gallery. The experimental results, carried out on the large 3D face databases-FRGC v2.0, demonstrate that our method significantly outperforms ICP method for both identification and authentication mode. It shows the GCD model is promising for coping with the shape distortion in 3D face recognition. Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001 |
CVPR | 2 |
| 2007 | Eyeblink-based Anti-Spoofing in Face Recognition from a Generic WebcameraabstractWe present a real-time liveness detection approach against photograph spoofing in face recognition, by recognizing spontaneous eyeblinks, which is a non-intrusive manner. The approach requires no extra hardware except for a generic webcamera. Eyeblink sequences often have a complex underlying structure. We formulate blink detection as inference in an undirected conditional graphical framework, and are able to learn a compact and efficient observation and transition potentials from data. For purpose of quick and accurate recognition of the blink behavior, eye closity, an easily-computed discriminative measure derived from the adaptive boosting algorithm, is developed, and then smoothly embedded into the conditional model. An extensive set of experiments are presented to show effectiveness of our approach and how it outperforms the cascaded Adaboost and HMM in task of eyeblink detection. Gang Pan 0001, Lin Sun 0009, Zhaohui Wu 0001, Shihong Lao |
ICCV | 1 |
| 2007 | ScudWare: A Semantic and Adaptive Middleware Platform for Smart Vehicle SpaceabstractCompared with smart spaces addressed previously, the smart vehicle space is quite special. First, unlike room, it is a high mobile space. Second, it requires frequent information exchange with the outer environment; for instance, it may need local traffic information and other local services. The complex vehicle space needs a software infrastructure of high adaptation to meet the complex and easy variational situations. This paper proposes a semantic and adaptive middleware platform, i.e., ScudWare, for smart vehicle space. In ScudWare, techniques of multiagent, context-aware, and adaptive component management are smoothly synthesized. It consists of three key components, namely 1) semantic virtual agents; 2) a semantic context management service; and 3) an adaptive component management service. This enables entities in smart vehicle space to interact autonomously, provides semantic-integration context awareness, and supports component-based adaptability and scalability. A test bed of smart car space is built to evaluate the ScudWare, which demonstrates the effectiveness and efficiency of ScudWare Zhaohui Wu 0001, Gang Pan 0001, Minde Zhao, Jie Sun 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2006 | Hallucinating 3D Faces
Shiqi Peng, Gang Pan 0001, Shi Han, Yueming Wang 0001 |
ACCV (2) | 2 |
| 2006 | Exploring Facial Expression Effects in 3D Face Recognition Using Partial ICP
Yueming Wang 0001, Gang Pan 0001, Zhaohui Wu 0001, Yigang Wang |
ACCV (1) | 2 |
| 2006 | Super-Resolution of 3D Face
Gang Pan 0001, Shi Han, Zhaohui Wu 0001, Yueming Wang 0001 |
ECCV (2) | 1 |
| 2005 | Learning-based super-resolution of 3D face modelabstractSuper resolution technique could produce a higher resolution image than the originally captured one. However, nearly all super-resolution algorithms arm at 2D images. In this paper, we focus on generating the 3D face model of higher resolution from one input of 3D face model. In our method, the 3D face models firstly are all regularized via resampling in cylindrical representation. The super resolution then performs in the regular domain of cylindrical coordinate. The experiments using USF HumanID 3D face database of 137 3D face models are carried out, and demonstrate the presented algorithm is promising. Shiqi Peng, Gang Pan 0001, Zhaohui Wu 0001 |
ICIP (2) | 2 |
| 2004 | 3d face recognition using local shape map
Zhaohui Wu 0001, Yueming Wang 0001, Gang Pan 0001 |
ICIP | 3 |
| 2003 | Automatic 3D face verification from range dataabstractWe present a novel approach for automatic 3D face verification from range data. The method consists of range data registration and comparison. There are two steps in the registration procedure: a coarse step conducting normalization by exploiting a priori knowledge of the human face and facial features; a fine step aligning the input data with the model stored in the database by the partial directed Hausdorff distance. To speed up the registration, a simplified version of the model is generated for each model in the model database. During the face comparison, the partial Hausdorff distance is employed as the similarity metric. Experiments have been carried out on a database with 30 individuals, and the best EER of 3.24% is achieved. Gang Pan 0001, Zhaohui Wu 0001, Yunhe Pan |
ICASSP (3) | 1 |
| 2003 | Automatic 3D face verification from range dataabstractIn this paper, we presented a novel approach for automatic 3D face verification from range data. The method consists of range data registration and comparison. There are two steps in registration procedure: the coarse step conducting the normalization by exploiting a priori knowledge of the human face and facial features, and the fine step aligning the input data with the model stored in the database by the partial directed Hausdorff distance. To speed up the registration, a simplified version of the model is generated for each model in the model database. During the face comparison, the partial Hausdorff distance is employed as the similarity metric. The experiments are carried out on a database with 30 individuals and the best EER of 3.24% is achieved. Gang Pan 0001, Zhaohui Wu 0001, Yunhe Pan |
ICME | 1 |
| 2003 | 3D face recognition by profile and surface matchingabstractIn this paper, we presented an approach for automatic face verification from range data. The method consists of profile and surface matching. The profile is extracted on the basis of symmetry of human face, and a global profile matching method based on k-th Hausdorff distance is used to align and compare profiles, without detection of fiducial points that is often unreliable. For each individual, a statistical model of facial surface is built to represent the distinct discriminative capability of the different parts in the facial surface. Then the model is incorporated into a weighted distance function to measure similarity of surfaces. Finally two experts are combined to give a decision. The comparable experimental results are obtained on a database with 180 pieces of range data of 30 individuals. Gang Pan 0001, Yijun Wu, Zhaohui Wu 0001, Wenyao Liu |
IJCNN | 1 |
| 2003 | Investigating profile extracted from range data for 3D face recognitionabstractIn this paper we investigate the discriminative capability of facial profile extracted from range data for 3D face recognition. A robust symmetry plane detection method is proposed for profile extracting. A global profile matching approach based on k-th Hausdorff metric is presented to align and compare profiles, without detection of fiducial points that is often unreliable. The experiment are carried out on a low-quality database with 180 pieces of range data of 30 individuals acquired by a structured light system. Based on the experimental results, we observe that the profile is quite valuable information in 3D face recognition, and the best EER of 2.22% is achieved. Gang Pan 0001, Yijun Wu, Zhaohui Wu 0001 |
SMC | 1 |
| 2003 | Pose-invariant detection of facial features from range dataabstractThis paper, firstly, presents a new signature representation for point in range data, called curgram. Curgram serves to describe the structural neighborhood of a point and establish a signature for each of the given 3D data points rather than information of this point position. It yields invariance under rigid motions and mirror imaging. Secondly, its application to detection of facial features from range data is performed by incorporating the metric for curgram into the conventional appearance-based face detection framework. Experimental results with fifteen face range images have demonstrated the validity and effectiveness of the proposed method. Gang Pan 0001, Yueming Wang 0001, Zhaohui Wu 0001 |
SMC | 1 |
| 2002 | A data hiding method for few-color imagesabstractThe few-color images are often synthetic graphics without complicated color and texture variation, which makes the embedding of invisible digital watermark difficult. This paper proposes a data hiding method that can hide a moderate amount of data in a few-color image, such as cartoon images, binary images, without introducing noticeable artifacts. To achieve least visual quality reduction, the prioritized pattern matching scheme is employed to embed the invisible data in the pixels those are close to the boundaries. The block permutation is also exploited before embedding. No additional color is introduced and the palette keeps unchanged after embedding. Extracting of the hidden data does not require the knowledge of the original image. The experimental results show that the approach achieves a quite great performance in visibility transparency. It is applicable to invisible annotation, covert communication etc. Gang Pan 0001, Zhaohui Wu 0001, Yunhe Pan |
ICASSP | 1 |
| 2001 | A Novel Data Hiding Method for Two-Color Images
Gang Pan 0001, Yijun Wu, Zhaohui Wu 0001 |
ICICS | 1 |