Xiaoyue Ji

dblp:297/5250 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-3526-5215ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 FE-SpikeFormer: A Camera-Based Facial Expression Recognition Method for Hospital Health Monitoring
abstract
Facial expression recognition has emerged as a critical research area in health monitoring, enabling healthcare professionals to assess patients' emotional and psychological states for timely intervention and personalized care. However, existing methods often struggle to balance computational accuracy with energy efficiency. To address this challenge, this paper proposes FE-SpikeFormer - a high-accuracy, low-energy, and deployment-friendly Spiking Neural Network (SNN) for facial emotion recognition. The proposed architecture comprises three key components: the initial convolution module, the spiking extraction block, and the spiking integration block. These three modules collectively support detailed and contextual feature extraction, promote spatial feature integration, and strengthen the representational capacity of spiking signals. Meanwhile, a joint verification is conducted in both controlled laboratory settings and real-world hospital scenarios. Experimental results demonstrate that FE-SpikeFormer achieves top-three recognition accuracy among state-of-the-art methods, while utilizing only 6.93 million parameters. Moreover, it exhibits strong robustness against various noise conditions, underscoring its potential for practical deployment in healthcare environments.
Zhekang Dong, Shiqi Zhou, Xiaoyue Ji, Chun Sing Lai, Minjiang Chen, Jiansong Ji
IEEE J. Biomed. Health Informatics4
2026 An Efficient Human Activity Recognition In-Memory Computing Architecture Development for Healthcare Monitoring
abstract
Human activity recognition has played a crucial role in healthcare information systems due to the fast adoption of artificial intelligence (AI) and the internet of thing (IoT). Most of the existing methods are still limited by computational energy, transmission latency, and computing speed. To address these challenges, we develop an efficient human activity recognition in-memory computing architecture for healthcare monitoring. Specifically, a mechanism-oriented model of Ag/a-Carbon/Ag memristor is designed, serving as the core circuit component of the proposed in-memory computing system. Then, one-transistor-two-memristor (1T2M) crossbar array is proposed to perform high-efficiency multiply-accumulate (MAC) operation and high-density memory in the proposed scheme. To facilitate understanding of the proposed efficient human activity recognition in-memory computing design, self-attention ConvLSTM module, multi-head convolutional attention module, and recognition module are proposed. Furthermore, the proposed system is applied to perform human activity recognition, which contains eleven different human activities, including five different postural falls, and six basic daily activities. The experimental results show that the proposed system has advantages in recognition performance (≥ 0.20% accuracy, ≥ 1.10% F1-score) and time consumption (approximately 8∼10 times speed up) compared to existing methods, indicating an advancement in smart healthcare applications.
Xiaoyue Ji, Zhekang Dong, Chenhao Hu, Chun Sing Lai
IEEE J. Biomed. Health Informatics1
2025 A Dual-Pathway Driver Emotion Classification Network Using Multitask Learning Strategy: A Joint Verification
abstract
Negative emotion (e.g., anger, fear) may influence normal driver behavior, resulting in serious traffic accidents. Thus, developing an automatic driver emotion classification method is necessary and urgent. Most of the existing methods are performed in realistic indoor environment and always lack effective utilization of heterogeneous information, resulting in low accuracy and reliability. In this paper, a novel dual-pathway driver emotion classification network using multi-task learning strategy is proposed. To illustrate the design of the proposed driver emotion classification network, three modules are constructed: 1) visual-facial data processing module; 2) driving behavioral data processing module; 3) fusion output module. Meanwhile, considering the influence of emotional states on driving behavior, a comprehensive analysis is conducted to distinguish the positive, neutral, and negative influence on driving behavior. Furthermore, a joint verification in both realistic indoor environment (i.e., laboratory simulation on the PPB-Emo dataset) and real-world outdoor scenario is performed. The experimental results illustrate that the proposed network exhibits superior performance in terms of classification accuracy and response time, achieving good balance between classification accuracy and running speed in internet of things scenarios.
Zhekang Dong, Chenhao Hu, Xiaoyue Ji, Chun Sing Lai
IEEE Internet Things J.4
2024 Time-Frequency Hybrid Neuromorphic Computing Architecture Development for Battery State-of-Health Estimation
abstract
With the rapid adoption of Internet of Things (IoT) and artificial intelligence (AI), lithium-ion battery state-of-health (SOH) estimation plays an important role in guaranteeing the secure and stable functioning of various domains. However, the majority of the existing methods are constrained by factors, such as transmission latency, computational energy, and computing speed. To address these challenges, we develop a time-frequency hybrid neuromorphic computing architecture for battery SOH estimation. Specifically, an eco-friendly, biodegradable memristor crossbar array is designed, enabling high-energy efficiency and high-performance density in the proposed system. To improve the understanding of the designed time-frequency hybrid neuromorphic computing system, a local information extraction module, a time-frequency feature fusion module, and a global information perception module are proposed. Furthermore, the proposed system is validated on two publicly available battery ageing data sets (i.e., the CALCE-CS2 data set and the National Aeronautics and Space Administration data set). The experimental results show that the system exhibits superior performance to that of the state-of-the-art (SOTA) methods in terms of estimation accuracy (highest estimation accuracy), time consumption (approximately 8–12 times faster), and transmission latency (approximately 10 times faster). This study is expected to promote the advancement and evolution of next-generation computing systems, enabling the realization of low-power consumption and high-density information processing in IoT scenarios.
Xiaoyue Ji, Junfan Wang, Guangdong Zhou, Chun Sing Lai, Zhekang Dong
IEEE Internet Things J.1
2024 SpikeTOD: A Biologically Interpretable Spike-Driven Object Detection in Challenging Traffic Scenarios
abstract
Artificial neural networks (ANN) have shown remarkable performance in intelligent transportation systems (ITS), especially for the traffic object detection. However, as the ITS is applied to a wider range of traffic scenarios, the increasing demand for the trade-off between detection performance and power resources has become inevitable. A biologically interpretable spike-driven traffic object detector for challenging scenarios is proposed in this paper, named SpikeTOD, achieving the trade-off between the accuracy and power consumption. Firstly, the spike neural network (SNN) is employed to realize energy-efficient object detection in traffic scenarios. And a local modulation-based integrate-and-fire (IF) neuron is designed, which provides an efficient way to convert the traffic detection model from ANN to SNN. Secondly, a biology-inspired detail-guided context-aware network (DCNet) is proposed to improve the detection performance. The integration of detail coherence and global priors is leveraged to selectively emphasize object features and improve the detection capabilities within challenging conditions. As far as we know, this is the first application of SNN in traffic object detection tasks. SpikeTOD achieved a mAP@50 of 46.11% on the BDD100K dataset with a power consumption of 4.73E-03J, demonstrating a more efficient trade-off in detection accuracy and power consumption. Notably, SpikeTOD maintained an average missed detection rate of 44.56%, further contributing to its overall efficacy in traffic object detection. Further, we conducted on road test by deploying SpikeTOD on Jetson Xavier NX and Loihi to demonstrate that model achieves a better balance between accuracy and power consumption.
Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Zhiwei He 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Vehicle-Mounted Adaptive Traffic Sign Detector for Small-Sized Signs in Multiple Working Conditions
abstract
Traffic sign detection is of great significance to the development of the Intelligent Transportation System (ITS) as a database for environmental awareness. The main challenges of existing traffic sign detection method are inaccurate small object detection, difficult mobile deployment, and complex working environment. Based on these, a vehicle-mounted adaptive traffic sign detector (VATSD) for small-sized signs in multiple working conditions is proposed in this paper. First, the Backbone of the detector is optimized. A feature tight fusion structure is designed to constitute a new feature extraction module, DCSP, which improves the feature extraction capability and the detection accuracy of small objects with negligible additional parameters. Second, an image enhancement network IENet with an adaptive joint filtering strategy is proposed. The IENet enables the dynamic selection of filters and thus adaptively optimizes low-quality images under multiple conditions to improve the accuracy of subsequent detection tasks. The proposed method has experimented on three traffic sign datasets and the detection accuracy increased by up to 7.6% compared to the original. The proposed detector demonstrates superiority over other state-of-the-art (SOTA) methods in terms of small object detection accuracy, detection speed, and environmental adaptability. Further, we deployed VATSD to Jetson Xavier NX and achieved a detection speed of 21.6 FPS, meeting real-time requirements.
Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Chun Sing Lai
IEEE Trans. Intell. Transp. Syst.3
2024 Metaverse Meets Intelligent Transportation System: An Efficient and Instructional Visual Perception Framework
abstract
The combination of the Metaverse and intelligent transportation systems (ITS) holds significant developmental promise, especially for visual perception tasks. However, the acquisition of high-quality scene data poses a challenging and expensive endeavor. Meanwhile, the visual disparity between the Metaverse and the physical world poses an impact on the practical applicability of the visual perception tasks. In this paper, a Metaverse Intelligent Traffic Visual Framework, MITVF, is developed to guide the implementation of visual perception tasks in the physical world. Firstly, a two-stage metadata optimization strategy is proposed that can efficiently provide diverse and high-quality scene data for traffic perception models. Specifically, an element reconfigurability strategy is proposed to flexibly combine dynamic and static traffic elements to enrich the data with a low cost. A diffusion model-based metadata optimization acceleration strategy is proposed to achieve efficient improvement of image resolution. Secondly, a Meta-Physical adaptive learning method is proposed, and further applied to visual perception tasks to compensate for the visual disparity between the Metaverse and the physical world. Experimental results show that MITVF achieves a 10$\times$acceleration in optimization speed, ensuring the image quality and reconstructing diverse. Further, MITVF is applied to the traffic object detection task to verify the effectiveness and validity. The performance of the model trained with 5k real data exceeded that of the model trained with 200k real data, with AP$_{50}$reaching 67.7%.
Junfan Wang, Xiaoyue Ji, Zhekang Dong, Mingyu Gao 0002, Chun Sing Lai
IEEE Trans. Intell. Transp. Syst.3
2024 MLG-NCS: Multimodal Local-Global Neuromorphic Computing System for Affective Video Content Analysis
abstract
Despite neuromorphic computing (NC) technologies offer tremendous potential in executing computationally intensive tasks with high efficiency and low latency, most of existing methods are still difficult to achieve software-comparable accuracy. To address this challenge, we develop a multimodal local–global NC system (MLG-NCS) that can capture local characteristics and exchange global cross-modal information sufficiently. Specifically, a high-density memristor crossbar array is prepared to perform efficient parallel in-memory operations, serving as the fundamental component of the proposed MLG-NCS. To facilitate understanding of the proposed MLG-NCS design, the local feature representation module, the global cross-modal interaction module, and the output module are designed. The experimental results show that the proposed system has advantages in classification accuracy (ranked top three), time consumption (approximately ten times speed up), and latency (about 1.2–15.3 times faster), enabling good inter-related tradeoffs between latency, efficiency, and accuracy. This study is expected to promote the revolution and development of next-generation computing system, which takes a firm step toward artificial general intelligence (AGI).
Xiaoyue Ji, Zhekang Dong, Guangdong Zhou, Chun Sing Lai, Donglian Qi
IEEE Trans. Syst. Man Cybern. Syst.1
2023 A Brain-Inspired Hierarchical Interactive In-Memory Computing System and Its Application in Video Sentiment Analysis
abstract
Video sentiment analysis can effectively establish the relationship between the emotion state and the multimodal information, while still suffer from intensive computation and low efficiency, due to the von Neumann computing architecture. Here, we present a brain-inspired hierarchical interactive in-memory computing (IMC) system, which can efficiently solve ‘von Neumann bottleneck’, enabling cross-modal interactions and semantic gap elimination. First, a 1T1M synapse array is fabricated using cost-effective, highly stable, flexible, and eco-friendly carbon materials, offering efficient analog multiply-accumulate operations. To illustrate the complexity of the proposed brain-inspired hierarchical interactive IMC system, three modules are proposed: 1) unimodal extraction module, 2) hierarchical interactive module, 3) output module. Furthermore, the proposed system is validated by applying it to video sentiment analysis. The experimental results demonstrate that the proposed system outperforms the existing state-of-the-art methods with high computational efficiency and good robustness. This work opens up a new way to achieve the deep integration of nanomaterials, deep learning, and modern electronics into IMC.
Xiaoyue Ji, Zhekang Dong, Yifeng Han, Chun Sing Lai, Donglian Qi
IEEE Trans. Circuits Syst. Video Technol.1