Yifan Zhuang

dblp:299/4119 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0003-3732-5215ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Human-computer interaction and pervasive computing
1 paper
Wearable and physiological sensing · 100%
Artificial intelligence
1 paper
Speech recognition and synthesis · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Wearable and physiological sensing
brain-computer interface
1.012026
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026
Wearable and physiological sensing
electromyography
1.012026
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026
Natural language and speech › Speech recognition and synthesis › automatic speech recognition
tone recognition
0.312026
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding · AAAI 2026

Methods — techniques the papers use, named apart from their topics

multimodal fusion · 2.0domain adversarial training · 2.0cross-attention fusion · 2.0
YearPublicationVenuePosition
2026 CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding
abstract
Brain-computer interface (BCI) speech decoding has emerged as a promising tool for assisting individuals with speech impairments. In this context, the integration of electroencephalography (EEG) and electromyography (EMG) signals offers strong potential for enhancing decoding performance. Mandarin tone classification presents particular challenges, as tonal variations convey distinct meanings even when phonemes remain identical. In this study, we propose a novel cross-subject multimodal BCI decoding framework that fuses EEG and EMG signals to classify four Mandarin tones under both audible and silent speech conditions. Inspired by the cooperative mechanisms of neural and muscular systems in speech production, our neural decoding architecture combines spatial-temporal feature extraction branches with a cross-attention fusion mechanism, enabling informative interaction between modalities. We further incorporate domain-adversarial training to improve cross-subject generalization. We collected 4,800 EEG trials and 4,800 EMG trials from 10 participants using only twenty EEG and five EMG channels, demonstrating the feasibility of minimal-channel decoding. Despite employing lightweight modules, our model outperforms state-of-the-art baselines across all conditions, achieving average classification accuracies of 87.83\% for audible speech and 88.08\% for silent speech. In cross-subject evaluations, it still maintains strong performance with accuracies of 83.27\% and 85.10\% for audible and silent speech, respectively. We further conduct ablation studies to validate the effectiveness of each component. Our findings suggest that tone-level decoding with minimal EEG-EMG channels is feasible and potentially generalizable across subjects, contributing to the development of practical BCI applications.
Yifan Zhuang, Calvin Huang, Zepeng Yu, Yongjie Zou 0001, Jiawei Ju
AAAI1
2022 Emolleia - Wearable Kinetic Flower Display for Expressing Emotions
abstract
What we wear (our clothes and wearable accessories) can represent our mood at the moment. We developed Emolleia to explore how to make aesthetic wears more expressive to become a novel form of non-verbal communication to express our emotional feelings. Emolleia is an open wearable kinetic display in form of three 3D printed flowers that can dynamically open and close at different speeds. With our open-source platform, users can define their own animated motions. In this paper, we described the prototype design, hardware considerations, and user surveys (n=50) to evaluate the expressiveness of 8 pre-defined animated motions of Emolleia. Our initial results showed animated motions are feasible to communicate different emotional feelings especially at the valence and arousal dimensions. Based on the findings, we mapped eight pre-defined animated motions to the reported, user-perceived valence, arousal and dominance and discussed possible directions for future work.
Yifan Zhuang, Keitaro Tsuchiya, Takuro Nakao, Jiawen Han, Megumi Isogai, Shinya Shimizu, Kai Kunze
TEI1
2022 Truck Parking Pattern Aggregation and Availability Prediction by Deep Learning
abstract
With the significant increase of e-commerce, freight transportation demand has surged significantly over the past decade. Most of the demand has been served by trucks in the United States. One major problem commonly identified across the country is the worsening truck parking availability because the increase of truck parking facilities has lagged behind the growth of trucking activities. The lack of parking spaces and real-time parking availability information greatly exacerbate the uncertainty of trips, and often results in illegal and potentially dangerous parking or overtime driving. This paper elaborates on pilot research on improving truck parking facilities cooperated with the Washington State Department of Transportation (WSDOT), building and testing the advanced Truck Parking Information and Management System (TPIMS) with the real-time user visualization and prediction function empowered by artificial intelligence. Furthermore, by analyzing the activities of truck drivers, the researchers aggregated the regularity of truck parking patterns by a customized sequential similarity methodology. A Truck Parking Occupancy Prediction (TPOP) neural network for time-variant occupancy prediction by deep learning and attributes embedding is proposed and integrated into the TPIMS. The TPOP achieves 5.82%, 5.07%, 4.84%, and 4.19% mean average percentage error (MAPE) for 16, 8, 4, and 2 minutes ahead of occupancy prediction respectively, significantly outperforms other state-of-the-art methods. Clearly, the proposed solutions can benefit both the truck drivers and government agencies by a more efficient and smart TPIMS.
Hao (Frank) Yang, Yifan Zhuang, Karthik Murthy, Ziyuan Pu, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.3
2021 A Smart, Efficient, and Reliable Parking Surveillance System With Edge Artificial Intelligence on IoT Devices
abstract
Cloud computing has been a main-stream computing service for years. Recently, with the rapid development in urbanization, massive video surveillance data are produced at an unprecedented speed. A traditional solution to deal with the big data would require a large amount of computing and storage resources. With the advances in Internet of things (IoT), artificial intelligence, and communication technologies, edge computing offers a new solution to the problem by processing all or part of the data locally at the edge of a surveillance system. In this study, we investigate the feasibility of using edge computing for smart parking surveillance tasks, specifically, parking occupancy detection using the real-time video feed. The system processing pipeline is carefully designed with the consideration of flexibility, online surveillance, data transmission, detection accuracy, and system reliability. It enables artificial intelligence at the edge by implementing an enhanced single shot multibox detector (SSD). A few more algorithms are developed either locally at the edge of the system or on the centralized data server targeting optimal system efficiency and accuracy. Thorough field tests were conducted in the Angle Lake parking garage for three months. The experimental results are promising that the final detection method achieves over 95% accuracy in real-world scenarios with high efficiency and reliability. The proposed smart parking surveillance system is a critical component of smart cities and can be a solid foundation for future applications in intelligent transportation systems.
Ruimin Ke, Yifan Zhuang, Ziyuan Pu, Yinhai Wang
IEEE Trans. Intell. Transp. Syst.2