Hongfei Xue

dblp:192/2552 · DBLP profile ↗
← Back
39ranked-venue papers
10as first author
34since 2021 · last 2026
0000-0001-9691-9668ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 20 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 18 since 2021Computer networks · 12 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WenetSpeech-Yue: A Large-Scale Cantonese Speech Corpus with Multi-dimensional Annotation
abstract
The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.
Longhao Li, Zhao Guo, Hongjie Chen 0001, Yuhang Dai, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Hui Bu, Jie Li 0001, Jian Kang 0006, Ruibin Yuan, Ziya Zhou, Wei Xue 0002, Lei Xie 0001
AAAI6
2026 Hearing More with Less: Multi-Modal Retrieval-and-Selection Augmented Conversational LLM-Based ASR
abstract
Automatic Speech Recognition (ASR) aims to convert human speech content into corresponding text. In conversational scenarios, effectively utilizing context can enhance its accuracy. Large Language Models' (LLMs) exceptional long-context understanding and reasoning abilities enable LLM-based ASR (LLM-ASR) to leverage historical context for recognizing conversational speech, which has a high degree of contextual relevance. However, existing conversational LLM-ASR methods use a fixed number of preceding utterances or the entire conversation history as context, resulting in significant ASR confusion and computational costs due to massive irrelevant and redundant information. This paper proposes a multi-modal retrieval-and-selection method named MARS that augments conversational LLM-ASR by enabling it to retrieve and select the most relevant acoustic and textual historical context for the current utterance. Specifically, multi-modal retrieval obtains a set of candidate historical contexts, each exhibiting high acoustic or textual similarity to the current utterance. Multi-modal selection calculates the acoustic and textual similarities for each retrieved candidate historical context and, by employing our proposed near-ideal ranking method to consider both similarities, selects the best historical context. Evaluations on the Interspeech 2025 Multilingual Conversational Speech Language Model Challenge dataset show that the LLM-ASR, when trained on only 1.5K hours of data and equipped with the MARS, outperforms the state-of-the-art top-ranking system trained on 179K hours of data.
Bingshen Mu, Hexin Liu, Hongfei Xue, Lei Xie 0001
AAAI3
2026 Lifelong Domain Adaptive 3D Human Pose Estimation
abstract
3D Human Pose Estimation (3D HPE) is vital in various applications, from person re-identification and action recognition to virtual reality. However, the reliance on annotated 3D data collected in controlled environments poses challenges for generalization to diverse in-the-wild scenarios. Existing domain adaptation (DA) paradigms like general DA and source-free DA for 3D HPE overlook the issues of non-stationary target pose datasets. To address these challenges, we propose a novel task named lifelong domain adaptive 3D HPE. To our knowledge, we are the first to introduce the lifelong domain adaptation to the 3D HPE task. In this lifelong DA setting, the pose estimator is pretrained on the source domain and subsequently adapted to distinct target domains. Moreover, during adaptation to the current target domain, the pose estimator cannot access the source and all the previous target domains. The lifelong DA for 3D HPE involves overcoming challenges in adapting to current domain poses and preserving knowledge from previous domains, particularly combating catastrophic forgetting. We present an innovative Generative Adversarial Network (GAN) framework, which incorporates 3D pose generators, a 2D pose discriminator, and a 3D pose estimator. This framework effectively mitigates domain shifts and aligns original and augmented poses. Moreover, we construct a novel 3D pose generator paradigm, integrating pose-aware, temporal-aware, and domain-aware knowledge to enhance the current domain's adaptation and alleviate catastrophic forgetting on previous domains. Our method demonstrates superior performance through extensive experiments on diverse domain adaptive 3D HPE datasets.
Qucheng Peng, Hongfei Xue, Pu Wang 0001, Chen Chen 0001
AAAI2
2026 Walk Before You Dance: High-fidelity and Editable Dance Synthesis via Generative Masked Motion Prior
abstract
Recent advances in dance generation have enabled the automatic synthesis of 3D dance motions. However, existing methods still face significant challenges in simultaneously achieving high realism, precise dance-music synchronization, diverse motion expression, and physical plausibility. To address these limitations, we propose a novel approach that leverages a generative masked text-to-motion model as a distribution prior to learn a probabilistic mapping from diverse guidance signals, including music, genre, and pose, into high-quality dance motion sequences. Our framework also supports semantic motion editing, such as motion inpainting and body part modification. Specifically, we introduce a multi-tower masked motion model that integrates a text-conditioned masked motion backbone with two parallel, modality-specific branches: a music-guidance tower and a pose-guidance tower. The model is trained using synchronized and progressive masked training, which allows effective infusion of the pretrained text-to-motion prior into the dance synthesis process while enabling each guidance branch to optimize independently through its own loss function, mitigating gradient interference. During inference, we introduce classifier-free logits guidance and pose-guided token optimization to strengthen the influence of music, genre, and pose signals. Extensive experiments demonstrate that our method sets a new state of the art in dance generation, significantly advancing the quality and editability over existing approaches.
Foram Niravbhai Shah, Parshwa Shah, Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang 0001, Hongfei Xue, Ahmed Helmy
AAAI6
2025 GenHMR: Generative Human Mesh Recovery
abstract
Human mesh recovery (HMR) is crucial in many computer vision applications; from health to arts and entertainment. HMR from monocular images has predominantly been addressed by deterministic methods that output a single prediction for a given 2D image. However, HMR from a single image is an ill-posed problem due to depth ambiguity and occlusions. Probabilistic methods have attempted to address this by generating and fusing multiple plausible 3D reconstructions, but their performance has often lagged behind deterministic approaches. In this paper, we introduce GenHMR, a novel generative framework that reformulates monocular HMR as an image-conditioned generative task, explicitly modeling and mitigating uncertainties in the 2D-to-3D mapping process. GenHMR comprises two key components: (1) a pose tokenizer to convert 3D human poses into a sequence of discrete tokens in a latent space, and (2) an image-conditional masked transformer to learn the probabilistic distributions of the pose tokens, conditioned on the input image prompt along with randomly masked token sequence. During inference, the model samples from the learned conditional distribution to iteratively decode high-confidence pose tokens, thereby reducing 3D reconstruction uncertainties. To further refine the reconstruction, a 2D pose-guided refinement technique is proposed to directly fine-tune the decoded pose tokens in the latent space, which forces the projected 3D body mesh to align with the 2D pose clues. Experiments on benchmark datasets demonstrate that GenHMR significantly outperforms state-of-the-art methods.
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Pu Wang 0001, Hongfei Xue, Srijan Das, Chen Chen 0001
AAAI4
2025 HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language Models
abstract
We introduce HOIGPT, a token-based generative method that unifies 3D hand-object interactions (HOI) perception and generation, offering the first comprehensive solution for captioning and generating high-quality 3D HOI sequences from a diverse range of conditional signals (e.g. text, objects, partial sequences). At its core, HOIGPT utilizes a large language model to predict the bidrectional transformation between HOI sequences and natural language descriptions. Given text inputs, HOIGPT generates a sequence of hand and object meshes; given (partial) HOI sequences, HOIGPT generates text descriptions and completes the sequences. To facilitate HOI understanding with a large language model, this paper introduces two key innovations: (1) a novel physically grounded HOI tokenizer, the hand-object decomposed VQ-VAE, for discretizing HOI sequences, and (2) a motion-aware language model trained to process and generate both text and HOI tokens. Extensive experiments demonstrate that HOIGPT sets new state-of-the-art performance on both text generation (+2.01% R Precision) and HOI generation (-2.56 FID) across multiple tasks and benchmarks.
Mingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang, Weiyao Wang 0001, Pierre Gleize, Hongfei Xue, Siwei Lyu, Kris Makoto Kitani, Matt Feiszli
CVPR9
2025 mmCooper: A Multi-Agent Multi-Stage Communication-Efficient and Collaboration-Robust Cooperative Perception Framework
abstract
Collaborative perception significantly enhances individual vehicle perception performance through the exchange of sensory information among agents. However, real-world deployment faces challenges due to bandwidth constraints and inevitable calibration errors during information exchange. To address these issues, we propose mmCooper, a novel multi-agent, multi-stage, communication-efficient, and collaboration-robust cooperative perception framework. Our framework leverages a multi-stage collaboration strategy that dynamically and adaptively balances intermediate- and late-stage information to share among agents, enhancing perceptual performance while maintaining communication efficiency. To support robust collaboration despite potential misalignments and calibration errors, our framework prevents misleading low-confidence sensing information from transmission and refines the received detection results from collaborators to improve accuracy. The extensive evaluation results on both real-world and simulated datasets demonstrate the effectiveness of the mmCooper framework and its components.
Bingyi Liu, Jian Teng, Hongfei Xue, Enshu Wang, Chuanhui Zhu, Pu Wang 0001
ICCV3
2025 MaskControl: Spatio-Temporal Control for Masked Motion Synthesis
Ekkasit Pinyoanuntapong, Muhammad Usama Saleem, Korrawe Karunratanakul, Pu Wang 0001, Hongfei Xue, Chen Chen 0001, Chuan Guo 0002, Junli Cao, Jian Ren 0005, Sergey Tulyakov
ICCV5
2025 MaskHand: Generative Masked Modeling for Robust Hand Mesh Reconstruction in the Wild
Muhammad Usama Saleem, Ekkasit Pinyoanuntapong, Mayur Jagdishbhai Patel, Hongfei Xue, Ahmed Helmy, Srijan Das, Pu Wang 0001
ICCV4
2025 Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-Based Action Recognition
abstract
Zero-shot skeleton-based action recognition aims to develop models capable of identifying actions beyond the categories encountered during training. Previous approaches have primarily focused on aligning visual and semantic representations but often overlooked the importance of fine-grained action patterns in the semantic space (e.g., the hand movements in drinking water and brushing teeth). To address these limitations, we propose a Frequency-Semantic Enhanced Variational Autoencoder (FS-VAE) to explore the skeleton semantic representation learning with frequency decomposition. FS-VAE consists of three key components: 1) a frequency-based enhancement module with high- and low-frequency adjustments to enrich the skeletal semantics learning and improve the robustness of zero-shot action recognition; 2) a semantic-based action description with multilevel alignment to capture both local details and global correspondence, effectively bridging the semantic gap and compensating for the inherent loss of information in skeleton sequences; 3) a calibrated cross-alignment loss that enables valid skeleton-text pairs to counterbalance ambiguous ones, mitigating discrepancies and ambiguities in skeleton and text features, thereby ensuring robust alignment. Evaluations on the benchmarks demonstrate the effectiveness of our approach, validating that frequency-enhanced semantic features enable robust differentiation of visually and semantically similar action clusters, improving zero-shot action recognition.
Zhishuai Guo, Chen Chen 0001, Hongfei Xue, Aidong Lu
ICCV4
2025 Physical Backdoor Attacks against mmWave-based Human Activity Recognition
abstract
Human Activity Recognition (HAR) using wireless signals like mmWave technology has promising applications in numerous scenarios, including monitoring and surveillance, healthcare, and smart home. Wireless HAR is non-intrusive and can operate in situations where traditional sensors or cameras may fail. However, these systems also introduce new attack surfaces alongside their benefits. Existing security research on wireless HAR primarily focuses on the vulnerabilities of the AI models used by these systems, without addressing the challenges of physically implementing these attacks in real-world scenarios. In this paper, we present the first physical backdoor attack for mmWave-based HAR systems, manipulating physical signals to deceive the systems into producing targeted outputs. Utilizing passive metal reflectors and optimized attacking strategies, our attack is efficient, stealthy, and easy to implement. Tailored experiments on a mmWave HAR prototype demonstrate the high effectiveness of the proposed attack.
Ziqian Bi, Amit Singha, Hongfei Xue, Tao Li 0042, Yimin Chen 0004
ICDCS3
2025 Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
Longhao Li, Yangze Li, Hongfei Xue, Jie Liu 0097, Shuai Fang, Lei Xie 0001
INTERSPEECH3
2025 Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty
Hongfei Xue, Yufeng Tang, Xuelong Geng, Lei Xie 0001
INTERSPEECH1
2025 Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
abstract
Large language models have been extended to the speech domain, leading to the development of speech large language models (SLLMs). While existing SLLMs demonstrate strong performance in speech instruction-following for core languages (e.g., English), they often struggle with non-core languages due to the scarcity of paired speech-text data and limited multilingual semantic reasoning capabilities. To address this, we propose the semi-implicit Cross-lingual Speech Chain-of-Thought (XS-CoT) framework, which integrates speech-to-text translation into the reasoning process of SLLMs. The XS-CoT generates four types of tokens: instruction and response tokens in both core and non-core languages, enabling cross-lingual transfer of reasoning capabilities. To mitigate inference latency in generating target non-core response tokens, we incorporate a semi-implicit CoT scheme into XS-CoT, which progressively compresses the first three types of intermediate reasoning tokens while retaining global reasoning logic during training. By leveraging the robust reasoning capabilities of the core language, XS-CoT improves responses for non-core languages by up to 45% in GPT-4 score when compared to direct supervised fine-tuning on two representative SLLMs, Qwen2-Audio and SALMONN. Moreover, the semi-implicit XS-CoT reduces token delay by more than 50% with a slight drop in GPT-4 scores. Importantly, XS-CoT requires only a small amount of high-quality training data for non-core languages by leveraging the reasoning capabilities of core languages. To support training, we also develop a data pipeline and open-source speech instruction-following datasets in Japanese, German, and French.
Hongfei Xue, Yufeng Tang, Hexin Liu, Xuelong Geng, Lei Xie 0001
ACM Multimedia1
2025 Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on
abstract
In this paper, we propose Argus, a wearable add-on system based on stripped-down (i.e., compact, lightweight, low-power, limited-capability) mmWave radars. It is the first to achieve egocentric human mesh reconstruction in a multi-view manner. Compared with conventional frontal-view mmWave sensing solutions, it addresses several pain points, such as restricted sensing range, occlusion, and the multipath effect caused by surroundings. To overcome the limited capabilities of the stripped-down mmWave radars (with only one transmit antenna and three receive antennas), we tackle three main challenges and propose a holistic solution, including tailored hardware design, sophisticated signal processing, and a deep neural network optimized for high-dimensional complex point clouds. Extensive evaluation shows that Argus achieves performance comparable to traditional solutions based on high-capability mmWave radars, with an average vertex error of 6.5 cm, solely using stripped-down radars deployed in a multi-view configuration. It presents robustness and practicality across conditions, such as with unseen users and different host devices.
Di Duan, Shengzhe Lyu, Mu Yuan, Hongfei Xue, Tianxing Li 0001, Weitao Xu, Kaishun Wu, Guoliang Xing
SenSys4
2025 RAM-Hand: Robust Acoustic Multi-Hand Pose Reconstruction Using a Microphone Array
abstract
Using 3D hand poses as the input of user interfaces can enable many novel human-computer interaction applications. However, conventional solutions for precisely reconstructing the hand poses are either vision-based, which are compute-intensive and may cause privacy issues, or wearable devices-based, which are intrusive to users. In this paper, we propose RAM-Hand, a Robust Acoustic 3D Multi-Hand pose reconstruction system built on a microphone array. Our RAM-Hand system can support multiple hands and is designed to be highly adaptable to new scenarios even when training data is limited. Specifically, it should robustly accommodate variations in environment, subject, and hand positions. To achieve this, on one hand, we propose a customized signal processing pipeline to segment multiple hands' reflections and extract the features corresponding to each hand, then feed those features into a transformer-based neural network for precise pose reconstruction. On the other hand, to tackle the challenge that the training data is limited, we propose a series of data augmentation methods to generate virtual training data, and utilize contrastive learning to ensure our model behaves well on new subjects. We conduct extensive experiments on a real-world microphone array testbed to evaluate the performance of the proposed system. The results show that our RAM-Hand system can localize each hand joint with an average error of 10.71 mm, handle multiple hands, and generalize well to the above mentioned new scenarios.
Henglin Pu, Qiming Cao, Tianci Liu 0003, Zhengxin Jiang, Hongfei Xue, Lu Su 0001
SenSys8
2025 BioPose: Biomechanically-Accurate 3D Pose Estimation from Monocular Videos
abstract
Recent advancements in 3D human pose estimation from single-camera images and videos have relied on parametric models, like SMPL. However, these models over-simplify anatomical structures, limiting their accuracy in capturing true joint locations and movements, which reduces their applicability in biomechanics, healthcare, and robotics. Biomechanically accurate pose estimation, on the other hand, typically requires costly marker-based motion capture systems and optimization techniques in specialized labs. To bridge this gap, we propose BioPose, a novel learning-based framework for predicting biomechanically accurate 3D human pose directly from monocular videos. BioPose includes three key components: a Multi-Query Human Mesh Recovery model (MQ-HMR), a Neural Inverse Kinematics (NeurIK) model, and a 2D-informed pose refinement technique. MQ-HMR leverages a multi-query deformable transformer to extract multi-scale fine-grained image features, enabling precise human mesh recovery. NeurIK treats the mesh vertices as virtual markers, applying a spatial-temporal network to regress biomechanically accurate 3D poses under anatomical constraints. To further improve 3D pose estimations, a 2D-informed refinement step optimizes the query tokens during inference by aligning the 3D structure with 2D pose observations. Experiments on benchmark datasets demonstrate that BioPose significantly outperforms state-of-the-art methods.
Farnoosh Koleini, Muhammad Usama Saleem, Pu Wang 0001, Hongfei Xue, Ahmed Helmy, Abbey Fenwick
WACV4
2025 mmHand: Toward Pixel-Level-Accuracy Hand Localization Using a Single Commodity mmWave Device
abstract
The hand localization problem has been a longstanding focus due to its many applications. The task involves modeling the hand as a singular point and determining its position within a defined coordinate system. However, due to data modality limitations, existing hand localization technologies face several challenges. For example, vision-based localization raises privacy concerns, while wearable-based methods compromise user comfort. In this article, we introduce mmHand, a new device-free, privacy-preserving dynamic hand localization system with pixel-level accuracy, using a single commodity mmWave device. We first propose a mmImage generation tool to fully extract spatial information from raw mmWave data and introduce a novel 2-D image-format representation of mmWave data. Next, we design a framework that provides a new quality evaluation method and pixel space labeling for the mmWave data. Finally, we present a cross-modality spatial feature-enhanced model with high spatial feature extraction capabilities, which can accurately localize hand positions at the pixel level in the mmWave radar U-V pixel coordinate system. We evaluate the system with experiments on 12 subjects in three scenarios, and the results across four metrics demonstrate the effectiveness of our hand localization system.
Zhengxiong Li, Chenhan Xu, Luchuan Song, Huining Li, Hongfei Xue, Yingxiao Wu, Wenyao Xu
IEEE Internet Things J.6
2024 Towards Robust mmWave-based Human Activity Recognition using Large Simulated Dataset for Model Pretraining
abstract
Human activity recognition (HAR) is crucial for real-world applications such as healthcare, surveillance, and smart homes. Among sensing technologies, millimeter wave (mmWave) sensors stand out due to their contactless nature, high sensitivity, and ability to operate in low-light environments while preserving privacy. However, the scarcity of mmWave sensing data limits the generalizability of mmWave-based HAR systems. To address this, we propose mmAP, a data augmentation and pretraining framework that synthesizes a large mmWave dataset using human mesh data, followed by pretraining a robust and general mmWave heatmap encoder using a multi-modal masked autoencoder framework using the synthesized data. We enhance the model’s robustness with heatmap-specific data perturbations and perform task-specific fine-tuning on a small real-world dataset. The experiment results over the baseline demonstrate the effectiveness of the proposed mmAP framework.
Vinay Joshi, Shengkai Xu, Qiming Cao, Yi Zhu 0012, Pu Wang 0001, Hongfei Xue
IEEE Big Data6
2024 Breakthrough from Nuance and Inconsistency: Enhancing Multimodal Sarcasm Detection with Context-Aware Self-Attention Fusion and Word Weight Calculation
abstract
Multimodal sarcasm detection has received considerable attention due to its unique role in social networks. Existing methods often rely on feature concatenation to fuse different modalities or model the inconsistencies among modalities. However, sarcasm is often embodied in local and momentary nuances in a subtle way, which causes difficulty for sarcasm detection. To effectively incorporate these nuances, this paper presents Context-Aware Self-Attention Fusion (CAAF) to integrate local and momentary multimodal information into specific words. Furthermore, due to the instantaneous nature of sarcasm, the connotative meanings of words post-multimodal integration generally deviate from their denotative meanings. Therefore, Word Weight Calculation (WWC) is presented to compute the weight of specific words based on CAAF’s fusion nuances, illustrating the inconsistency between connotation and denotation. We evaluate our method on the MUStARD dataset, achieving an accuracy of 76.9 and an F1 score of 76.1, which surpasses the current state-of-the-art IWAN model by 1.7 and 1.6 respectively.
Hongfei Xue, Linyan Xu, Jiali Lin, Dazhi Jiang
LREC/COLING1
2024 SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
abstract
Multilingual automatic speech recognition (ASR) systems have garnered attention for their potential to extend language coverage globally. While self-supervised learning (SSL) models, like MMS, have demonstrated their effectiveness in multilingual ASR, it is worth noting that various layers’ representations potentially contain distinct information that has not been fully leveraged. In this study, we propose a novel method that leverages self-supervised hierarchical representations (SSHR) to fine-tune the MMS model. We first analyze the different layers of MMS and show that the middle layers capture language-related information, and the high layers encode content-related information, which gradually decreases in the final layers. Then, we extract a language-related frame from correlated middle layers and guide specific language extraction through self-attention mechanisms. Additionally, we steer the model toward acquiring more content-related information in the final layers using our proposed Cross-CTC. We evaluate SSHR on two multilingual datasets, Common Voice and ML-SUPERB, and the experimental results demonstrate that our method achieves state-of-the-art performance to the best of our knowledge.
Hongfei Xue, Qijie Shao, Kaixun Huang, Peikun Chen, Jie Liu 0097, Lei Xie 0001
ICME1
2024 AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
Rong Gong, Hongfei Xue, Lezhi Wang, Qisheng Li, Lei Xie 0001, Hui Bu, Shaomei Wu, Jiaming Zhou 0001, Jun Du 0002, Jia Bin, Ming Li 0026
INTERSPEECH2
2024 Malicious Attacks against Multi-Sensor Fusion in Autonomous Driving
abstract
Multi-sensor fusion has been widely used by autonomous vehicles (AVs) to integrate the perception results from different sensing modalities including LiDAR, camera and radar. Despite the rapid development of multi-sensor fusion systems in autonomous driving, their vulnerability to malicious attacks have not been well studied. Although some prior works have studied the attacks against the perception systems of AVs, they only consider a single sensing modality or a camera-LiDAR fusion system, which can not attack the sensor fusion system based on LiDAR, camera, and radar. To fill this research gap, in this paper, we present the first study on the vulnerability of multi-sensor fusion systems that employ LiDAR, camera, and radar. Specifically, we propose a novel attack method that can simultaneously attack all three types of sensing modalities using a single type of adversarial object. The adversarial object can be easily fabricated at low cost, and the proposed attack can be easily performed with high stealthiness and flexibility in practice. Extensive experiments based on a real-world AV testbed show that the proposed attack can continuously hide a target vehicle from the perception system of a victim AV using only two small adversarial objects.
Yi Zhu 0012, Chenglin Miao, Hongfei Xue, Yunnan Yu, Lu Su 0001, Chunming Qiao
MobiCom3
2024 mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment
abstract
Millimeter-wave (mmWave) based human activity recognition (HAR) systems have demonstrated promising performance in various applications, leveraging the power of deep neural networks. However, these systems are suffering from the scarcity of available mmWave data for model training. To address this challenge, we explore the possibility of transferring knowledge from large AI models built on massive text and visual data to enhance the generalizability of mmWave-based HAR models. Towards this end, we introduce mmCLIP, a novel system that aligns mmWave signal space and text space to facilitate zero-shot recognition for unseen activities. To enable this alignment, we employ cross-modality signal synthesis to augment mmWave signal data using large human mesh datasets and design an activity attribute decomposition and recomposition approach to characterize the semantic interconnections among activities. We conducted extensive experiments to demonstrate the effectiveness of our proposed framework.
Qiming Cao, Hongfei Xue, Tianci Liu 0003, Haoyu Wang 0004, Xincheng Zhang, Lu Su 0001
SenSys2
2024 Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
abstract
The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies.
Hongfei Xue, Rong Gong, Mingchen Shao, Lezhi Wang, Lei Xie 0001, Hui Bu, Jiaming Zhou 0001, Jun Du 0002, Ming Li 0026
SLT1
2024 Towards Smartphone-based 3D Hand Pose Reconstruction Using Acoustic Signals
abstract
Accurately reconstructing 3D hand poses is a pivotal element for numerous Human-Computer Interaction applications. In this work, we propose SonicHand, the first smartphone-based 3D hand pose reconstruction system using purely inaudible acoustic signals. SonicHand incorporates signal processing techniques and a deep learning framework to address a series of challenges. First, it encodes the topological information of the hand skeleton as prior knowledge and utilizes a deep learning model to realistically and smoothly reconstruct the hand poses. Second, the system employs adversarial training to enhance the generalization ability of our system to be deployed in a new environment or for a new user. Third, we adopt a hand tracking method based on channel impulse response estimation. It enables our system to handle the scenario where the hand performs gestures while moving arbitrarily as a whole. We conduct extensive experiments on a smartphone testbed to demonstrate the effectiveness and robustness of our system from various dimensions. The experiments involve 10 subjects performing up to 12 different hand gestures in three distinctive environments. When the phone is held in one of the user’s hands, the proposed system can track joints with an average error of 18.64 mm.
Chenglin Miao, Qiming Cao, Haoyu Wang 0004, Ke Sun 0012, Hongfei Xue, Lu Su 0001
ACM Trans. Sens. Networks8
2023 BA-MoE: Boundary-Aware Mixture-of-Experts Adapter for Code-Switching Speech Recognition
abstract
Mixture-of-experts based models, which use language experts to extract language-specific representations effectively, have been well applied in code-switching automatic speech recognition. However, there is still substantial space to improve as similar pronunciation across languages may result in ineffective multi-language modeling and inaccurate language boundary estimation. To eliminate these drawbacks, we propose a cross-layer language adapter and a boundary-aware training method, namely Boundary-Aware Mixture-of-Experts (BA-MoE). Specifically, we introduce language-specific adapters to separate language-specific representations and a unified gating layer to fuse representations within each encoder layer. Second, we compute language adaptation loss of the mean output of each language-specific adapter to improve the adapter module’s language-specific representation learning. Besides, we utilize a boundary-aware predictor to learn boundary representations for dealing with language boundary confusion. Our approach achieves significant performance improvement, reducing the mixture error rate by 16.55% compared to the baseline on the ASRU 2019 Mandarin-English code-switching challenge dataset.
Peikun Chen, Fan Yu 0002, Yuhao Liang, Hongfei Xue, Xucheng Wan, Naijun Zheng, Huan Zhou 0004, Lei Xie 0001
ASRU4
2023 TileMask: A Passive-Reflection-based Attack against mmWave Radar Object Detection in Autonomous Driving
abstract
In autonomous driving, millimeter wave (mmWave) radar has been widely adopted for object detection because of its robustness and reliability under various weather and lighting conditions. For radar object detection, deep neural networks (DNNs) are becoming increasingly important because they are more robust and accurate, and can provide rich semantic information about the detected objects, which is critical for autonomous vehicles (AVs) to make decisions. However, recent studies have shown that DNNs are vulnerable to adversarial attacks. Despite the rapid development of DNN-based radar object detection models, there have been no studies on their vulnerability to adversarial attacks. Although some spoofing attack methods are proposed to attack the radar sensor by actively transmitting specific signals using some special devices, these attacks require sub-nanosecond-level synchronization between the devices and the radar and are very costly, which limits their practicability in real world. In addition, these attack methods can not effectively attack DNN-based radar object detection. To address the above problems, in this paper, we investigate the possibility of using a few adversarial objects to attack the DNN-based radar object detection models through passive reflection. These objects can be easily fabricated using 3D printing and metal foils at low cost. By placing these adversarial objects at some specific locations on a target vehicle, we can easily fool the victim AV's radar object detection model. The experimental results demonstrate that the attacker can achieve the attack goal by using only two adversarial objects and conceal them as car signs, which have good stealthiness and flexibility. To the best of our knowledge, this is the first study on the passive-reflection-based attacks against the DNN-based radar object detection models using low-cost, readily-available and easily concealable geometric shaped objects.
Yi Zhu 0012, Chenglin Miao, Hongfei Xue, Zhengxiong Li, Yunnan Yu, Wenyao Xu, Lu Su 0001, Chunming Qiao
CCS3
2023 TranUSR: Phoneme-to-word Transcoder Based Unified Speech Representation Learning for Cross-lingual Speech Recognition
Hongfei Xue, Qijie Shao, Peikun Chen, Lei Xie 0001, Jie Liu 0097
INTERSPEECH1
2023 Macular: A Multi-Task Adversarial Framework for Cross-Lingual Natural Language Understanding
abstract
Cross-lingual natural language understanding~(NLU) aims to train NLU models on a source language and apply the models to NLU tasks in target languages, and is a fundamental task for many cross-language applications. Most of the existing cross-lingual NLU models assume the existence of parallel corpora so that words and sentences in source and target languages could be aligned. However, the construction of such parallel corpora is expensive and sometimes infeasible. Motivated by this challenge, recent works propose data augmentation or adversarial training methods to reduce the reliance on external parallel corpora. In this paper, we propose an orthogonal and novel perspective to tackle this challenging cross-lingual NLU task (i.e., when parallel corpora are unavailable). We propose to conduct multi-task learning across different tasks for mutual performance improvement on both source and target languages. The proposed multi-task learning framework is complementary to existing studies and could be integrated with existing methods to further improve their performance on challenging cross-lingual NLU tasks.
Haoyu Wang 0004, Yaqing Wang 0001, Feijie Wu, Hongfei Xue, Jing Gao 0004
KDD4
2023 Towards Generalized mmWave-based Human Pose Estimation through Signal Augmentation
abstract
The unprecedented advance of wireless human sensing is enabled by the proliferation of the deep learning techniques, which, however, rely heavily on the completeness and representativeness of the data patterns contained in the training set. Thus, deep learning based wireless human perception models usually fail when the human subject is conducting activities that are unseen during the model training. To address this problem, we propose a novel wireless signal augmentation framework, named mmGPE, for Generalized mmWave-based Pose Estimation. In mmGPE, we adopt a physical simulator to generate mmWave FMCW signals. However, due to the imperfect simulation of the physical world, there is a big gap between the signals generated by the physical simulator and the real-world signals collected by the mmWave radar. To tackle this challenge, we propose to integrate the physical signal simulation with deep learning techniques. Specifically, we develop a deep learning-based signal refiner in mmGPE that is capable of bridging the gap and generating realistic signal data. Through extensive evaluations on a COTS mmWave testbed, our mmGPE system demonstrates high accuracy in generating human meshes for unseen activities.
Hongfei Xue, Qiming Cao, Chenglin Miao, Yan Ju, Haochen Hu, Aidong Zhang 0001, Lu Su 0001
MobiCom1
2022 Fusing Global and Local Features for Generalized AI-Synthesized Image Detection
abstract
With the development of the Generative Adversarial Networks (GANs) and DeepFakes, AI-synthesized images are now of such high quality that humans can hardly distinguish them from real images. It is imperative for media forensics to develop detectors to expose them accurately. Existing detection methods have shown high performance in generated images detection, but they tend to generalize poorly in the real-world scenarios, where the synthetic images are usually generated with unseen models using unknown source data. In this work, we emphasize the importance of combining information from the whole image and informative patches in improving the generalization ability of AI-synthesized image detection. Specifically, we design a two-branch model to combine global spatial information from the whole image and local informative features from multiple patches selected by a novel patch selection module. Multi-head attention mechanism is further utilized to fuse the global and local features. We collect a highly diverse dataset synthesized by 19 models with various objects and resolutions to evaluate our model. Experimental results demonstrate the high accuracy and good generalization ability of our method in detecting generated images. Our code is available at https://github.com/littlejuyan/FusingGlobalandLocal.
Yan Ju, Shan Jia, Lipeng Ke, Hongfei Xue, Koki Nagano, Siwei Lyu
ICIP4
2022 M4esh: mmWave-Based 3D Human Mesh Construction for Multiple Subjects
abstract
The recent proliferation of various wireless sensing systems and applications demonstrates the advantages of radio frequency (RF) signals over traditional camera-based solutions that are faced with various challenges, such as occlusions and poor lighting conditions. Towards the ultimate goal of imaging human body using RF signals, researchers have been exploring the possibility of constructing the human mesh, a structure capturing not only the pose but also the shape of the human body, from RF signals. In this paper, we introduce M4esh, a novel system that utilizes commercial millimeter wave (mmWave) radar for multi-subject 3D human mesh construction. Our M4esh system can detect and track the subjects on a 2D energy map by predicting the subject bounding boxes on the map, and tackle the subjects' mutual occlusion through utilizing the location, velocity and size information of the subjects' bounding boxes from the previous frames as a clue to estimate the bounding box in the current frame. Through extensive experiments on a real-world COTS millimeter-wave testbed, we show that our proposed M4esh system can accurately localize the subjects and generate their human meshes, which demonstrate the superior effectiveness of the proposed M4esh system.
Hongfei Xue, Qiming Cao, Yan Ju, Haochen Hu, Haoyu Wang 0004, Aidong Zhang 0001, Lu Su 0001
SenSys1
2021 mmMesh: towards 3D real-time dynamic human mesh construction using millimeter-wave
abstract
In this paper, we present mmMesh, the first real-time 3D human mesh estimation system using commercial portable millimeter-wave devices. mmMesh is built upon a novel deep learning framework that can dynamically locate the moving subject and capture his/her body shape and pose by analyzing the 3D point cloud generated from the mmWave signals that bounce off the human body. The proposed deep learning framework addresses a series of challenges. First, it encodes a 3D human body model, which enables mmMesh to estimate complex and realistic-looking 3D human meshes from sparse point clouds. Second, it can accurately align the 3D points with their corresponding body segments despite the influence of ambient points as well as the error-prone nature and the multi-path effect of the RF signals. Third, the proposed model can infer missing body parts from the information of the previous frames. Our evaluation results on a commercial mmWave sensing testbed show that our mmMesh system can accurately localize the vertices on the human mesh with an average error of 2.47 cm. The superior experimental results demonstrate the effectiveness of our proposed human mesh construction system.
Hongfei Xue, Yan Ju, Chenglin Miao, Yijiang Wang, Aidong Zhang 0001, Lu Su 0001
MobiSys1
2020 Towards 3D human pose construction using wifi
abstract
This paper presents WiPose, the first 3D human pose construction framework using commercial WiFi devices. From the pervasive WiFi signals, WiPose can reconstruct 3D skeletons composed of the joints on both limbs and torso of the human body. By overcoming the technical challenges faced by traditional camera-based human perception solutions, such as lighting and occlusion, the proposed WiFi human sensing technique demonstrates the potential to enable a new generation of applications such as health care, assisted living, gaming, and virtual reality. WiPose is based on a novel deep learning model that addresses a series of technical challenges. First, WiPose can encode the prior knowledge of human skeleton into the posture construction process to ensure the estimated joints satisfy the skeletal structure of the human body. Second, to achieve cross environment generalization, WiPose takes as input a 3D velocity profile which can capture the movements of the whole 3D space, and thus separate posture-specific features from the static objects in the ambient environment. Finally, WiPose employs a recurrent neural network (RNN) and a smooth loss to enforce smooth movements of the generated skeletons. Our evaluation results on a real-world WiFi sensing testbed with distributed antennas show that WiPose can localize each joint on the human skeleton with an average error of 2.83cm, achieving a 35% improvement in accuracy over the state-of-the-art posture construction model designed for dedicated radar sensors.
Hongfei Xue, Chenglin Miao, Sen Lin 0009, Chong Tian, Srinivasan Murali, Haochen Hu, Lu Su 0001
MobiCom2
2019 Deep Metric Learning: The Generalization Analysis and an Adaptive Algorithm
abstract
As an effective way to learn a distance metric between pairs of samples, deep metric learning (DML) has drawn significant attention in recent years. The key idea of DML is to learn a set of hierarchical nonlinear mappings using deep neural networks, and then project the data samples into a new feature space for comparing or matching. Although DML has achieved practical success in many applications, there is no existing work that theoretically analyzes the generalization error bound for DML, which can measure how good a learned DML model is able to perform on unseen data. In this paper, we try to fill up this research gap and derive the generalization error bound for DML. Additionally, based on the derived generalization bound, we propose a novel DML method (called ADroDML), which can adaptively learn the retention rates for the DML models with dropout in a theoretically justified way. Compared with existing DML works that require predefined retention rates, ADroDML can learn the retention rates in an optimal way and achieve better performance. We also conduct experiments on real-world datasets to verify the findings derived from the generalization error bound and demonstrate the effectiveness of the proposed adaptive DML method.
Mengdi Huai, Hongfei Xue, Chenglin Miao, Liuyi Yao, Lu Su 0001, Changyou Chen, Aidong Zhang 0001
IJCAI2
2019 On the Estimation of Treatment Effect with Text Covariates
abstract
Estimating the treatment effect benefits decision making in various domains as it can provide the potential outcomes of different choices. Existing work mainly focuses on covariates with numerical values, while how to handle covariates with textual information for treatment effect estimation is still an open question. One major challenge is how to filter out the nearly instrumental variables which are the variables more predictive to the treatment than the outcome. Conditioning on those variables to estimate the treatment effect would amplify the estimation bias. To address this challenge, we propose a conditional treatment-adversarial learning based matching method (CTAM). CTAM incorporates the treatment-adversarial learning to filter out the information related to nearly instrumental variables when learning the representations, and then it performs matching among the learned representations to estimate the treatment effects. The conditional treatment-adversarial learning helps reduce the bias of treatment effect estimation, which is demonstrated by our experimental results on both semi-synthetic and real-world datasets.
Liuyi Yao, Sheng Li 0001, Yaliang Li, Hongfei Xue, Jing Gao 0004, Aidong Zhang 0001
IJCAI4
2019 DeepFusion: A Deep Learning Framework for the Fusion of Heterogeneous Sensory Data
abstract
In recent years, significant research efforts have been spent towards building intelligent and user-friendly IoT systems to enable a new generation of applications capable of performing complex sensing and recognition tasks. In many of such applications, there are usually multiple different sensors monitoring the same object. Each of these sensors can be regarded as an information source and provides us a unique "view" of the observed object. Intuitively, if we can combine the complementary information carried by multiple sensors, we will be able to improve the sensing performance. Towards this end, we propose DeepFusion, a unified multi-sensor deep learning framework, to learn informative representations of heterogeneous sensory data. DeepFusion can combine different sensors' information weighted by the quality of their data and incorporate cross-sensor correlations, and thus can benefit a wide spectrum of IoT applications. To evaluate the proposed DeepFusion model, we set up two real-world human activity recognition testbeds using commercialized wearable and wireless sensing devices. Experiment results show that DeepFusion can outperform the state-of-the-art human activity recognition methods.
Hongfei Xue, Chenglin Miao, Ye Yuan 0006, Fenglong Ma, Xin Ma 0006, Yijiang Wang, Shuochao Yao, Wenyao Xu, Aidong Zhang 0001, Lu Su 0001
MobiHoc1
2018 Towards Environment Independent Device Free Human Activity Recognition
abstract
Driven by a wide range of real-world applications, significant efforts have recently been made to explore device-free human activity recognition techniques that utilize the information collected by various wireless infrastructures to infer human activities without the need for the monitored subject to carry a dedicated device. Existing device free human activity recognition approaches and systems, though yielding reasonably good performance in certain cases, are faced with a major challenge. The wireless signals arriving at the receiving devices usually carry substantial information that is specific to the environment where the activities are recorded and the human subject who conducts the activities. Due to this reason, an activity recognition model that is trained on a specific subject in a specific environment typically does not work well when being applied to predict another subject's activities that are recorded in a different environment. To address this challenge, in this paper, we propose EI, a deep-learning based device free activity recognition framework that can remove the environment and subject specific information contained in the activity data and extract environment/subject-independent features shared by the data collected on different subjects under different environments. We conduct extensive experiments on four different device free activity recognition testbeds: WiFi, ultrasound, 60 GHz mmWave, and visible light. The experimental results demonstrate the superior effectiveness and generalizability of the proposed EI framework.
Chenglin Miao, Fenglong Ma, Shuochao Yao, Yaqing Wang 0001, Ye Yuan 0006, Hongfei Xue, Chen Song 0001, Xin Ma 0006, Dimitrios Koutsonikolas, Wenyao Xu, Lu Su 0001
MobiCom7