Xiaojing Yu

dblp:08/1625 · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Multi-level dynamic fusion and temporal role-aware network for diagnosis prediction
Zhuang Fu, Youshen Chi, Xiaojing Yu, Yuliang Shi, Lin Cheng 0007, Hui Li 0048
J. Biomed. Informatics3
2026 MGFNet: Multi-granularity medical pattern fusion network for patient risk prediction
Lin Cheng 0007, Yuliang Shi, Xiaojing Yu, Xinjun Wang 0003, Zhongmin Yan
Pattern Recognit.3
2025 TypeFly: Low-Latency Drone Planning With Large Language Models
abstract
Recent advancements in robot planning using large language models (LLMs) have demonstrated significant potential, primarily due to LLMs' capabilities to understand natural language commands and generate executable plans in various languages. However, in time-sensitive and interactive applications involving mobile robots, particularly drones, the sequential token generation process inherent to LLMs introduces substantial latency, i.e., response time, during the control plan generation. In this paper, we present a system called ChatFly that tackles this latency problem using a combination of a novel programming language called MiniSpec and its runtime to reduce both the response time and generation time for the robot plan. That is, instead of asking an LLM to write a program (robotic plan) in the popular but verbose Python, ChatFly gets it to do it in MiniSpec specially designed for token efficiency and stream interpreting. Using a set of challenging drone tasks, we show that design choices made by ChatFly can reduce the average response time to 74% compared to existing works and provide a more consistent user experience, enabling responsive and intelligent LLM-based drone control.
Xiaojing Yu, Neiwen Ling, Lin Zhong 0001
IEEE Trans. Mob. Comput.2
2025 IUAC: Inaudible Universal Adversarial Attacks Against Smart Speakers
abstract
Intelligent voice systems are widely utilized to control smart home applications, which raises significant privacy and security concerns. Recent studies have revealed their vulnerability to adversarial attacks, replay attacks, and so on. However, these attacks rely on the victim’s voice data. In our work, we investigate a stealthy and command-independent attack that does not necessitate collecting victims’ voices. Our proposed attack, IUAC, misleads the voice system to go against the victim’s will, regardless of the commands delivered. Our core concept is to train highly robust attack commands through the construction of diverse data, rendering the user’s commands negligible. To achieve stealthy attacks, we leverage a high-frequency carrier to construct an inaudible universal adversarial command. Extensive experiments conducted with real-world datasets demonstrate that our attack system attains an average attack success rate of 96% while resisting environmental interference. Moreover, our attack success rate against real-world voice systems is 4.52× higher than the state-of-the-art. Finally, we propose an effective defense mechanism and provide experimental tests to validate its efficacy.
Haifeng Sun 0005, Haohua Du, Xiaojing Yu, Jiahui Hou, Lan Zhang 0002, Xiang-Yang Li 0001
ACM Trans. Sens. Networks3
2024 ThumbUp: Secure Smartwatch Controller for Smart Homes Using Simple Hand Gestures
abstract
The development of creative applications and intelligent gadgets requires a secure and straightforward interface with human users. We propose, design, and implement ThumbUp, a smartwatch-based two-factor real-time identification and authentication system in which smartwatch users can identify and authenticate themselves using some simple hand and finger movements, such as thumb-up. ThumbUp leverages the signal from the Inertial Measurement Unit (IMU) in in Commercial-Off-The-Shelf (COTS) smart devices to discover the unique pattern generated by each user's simple gestures using a carefully constructed deep learning model. Smart homes provide a comfortable, safe, and efficient living environment, epecially help the sick and aged. We propose strategies for convenient and reliable control in smart homes with gesture command recognition. We build an Auto-Encoder-based filter that reconstructs the raw data to improve the representation of gesture features. Moreover, we adopt the random forest method to analyse the contextual command correlation. And we employ the authentication system based on smartwatch for personalized command feedback and ensure that illegals cannot use the device. We implement our system and undertake rigorous studies to determine its usefulness and efficiency over a three-month period with 65 users. It achieves a 97% accuracy for user classification and an EER of 0.014 for authentication task with a single simple gesture. And our method achieves 91% accuracy for command recognition and 96% command accuracy with contextual informations. Additionally, we conduct a study of user acceptability of our system and explain how gesture proficiency influences authentication accuracy.
Xiaojing Yu, Zhijun Zhou, Lan Zhang 0002, Xiang-Yang Li 0001
IEEE Trans. Mob. Comput.1
2023 Differentially Private Distributed Online Convex Optimization Towards Low Regret and Communication Cost
abstract
Distributed online convex optimization (DOCO) has emerged as a promising approach in scenarios where multiple learners collaboratively serve sequential (possibly untrusted) clients using AI models. However, ensuring clients' privacy and minimizing regret while keeping the communication cost reasonable poses a significant challenge. To address this issue, we propose private DOCO algorithms, termed PDOM, for both oblivious and stochastic settings. Our approach involves a mini-batch strategy that optimally balances the effects of slower model updates and differential privacy (DP) perturbation. Our theoretical analysis shows that compared to state-of-the-art algorithms, PDOM reduces the regret bounds and communication cost by an O(d/ε) factor for the oblivious setting. For the stochastic setting, the impact of DP perturbation becomes negligible if the learning time T = Ω(d4/(nε4)) provided that the loss functions are Lipschitz, convex, and smooth, where d is the model dimension, n is the number of learners, and ε is the privacy budget. Our evaluations validate these results, demonstrating that PDOM can reduce classification error rates of the state-of-the-art methods by up to 20% in distributed online logistic regression tasks while achieving communication savings of above 90%.
Lan Zhang 0002, Xiaojing Yu, Xiang-Yang Li 0001
MobiHoc3
2023 E-Talk: Accelerating Active Speaker Detection with Audio-Visual Fusion and Edge-Cloud Computing
abstract
Active Speaker Detection (ASD) aims to enhance communication and interaction in various scenarios, including meetings, group discussions, and security surveillance systems. The primary objective of ASD is to identify and label the position of the main active speaker. In large-scale surveillance systems, real-time ASD can pose network congestion issues due to the extensive video data uploaded from numerous cameras. To address this challenge, we propose a collaborative edge-cloud solution called E-TALK for ASD. E-TALK leverages the simplicity of voiceprint comparison and processing, as opposed to analyzing video sequences. It utilizes voiceprint consistency as the criterion for determining if there has been a change in the active speaker. Our research focuses on evaluating the performance and computational costs of different voiceprint features and recognition models in speaker identification tasks. Additionally, E-TALK introduces a potential speaker tracking scheme for fixed-angle cameras, in conjunction with foreground extraction algorithms. Moreover, E-TALK incorporates a cloud-based high-precision facial ASD model, which utilizes historical information to determine the active speaker in real-time. We conducted experiments to evaluate the performance of our proposed solution in various scenarios and settings. The results demonstrate the effectiveness of the E-TALK approach in improving active speaker detection, highlighting its potential for practical application in surveillance systems.
Xiaojing Yu, Lan Zhang 0002, Xiang-Yang Li 0001
SECON1
2023 A drug molecular classification model based on graph structure generation
Lixuan Che, Yide Jin, Yuliang Shi, Xiaojing Yu, Hongfeng Sun
J. Biomed. Informatics4
2022 ANTIGONE: Accurate Navigation Path Caching in Dynamic Road Networks leveraging Route APIs
abstract
Navigation paths and corresponding travel times play a key role in location-based services (LBS) of which large-scale navigation path caching constitutes a fundamental component. In view of the highly dynamic real-time traffic changes in road networks, the main challenge amounts to updating paths in the cache in a fashion that incurs minimal costs due to querying external map service providers and cache maintenance. In this paper, we propose a hybrid graph approach in which an LBS provider maintains a dynamic graph with edge weights representing travel times, and queries the external map server so as to ascertain high fidelity of the cached paths subject to stringent limitations on query costs. We further deploy our method in one of the biggest on-demand food delivery platforms and evaluate the performance against state-of-the-art methods. Our experimental results demonstrate the efficacy of our approach in terms of both substantial savings in the number of required queries and superior fidelity of the cached paths.
Xiaojing Yu, Xiang-Yang Li 0001, Guobin Shen, Nikolaos M. Freris, Lan Zhang 0002
INFOCOM1
2021 Expanding, Retrieving and Infilling: Diversifying Cross-Domain Question Generation with Flexible Templates
abstract
Sequence-to-sequence based models have recently shown promising results in generating high-quality questions.However, these models are also known to have main drawbacks such as lack of diversity and bad sentence structures.In this paper, we focus on question generation over SQL database and propose a novel framework by expanding, retrieving, and infilling that first incorporates flexible templates with a neural-based model to generate diverse expressions of questions with guidance of sentence structure.Furthermore, a new activation/deactivation mechanism is proposed for template-based sequenceto-sequence generation, which learns to discriminate template patterns and content patterns, thus further improves generation quality.We conduct experiments on two largescale cross-domain datasets.The experiments show that the superiority of our question generation method in producing more diverse questions while maintaining high quality and consistency under both automatic evaluation and human evaluation.
Xiaojing Yu, Anxiao Jiang
EACL1
2020 Entropy Repulsion for Semi-supervised Learning Against Class Mismatch
Xuanke You, Lan Zhang 0002, Linzhuo Yang, Xiaojing Yu, Kebin Liu 0001
ICONIP (2)4
2020 Dataset and Enhanced Model for Eligibility Criteria-to-SQL Semantic Parsing
abstract
Clinical trials often require that patients meet eligibility criteria (e.g., have specific conditions) to ensure the safety and the effectiveness of studies. However, retrieving eligible patients for a trial from the electronic health record (EHR) database remains a challenging task for clinicians since it requires not only medical knowledge about eligibility criteria, but also an adequate understanding of structured query language (SQL). In this paper, we introduce a new dataset that includes the first-of-its-kind eligibility-criteria corpus and the corresponding queries for criteria-to-sql (Criteria2SQL), a task translating the eligibility criteria to executable SQL queries. Compared to existing datasets, the queries in the dataset here are derived from the eligibility criteria of clinical trials and include Order-sensitive, Counting-based, and Boolean-type cases which are not seen before. In addition to the dataset, we propose a novel neural semantic parser as a strong baseline model. Extensive experiments show that the proposed parser outperforms existing state-of-the-art general-purpose text-to-sql models while highlighting the challenges presented by the new dataset. The uniqueness and the diversity of the dataset leave a lot of research opportunities for future improvement.
Xiaojing Yu, Tianlong Chen 0001, Zhengjie Yu, Xiaoqian Jiang, Anxiao Jiang
LREC1
2020 ThumbUp: Identification and Authentication by Smartwatch using Simple Hand Gestures
abstract
The widespread creative application and smart devices call for convenient and secure interaction with human users. We propose, design, and implement a smartwatch-based two-factor real-time identification and authentication system named ThumbUp, where smartwatch users can identify and authenticate themselves by some simple hand and finger gestures, such as thumb-up. ThumbUp leverages the signal collected from the Inertial Measurement Unit (IMU) in Commercial Off-The-Shelf (COTS) smart devices and discovers the unique fingerprint pattern produced by each user’s simple hand gestures using a carefully crafted deep learning model. We implement our system and conduct extensive experiments to evaluate its efficacy and efficiency with 65 different users over a period of more than 3 months. It reaches an accuracy of 97% for identification, and EER 0.014 for authentication using only one simple gesture. We also survey the users’ acceptance of our system and discuss how the proficiency of gestures affects authentication accuracy.
Xiaojing Yu, Zhijun Zhou, Mingxue Xu, Xuanke You, Xiang-Yang Li 0001
PerCom1
2020 XHAR: Deep Domain Adaptation for Human Activity Recognition with Smart Devices
abstract
To further improve the convenience and effectiveness of human computer interaction (HCI) with smart devices, human activity recognition (HAR) has been widely studied from various aspects. Unfortunately, deep learning based methods often suffer from either expensive labeling efforts or weak generalization ability. Inspired by recently developed domain adaptation strategies, we propose XHAR, a novel adversarial deep domain adaptation framework for HAR using smart devices, providing better device and user adaptation. XHAR first selects the most similar source dataset (with label), then extracts device and user independent spatial-temporal features through the combinations of Convolutional Neural Networks (CNN) and Bidirectional Gated Recurrent Units (BiGRU) feature extractors. Moreover, it removes the distribution discrepancy using multiple domain discriminators, and finally performs adaptation on the target dataset (without label) to obtain the predicted labels. We conduct extensive experiments on 50 users (i.e., of different ages, genders, and body shapes) and 4 smart devices with two kinds of datasets (i.e., gesture activities and sport activities). We compare our method with the source-only model and several state-of-the-art domain adaptation models. The results show that XHAR increases the classification accuracy by at least 4.81% (to 74.50%) on the adaptation between different users, and accordingly by at least 9.25% (to 69.23%) between different devices.
Zhijun Zhou, Yingtian Zhang, Xiaojing Yu, Panlong Yang, Xiang-Yang Li 0001, Hao Zhou 0001
SECON3
2006 Classification of protein quaternary structure by functional domain composition
abstract
BACKGROUND: The number and the arrangement of subunits that form a protein are referred to as quaternary structure. Quaternary structure is an important protein attribute that is closely related to its function. Proteins with quaternary structure are called oligomeric proteins. Oligomeric proteins are involved in various biological processes, such as metabolism, signal transduction, and chromosome replication. Thus, it is highly desirable to develop some computational methods to automatically classify the quaternary structure of proteins from their sequences. RESULTS: To explore this problem, we adopted an approach based on the functional domain composition of proteins. Every protein was represented by a vector calculated from the domains in the PFAM database. The nearest neighbor algorithm (NNA) was used for classifying the quaternary structure of proteins from this information. The jackknife cross-validation test was performed on the non-redundant protein dataset in which the sequence identity was less than 25%. The overall success rate obtained is 75.17%. Additionally, to demonstrate the effectiveness of this method, we predicted the proteins in an independent dataset and achieved an overall success rate of 84.11% CONCLUSION: Compared with the amino acid composition method and Blast, the results indicate that the domain composition approach may be a more effective and promising high-throughput method in dealing with this complicated problem in bioinformatics.
Xiaojing Yu
BMC Bioinform.1