Siyang Jiang

dblp:260/3679 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
abstract
Multimodal human action recognition (HAR) utilizes complementary data for activity classification. Built on traditional HAR tasks, recent advances in Large Language Models (LLMs) enable detailed descriptions and causal reasoning of human actions, advancing new tasks of human action understanding (HAU) and human action reasoning (HARn). However, most LLMs, especially multimodal Large Vision-Language Models (LVLMs), struggle with modalities other than RGB images, like depth, IMU, ormmWave, due to a lack of large-scale datasets in these task domains. Existing HAR datasets provide only coarse-grained annotations, in-sufficient for depicting the detailed action dynamics required in HAU and HARn tasks. Simply combining annotations and generating captions with LLMs often lacks necessary logical and spatiotemporal consistency. In this paper, we introduce CUHK-X, a large-scale multi-modal dataset and benchmarks for HAR, HAU, and HARn. It includes 64,267 samples of 40 actions performed by 30 participants across two indoor environments, covering diverse daily scenarios. To address the challenge of spatiotemporal inconsistencies in captions, we propose a prompt-based scene creation method that leverages LLMs to generate logically connected activity sequences. CUHK-X also includes three benchmarks with six tasks to evaluate state-of-the-art models. Experimental results show average accuracies of 76.52% for HAR, 40.76% for HAU, and 70.25% for HARn. This large-scale multimodal dataset aims to empower the research community to apply, develop, and adapt data-intensive learning techniques for a wide range of human activity-related tasks.
Siyang Jiang, Mu Yuan, Bufang Yang, Lilin Xu, Yang Li 0147, Yuting He 0006, Liran Dong, Wenrui Lu, Zhenyu Yan 0002, Xiaofan Jiang 0001, Wei Gao 0006, Hongkai Chen 0001, Guoliang Xing
MobiSys1
2025 Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
abstract
In-home resistance training (RT) is a convenient and effective way to maintain health and well-being. However, incorrect exercise execution can result in unintended muscle engagement and an increased risk of injury. Without access to professional coaching, an accurate muscle-aware motion feedback system becomes essential for safe and effective training. However, existing visual language models (VLMs) struggle to provide accurate and effective muscle-aware movement guidance due to their limited understanding of RT motion and the absence of related expert knowledge. In this work, we introduce Myo-Trainer, the first vision-based muscle-aware motion feedback system that uses explicit muscle-aware motion analysis and domain-specific expert knowledge to provide corrective guidance on muscle engagement and movement execution. Also, we propose a novel DAGCN-Former network that integrates both spatial and temporal modeling capabilities to capture the complex dynamics of human RT motion. Experiments involving 26 subjects and 1000+ minutes of RT demonstrate that Myo-Trainer improves the accuracy of motion analysis by 17.22%, achieves a 2.5x reduced inference latency and a BertScore of 85.88% of generated feedback compared to those provided by experienced certified trainers, outperforming existing solutions. Additionally, Myo-Trainer received higher satisfaction ratings from participants compared to other AI trainers and video tutorials, highlighting its potential for real-world applications.
Yuting He 0006, Xinyan Wang 0003, Mu Yuan, Bufang Yang, Siyang Jiang, Yihua Huang 0002, Doris Sau-Fung Yu, Guoliang Xing, Hongkai Chen 0001
MobiCom5
2025 Dual Alignment Framework for Few-shot Learning with Inter-Set and Intra-Set Shifts
abstract
Few-shot learning (FSL) aims to classify unseen examples (query set) into labeled data (support set) through low-dimensional embeddings. However, the diversity and unpredictability of environments and capture devices make FSL more challenging in real-world applications. In this paper, we propose Dual Support Query Shift (DSQS), a novel challenge in FSL that integrates two key issues: inter-set shifts (between support and query sets) and intra-set shifts (within each set), which significantly hinder model performance. To tackle these challenges, we introduce a Dual Alignment framework (DUAL), whose core insight is that clean features can improve optimal transportation (OT) alignment. Firstly, DUAL leverages a robust embedding function enhanced by a repairer network trained with perturbed and adversarially generated “hard” examples to obtain clean features. Additionally, it incorporates a two-stage OT approach with a negative entropy regularizer, which aligns support set instances, minimizes intra-class distances, and uses query data as anchor nodes to achieve effective distribution alignment. We provide a theoretical bound of DUAL and experimental results on three image datasets, compared against 10 state-of-the-art baselines, showing that DUAL achieves a remarkable average performance improvement of 25.66%. Our code is available at https://github.com/siyang-jiang/DUAL.
Siyang Jiang, Rui Fang 0002, Hsi-Wen Chen, Guoliang Xing, Ming-Syan Chen
NeurIPS1
2025 ContextAgent: Context-Aware Proactive LLM Agents with Open-world Sensory Perceptions
abstract
Recent advances in Large Language Models (LLMs) have propelled intelligent agents from reactive responses to proactive support. While promising, existing proactive agents either rely exclusively on observations from enclosed environments (e.g., desktop UIs) with direct LLM inference or employ rule-based proactive notifications, leading to suboptimal user intent understanding and limited functionality for proactive service. In this paper, we introduce ContextAgent, the first context-aware proactive agent that incorporates extensive sensory contexts surrounding humans to enhance the proactivity of LLM agents. ContextAgent first extracts multi-dimensional contexts from massive sensory perceptions on wearables (e.g., video and audio) to understand user intentions. ContextAgent then leverages the sensory contexts and personas from historical data to predict the necessity for proactive services. When proactive assistance is needed, ContextAgent further automatically calls the necessary tools to assist users unobtrusively. To evaluate this new task, we curate ContextAgentBench, the first benchmark for evaluating context-aware proactive LLM agents, covering 1,000 samples across nine daily scenarios and twenty tools. Experiments on ContextAgentBench show that ContextAgent outperforms baselines by achieving up to 8.5% and 6.0% higher accuracy in proactive predictions and tool calling, respectively. We hope our research can inspire the development of more advanced, human-centric, proactive AI assistants. The code and dataset are publicly available at https://github.com/openaiotlab/ContextAgent.
Bufang Yang, Lilin Xu, Liekang Zeng, Kaiwei Liu 0001, Siyang Jiang, Wenrui Lu, Hongkai Chen 0001, Xiaofan Jiang 0001, Guoliang Xing, Zhenyu Yan 0002
NeurIPS5
2025 SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer Serving
abstract
Transformer models have revolutionized fields like natural language processing and computer vision but face privacy concerns in sensitive applications such as medical diagnostics. Existing confidential serving methods, including cryptography-based, memory isolation-based, and access control-based, offer trade-offs between privacy and efficiency but often struggle with high latency or hardware dependencies. This work proposes stateless KV-cache encoding (SCX), a novel framework that encodes the intermediate key-value cache during Transformer inference using user-controlled keys. SCX ensures that the cloud can neither recover the input nor independently complete the next token prediction, effectively preserving privacy. By introducing efficient encoding and decoding schemes, SCX addresses communication complexity and attack vulnerabilities while ensuring zero loss of inference quality. Experiments on large Transformer models demonstrate that SCX achieves lower latency (e.g., 36ms for LLaMA-7B), outperforming state-of-the-art cryptography and memory isolation methods by orders of magnitude. Moreover, SCX can complementarily work with advanced KV-cache management techniques to further enhance KV-cache communication efficiency by 85%, marking a significant step toward practical, privacy-preserving large Transformer serving.
Mu Yuan, Lan Zhang 0002, Liekang Zeng, Siyang Jiang, Bufang Yang, Di Duan, Guoliang Xing
SIGCOMM4
2025 A Safe-Critical and Efficient Self-Merging Strategy for CAVs in Mixed Traffic Scenarios
abstract
Connected and autonomous vehicles (CAVs) are emerging as a potential solution to merging safety problems. However, in mixed traffic scenarios where CAVs coexist with human-driven vehicles (HVs), challenges arise due to the lack of proactive cooperation and the limited length of the acceleration lane, complicating the merging processes of CAVs. In these cases, CAVs should actively seize the transit opportunity and perform safe and efficient merging. Failure to do so can lead to decreased traffic efficiency, increased fuel consumption and emissions, compromised self-merging capacities, and heightened crash risk. Therefore, this article employs the roadside unit and proposes a two-level hierarchical self-merging strategy for CAVs to increase the merging efficiency while ensuring high safety. Since the surrogate safety measures (SSMs) can formulate reliable safety assessment and identify the merging conflict risk by setting appropriate threshold, the upper level uses a novel SSM-based method (i.e., the minimum acceleration rate, MIAR) to determine the merging sequence (MS). A theoretical model for merging safety (TMMS) is developed to estimate the MIAR value, and the MIAR threshold is determined using signal detection theory (SDT). At the lower level, the strategy recommends the optimized merging maneuvers, pregenerated using sequential quadratic programming-model predictive control (SQP-MPC), based on the determined MS and the CAV’s velocity. A case study at a real-world freeway merging area demonstrates the effectiveness of the MIAR in measuring merging conflict risk. Besides, numerous simulations are conducted, and the results demonstrate that the proposed strategy significantly improves merging success rate and overall traffic efficiency.
Siyang Jiang, Menglu Gu, Yanqi Su, Chang Wang 0002, Wenhui Wei
IEEE Internet Things J.1
2024 A Survey on Model-heterogeneous Federated Learning: Problems, Methods, and Prospects
abstract
As privacy concerns continue to grow, federated learning (FL) has gained significant attention as a promising privacy-preserving technology, leading to considerable advancements in recent years. Unlike traditional machine learning, which requires central data collection, FL keeps data localized on user devices. However, conventional FL assumes that all clients operate with identical model structures initialized by the server. In real-world applications, system heterogeneity is common, with clients possessing varying computational capabilities. This disparity can hinder training for resource-limited clients and result in inefficient resource use for those with greater processing power. To address this challenge, model-heterogeneous FL has been introduced, enabling clients to train models of varying complexity based on their hardware resources. This paper reviews state-of-the-art approaches in model-heterogeneous FL, analyzing their strengths and weaknesses, while identifying open challenges and future research directions. To the best of our knowledge, this is the first survey to specifically focus on model-heterogeneous FL.
Boyu Fan, Siyang Jiang, Xiang Su 0001, Sasu Tarkoma, Pan Hui 0001
IEEE Big Data2
2024 ArtFL: Exploiting Data Resolution in Federated Learning for Dynamic Runtime Inference via Multi-Scale Training
abstract
Federated Learning (FL) has emerged as a prominent paradigm for distributed machine learning, crucial for mission-critical applications such as autonomous driving and smart health. However, existing FL systems have not adequately addressed the dynamic real-time requirements of these applications due to stringent inference deadlines and resource limitations on edge devices. In this paper, we propose ArtFL, a novel federated learning system designed to support dynamic runtime inference through multi-scale training. The key idea of ArtFL is to utilize the data resolution, i.e., frame resolution of videos, as a knob to accommodate dynamic inference latency requirements. Specifically, we initially propose data-utility-based multi-scale training, allowing the trained model to process data of varying resolutions during inference. Subsequently, we introduce an innovative strategy for frame resolution selection in inference, based on the similarity of adjacent frames. Finally, leveraging latency-based dynamic data dropping, we propose a systematic scheme to reduce the overall training time by shortening the waiting time in FL. For evaluation, we build two real-world FL testbeds for smart vehicles and healthcare applications, utilizing a heterogeneous edge platform. Extensive experiments across our testbeds and three public datasets show that ArtFL outperforms state-of-the-art baselines in overall accuracy and system performance up to 36.36% and 47.81%, respectively. A demo video of ArtFL on our smart vehicle testbed is available at https://youtu.be/eeK6yRVEG3U, and our code is available at https://github.com/siyang-jiang/ArtFL.git.CCS CONCEPTS• Computing methodologies → Machine learning.
Siyang Jiang, Xian Shuai, Guoliang Xing
IPSN1
2024 Demo Abstract: CaringFM: An Interactive In-home Healthcare System Empowered by Large Foundation Models
abstract
The demand for fully on-device health monitoring is huge and urgent. However, deploying Large Foundation Models conventionally relies on cloud-based computing services, which poses privacy concerns. Driven by the belief of delivering personalised healthcare to family members, this study presents the development of an innovative on-device machine learning system, CaringFM. This family caring system utilizes privacy-protecting sensors and an edge-deployed Foundation Model(FM) to offer a convenient and low-cost solution for chronic disease prediction and health condition monitoring at home. In particular, CaringFM provides general health suggestions and personalized medical information while ensuring high privacy by processing and preserving all data locally.
Kaiwei Liu 0001, Siyang Jiang, Zhenyu Yan 0002, Guoliang Xing
IPSN3
2024 Poster Abstract: Threshold Cryptography-based Authentication Protocol for Remote Healthcare
abstract
With the advancement of the Internet of Medical Things (IoMT) and cryptographic technologies, remote healthcare services have become more widespread, presenting new challenges for patient privacy and data security. Conventional security mechanisms, such as centralized authentication and key distribution systems, are susceptible to single points of failure and significant management burdens, potentially leading to compromised authentication centers and internal security threats. In response, this study presents a threshold signature algorithm, it uses Distributed Key Generation (DKG) that distributes private keys without the need for a trusted key distributor, requiring the cooperative signature of at least two nodes for authentication. This approach not only circumvents the risk of single points of failure but also enhances the system’s robustness and efficiency. The experimental results validate its prospective utility in safeguarding remote healthcare data.
Qipeng Xie, Linshan Jiang, Siyang Jiang, Salabat Khan, Weizheng Wang 0001, Kaishun Wu
IPSN4
2024 Efficiency Optimization Techniques in Privacy-Preserving Federated Learning With Homomorphic Encryption: A Brief Survey
abstract
Federated learning (FL) offers distributed machine learning on edge devices. However, the FL model raises privacy concerns. Various techniques, such as homomorphic encryption (HE), differential privacy, and multiparty cooperation, are used to address the privacy issues of the FL model. Among them, HE ensures greater security and privacy since end-to-end encryption maintains data privacy throughout the computation process. Compared with other privacy-preserving techniques, HE does not require the establishment of a trusted environment or protocol among multiple parties and does not involve any artificial noise that can impair system performance. Unfortunately, it suffers from efficiency overhead when applied to privacy-preserving FL (PPFL). Some existing surveys on PPFL discuss the generic construction and organization of PPFL from the perspective of practical HE deployment in PPFL. However, none of them covers the efficiency optimization of HE when applied to PPFL. This article conducts a comprehensive review of the efficiency optimization of HE when applied to PPFL. First, we review general optimization strategies and discuss their limitations when applied directly to HE-based PPFL. Second, an overview of algorithmic, hardware, and hybrid optimizations is provided, along with a discussion of their adaptation. Additionally, we provide a detailed taxonomy of optimizations. Finally, we suggest future HE-based PPFL research directions.
Qipeng Xie, Siyang Jiang, Linshan Jiang, Yongzhi Huang 0002, Salabat Khan, Wangchen Dai, Zhe Liu 0001, Kaishun Wu
IEEE Internet Things J.2
2023 Incremental Reinforcement Learning with Dual-Adaptive ε-Greedy Exploration
abstract
Reinforcement learning (RL) has achieved impressive performance in various domains. However, most RL frameworks oversimplify the problem by assuming a fixed-yet-known environment and often have difficulty being generalized to real-world scenarios. In this paper, we address a new challenge with a more realistic setting, Incremental Reinforcement Learning, where the search space of the Markov Decision Process continually expands. While previous methods usually suffer from the lack of efficiency in exploring the unseen transitions, especially with increasing search space, we present a new exploration framework named Dual-Adaptive ϵ-greedy Exploration (DAE) to address the challenge of Incremental RL. Specifically, DAE employs a Meta Policy and an Explorer to avoid redundant computation on those sufficiently learned samples. Furthermore, we release a testbed based on a synthetic environment and the Atari benchmark to validate the effectiveness of any exploration algorithms under Incremental RL. Experimental results demonstrate that the proposed framework can efficiently learn the unseen transitions in new environments, leading to notable performance improvement, i.e., an average of more than 80%, over eight baselines examined.
Siyang Jiang, Hsi-Wen Chen, Ming-Syan Chen
AAAI2
2023 Poster Abstract: CNN-guardian: Secure Neural Network Inference Acceleration on Edge GPU
abstract
The rapid development of AI applications powered by deep learning in edge devices boosts the opportunity for real-time health monitoring. To address the potential privacy concern in the inference phase, homomorphic encryption (HE) is an alternative solution that encrypts inference data without exposing raw data and has several distinct advantages, (i.e., single-round communication, lightweight bandwidth consumption, and non-interactive computation). However, the computational overhead on the current HE-based privacy-preserving inference necessitates a substantial amount of time, which is not feasible for some real-time applications on edge devices. To address this issue, we propose CNN-guardian, a unified and compact neural network structure for real-time inference in HE-based inference on edge GPU. CNN-guardian designs a HE-friendly neural network and GPU engine that optimizes HE operations to accelerate the inference in the HE domain.
Qipeng Xie, Hao Yang 0062, Linshan Jiang, Siyang Jiang, Shiyu Shen 0001, Salabat Khan, Zhe Liu 0001, Kaishun Wu
SenSys5
2023 Pairwise learning problems with regularization networks and Nyström subsampling approach
Ting Hu 0002, Siyang Jiang
Neural Networks3
2022 Dual-Triangular QR Decomposition with Global Acceleration and Partially Q-Rotation Skipping
abstract
Efficient matrix operations have been deemed keys to efficient data analysis. Dual-Triangular QR Decomposition (DT-QRD) is a critical component in Tall and skinny QR decomposition (TS-QRD), which is a widely-used matrix operation with various applications, such as data compression and feature extraction. In order to accelerate DT-QRD, in this paper, we propose a new acceleration framework, including Global Acceleration Schemes, and Partially$\boldsymbol{Q}$-rotation Skipping, which utilize the special DT structure in both$\mathbf{Q}$and$\mathbf{R}$matrix to reduce the latency and computation resource. Further, we employ the Systolic-Array Based Architecture (1D & 2D) for implementation to reduce the memory usage. Experimental results manifest that our framework achieves$169.70\times\ (\mathbf{1}\mathbf{D})$and$250.13\times\ (\mathbf{2}\mathbf{D})$speedup.
Rui Fang 0002, Siyang Jiang, Hsi-Wen Chen, Ming-Syan Chen
FPT2
2022 BalanceFL: Addressing Class Imbalance in Long-Tail Federated Learning
abstract
Federated Learning (FL) is an emerging learning paradigm that enables the collaborative learning of different nodes without ex-posing the raw data. However, a critical challenge faced by the current federated learning algorithms in real-world applications is the long-tailed data distribution, i.e., in both local and global views, the numbers of classes are often highly imbalanced. This would lead to poor model accuracy on some rare but vital classes, e.g., those related to safety in health and autonomous driving applications. In this paper, we propose BalanceFL, a federated learning frame-work that can robustly learn both common and rare classes from a long-tailed real-world dataset, addressing both the global and local data imbalance at the same time. Specifically, instead of letting nodes upload a class-drifted model trained on imbalanced private data, we design a novel local update scheme that rectifies the class imbalance, forcing the local model to behave as if it were trained on ideal uniform distributed data. To evaluate the performance of BalanceFL, we first adapt two public datasets to the long-tailed federated learning setting, and then collect a real-life IMU dataset for action recognition, which includes more than 10,000 data sam-ples and naturally exhibits the global long tail effect and the local imbalance. On all of these three datasets, BalanceFL outperforms state-of-the-art federated learning approaches by a large margin.
Xian Shuai, Yulin Shen 0001, Siyang Jiang, Zhenyu Yan 0002, Guoliang Xing
IPSN3
2022 PGADA: Perturbation-Guided Adversarial Alignment for Few-Shot Learning Under the Support-Query Shift
Siyang Jiang, Hsi-Wen Chen, Ming-Syan Chen
PAKDD (1)1
2021 Dataflow Systolic Array Implementations of Exploring Dual-Triangular Structure in QR Decomposition Using High-Level Synthesis
abstract
Tall and skinny QR (TSQR) decomposition is an essential matrix operation with various applications in edge computing, including data compression, subspace projection, and dimension reduction. As a critical component in TSQR, Dual-Triangular QR (DTQR) decomposition is solved by the Normal QR method in most works without utilizing the dual-triangular structure. Therefore, we propose a novel DTQR accelerator by recursively exploring the DT structure and propose three acceleration strategies with the systolic array to achieve higher parallelism. Experimental results manifest that our algorithm achieves 21.55x on average speedup compared with the baselines.
Siyang Jiang, Hsi-Wen Chen, Ming-Syan Chen
FPT1
2020 Defect Detection of Pantograph Slide Based on Deep Learning and Image Processing Technology
abstract
Pantograph is one of the most important components in electrical railway vehicles. To guarantee steady power supply for the train, the surface of the pantograph slide plate should be smooth enough so that the catenary can move on it from one side to the other side steadily with low friction. In addition, the thickness of the pantograph slide plate cannot be smaller than the lower limit for the sake of safety. Therefore, periodical inspection and maintenance of the pantograph slide plate are significant in terms of safe and stable operation. In this paper, an innovative and intelligent method based on deep learning and image processing technologies is proposed for the online condition monitoring of the pantograph slide plate. In the first stage, the surface defect detection and recognition method of the pantograph slide plate is proposed. Four typical surface defects of the slide are considered, and a deep learning model, pantograph defect detection neural network (PDDNet), is trained for the defect detection and recognition. In the second stage, five key criteria for qualifying the wear condition are proposed. The wear edge estimation based on image processing technology is investigated in detail. Furthermore, they are used to calculate the wear depth and evaluate the wear condition of the pantograph slide. The experiment results demonstrate that the proposed PDDNet can detect the surface defects and also recognize the four kinds of defects with a sound accuracy. The wear depth estimation results are compared with on-site measurement data, and the proposed method can achieve high estimation accuracy.
Xiukun Wei, Siyang Jiang, Chenliang Li 0004, Limin Jia 0002, Yongguang Li
IEEE Trans. Intell. Transp. Syst.2