Huanghuang Liang

dblp:228/4696 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-2847-0285ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 REMISVFU: Vertical Federated Unlearning via Representation Misdirection for Intermediate Output Feature
abstract
Data-protection regulations such as the GDPR grant every participant in a federated system a right to be forgotten. Federated unlearning has therefore emerged as a research frontier, aiming to remove a specific party's contribution from the learned model while preserving the utility of the remaining parties. However, most unlearning techniques focus on Horizontal Federated Learning (HFL), where data are partitioned by samples. In contrast, Vertical Federated Learning (VFL) allows organizations that possess complementary feature spaces to train a joint model without sharing raw data. The resulting feature-partitioned architecture renders HFL-oriented unlearning methods ineffective. In this paper, we propose ReMisVFU, a plug-and-play representation-misdirection framework that enables fast, client-level unlearning in splitVFL systems. When a deletion request arrives, the forgetting party collapses its encoder output to a randomly sampled anchor on the unit sphere, severing the statistical link between its features and the global model. To maintain utility for the remaining parties, the server jointly optimizes a retention loss and a forgetting loss, aligning their gradients via orthogonal projection to eliminate destructive interference. Evaluations on public benchmarks show that ReMisVFU suppresses back-door attack success to the natural class-prior level and sacrifices only about 2.5% points of clean accuracy, outperforming state-of-the-art baselines.
Huanghuang Liang, Yili Gong, Jiawei Jiang 0001, Chuang Hu, Dazhao Cheng
AAAI3
2026 Nexus: Communication-Aware Role Differentiation for Adaptive Multi-Robot Exploration
Rui Ge 0010, Huanghuang Liang, Jianqi Ma, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng
INFOCOM2
2026 Metadata-guided multi-task transfer learning for thickness deviation detection in aluminum cold rolling: System design and real-world deployment
Rui Ge 0010, Huanghuang Liang, Qing Shen 0001, Jiawei Jiang 0001, Chuang Hu, Dazhao Cheng
Expert Syst. Appl.3
2026 FlePo: GPU Multitask Scheduling Optimization Framework for Dynamic Scenes
abstract
Deep Neural Networks (DNNs) are widely used in intelligent applications, driving increasing computational demands on GPUs. However, modern GPU multitasking scheduling algorithms fail to effectively balance real-time task performance and resource utilization, especially under dynamic workloads with highly variable DNN computational demands. The complex and workload-dependent execution times of DNN kernels often lead to inefficient resource allocation, degraded system throughput, and missed real-time constraints. To address these challenges, we propose Flexible Parallel Orchestrator (FlePo), a GPU multitasking scheduling framework designed to optimize resource utilization and maintain real-time task performance within acceptable limits for soft real-time systems. FlePo integrates two key techniques: Adaptive Padding Dispatch (APD), which dynamically schedules best-effort tasks while leveraging the predictable execution characteristics of DNN kernels to maintain real-time predictability; and Dynamic Parallel Fusion (DPF), which employs kernel fusion to create computational isolation, reducing interference in parallel job execution. By combining offline profiling with online adaptation, FlePo efficiently responds to workload variations. We evaluate FlePo on two heterogeneous GPU platforms, NVIDIA Tesla V100 and AMD MI50, achieving up to a 50% increase in throughput while keeping real-time overhead below 2%. Our work enhances GPU multitasking in dynamic environments, with potential applications in autonomous driving, smart homes, and intelligent healthcare.
Huanghuang Liang, Rui Ge 0010, Yaqi Xia, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng
ACM Trans. Auton. Adapt. Syst.1
2025 Understanding the Challenges Students Face in Non-English Programming Environments Due to the Programming Language Transition: A Case Study of Keywords in the Chinese Version of Scratch
Janice Jianing Si, Huanghuang Liang, Chuang Hu, Yujun Zhu, Xiaobo Zhou 0002, Kanye Ye Wang, Dazhao Cheng
CHI3
2025 LightTrace: A Versatile Ebpf-Enabled Toolkit for Lightweight Distributed Tracing
abstract
Distributed tracing is widely employed for troubleshooting distributed systems such as microservices. Existing tracing systems typically improve one or more of data completeness, non-intrusiveness, or lightweight operation through various data generation strategies. However, no current solution optimizes all three aspects simultaneously, which can introduce significant overhead in I/O-intensive environments that demand both high completeness and minimal intrusion. In this paper, we introduce LightTrace, a novel, eBPF-enabled toolkit that optimizes the data transmission mechanism in distributed tracing. LightTrace integrates seamlessly with mainstream tracing systems and leverages eBPF to reduce end-to-end latency and lower overall system overhead. LightTrace is implemented using a combination of kernel-level eBPF and user-space Golang components. Our evaluation demonstrates that LightTrace decreases the average latency overhead by up to 22.8 % and improves the peak throughput of microservice systems by between$\mathbf{1 2. 3 \%}$and$\mathbf{1 8. 6 \%}$. Furthermore, LightTrace's adaptability across diverse tracing platforms underscores its versatility in various microservice environments.
Yanze Zhang, Kanye Ye Wang, Shufan Gong, Huanghuang Liang, Chuang Hu, Xiaobo Zhou 0002
ICPADS4
2025 Zero-shot Federated Unlearning via Transforming from Data-Dependent to Personalized Model-Centric
abstract
Federated Unlearning (FU) addresses the "right to be forgotten" in federated learning by removing specific client data's contribution without retraining from scratch. Existing FUs are data-dependent, which make the assumption that systems can access original training data or stored historical parameter updates during unlearning. However, the assumption cannot always hold in practice, as users usually request the deletion of client data and historical parameter updates due to privacy concerns or storage limitations. Therefore, it is crucial to develop a zero-shot FU method without such data access. The key challenge is how to distinguish and remove the impact of target clients without data-level information. Motivated by the idea that if we can learn client-specific personalized information from the model instead of data, FU can be model-centric and data-free, we present the first zero-shot FU framework ZeroFU. By embedding client contributions into the model during learning via condition computation, ZeroFU enables the model to possess personalized features for unlearning. The unlearning is achieved using a proposed GAN-based distillation framework that obfuscates the personalized feature of the target client. Evaluations demonstrate its effectiveness in unlearning under non-IID settings.
Huanghuang Liang, Jingling Yuan, Jiawei Jiang 0001, Kanye Ye Wang, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng
IJCAI2
2025 Streamlining Data Transfer in Collaborative SLAM Through Bandwidth-Aware Map Distillation
abstract
Edge intelligence offers a promising solution for Simultaneous Localization and Mapping (SLAM) in large-scale scenarios, where multiple robots collaboratively perceive the environment and upload their local maps to an edge server. However, maintaining mapping accuracy under constrained and dynamic communication resources remains a significant challenge for the practical deployment of robot swarms. Concurrent data uploads from multiple agents can exacerbate network congestion, leading to the loss of critical information, delayed updates, and, ultimately, the inconsistency of the generated maps. This paper presents Hermes, an edge-assisted collaborative mapping system designed for communication-constrained environments. Hermes streamlines data transfer through bandwidth-aware map distillation, ensuring only the most crucial messages are transmitted to the edge server. We quantify the importance of keyframes and landmarks based on their information entropy gain in pose estimation. By selectively sharing essential submaps, Hermes adaptively balances communication bandwidth and information richness during the mapping process. We implemented Hermes on heterogeneous platforms and conducted experiments using public datasets and self-collected campus data. Hermes exceeds SwarmMap by 50% in bandwidth utilization with similar accuracy and surpasses COVINS-G by 65% in trajectory error under highly constrained network resources.
Rui Ge 0010, Huanghuang Liang, Chuang Hu, Xiaobo Zhou 0002, Dazhao Cheng
IEEE Trans. Mob. Comput.2
2025 Mimir: Data-Free Federated Unlearning Through Client-Specific Prompt Generation for Personalized Models
abstract
Federated unlearning (FU) has become an important area of research due to an increasing need for federated learning (FL) applications to comply with emerging data privacy regulations such as GDPR. It facilitates the removal of certain clients' data from an already trained FL model while preserving the performance on the remaining client without the need to retrain from scratch. Existing FU methods typically require clients to have access to their training data or historical model updates, which may be impractical in real-world scenarios due to privacy constraints and changes in data availability. Moreover, FU methods may cause catastrophic unlearning, where removing a client's data from heterogeneous, non-IID settings can negatively impact the model's performance on data from retained clients. To address the aforementioned issues and leverage the capabilities of personalized federated learning (pFL) in handling non-IID data distributions, this paper introduce Mimir, a novel data-free federated unlearning framework designed for pFL settings. Mimir integrates both learning and unlearning phases by utilizing personalized prompts for each client. We design a distillation structure based on Generative Adversarial Networks (GANs) for client-level unlearning that does not require access to original data or historical updates. By leveraging client-specific prompts generated during the pFL phase, Mimir adapts to heterogeneous data distributions and mitigates catastrophic unlearning on the retained data. We demonstrate the effectiveness of Mimir through extensive experiments on benchmark datasets, showing its ability to forget target client data while preserving model accuracy on the remaining clients.
Huanghuang Liang, Tianyu Tu, Jiawei Jiang 0001, Chuang Hu, Dazhao Cheng
IEEE Trans. Mob. Comput.2
2025 Spread+: Scalable Model Aggregation in Federated Learning With Non-IID Data
abstract
Federated learning (FL) addresses privacy concerns by training models without sharing raw data, overcoming the limitations of traditional machine learning paradigms. However, the rise of smart applications has accentuated the heterogeneity in data and devices, which presents significant challenges for FL. In particular, data skewness among participants can compromise model accuracy, while diverse device capabilities lead to aggregation bottlenecks, causing severe model congestion. In this article, we introduce Spread+, a hierarchical system that enhances FL by organizing clients into clusters and delegating model aggregation to edge devices, thus mitigating these challenges. Spread+ leverages hedonic coalition formation game to optimize customer organization and adaptive algorithms to regulate aggregation intervals within and across clusters. Moreover, it refines the aggregation algorithm to boost model accuracy. Our experiments demonstrate that Spread+ significantly alleviates the central aggregation bottleneck and surpasses mainstream benchmarks, achieving performance improvements of 49.58% over FAVG and 22.78% over Ring-allreduce.
Huanghuang Liang, Boan Liu, Chuang Hu, Dan Wang 0002, Xiaobo Zhou 0002, Dazhao Cheng
IEEE Trans. Parallel Distributed Syst.1
2024 A unified hybrid memory system for scalable deep learning and big data applications
Wei Rang, Huanghuang Liang, Kanye Ye Wang, Xiaobo Zhou 0002, Dazhao Cheng
J. Parallel Distributed Comput.2
2024 A Survey on Spatio-Temporal Big Data Analytics Ecosystem: Resource Management, Processing Platform, and Applications
abstract
With the rapid evolution of the Internet, Internet of Things (IoT), and geographic information systems (GIS), spatio-temporal Big Data (STBD) is experiencing exponential growth, marking the onset of the STBD era. Recent studies have concentrated on developing algorithms and techniques for the collection, management, storage, processing, analysis, and visualization of STBD. Researchers have made significant advancements by enhancing STBD handling techniques, creating novel systems, and integrating spatio-temporal support into existing systems. However, these studies often neglect resource management and system optimization, crucial factors for enhancing the efficiency of STBD processing and applications. Additionally, the transition of STBD to the innovative Cloud-Edge-End unified computing system needs to be noticed. In this survey, we comprehensively explore the entire ecosystem of STBD analytics systems. We delineate the STBD analytics ecosystem and categorize the technologies used to process GIS data into five modules: STBD, computation resources, processing platform, resource management, and applications. Specifically, we subdivide STBD and its applications into geoscience-oriented and human-social activity-oriented. Within the processing platform module, we further categorize it into the data management layer (DBMS-GIS), data processing layer (BigData-GIS), data analysis layer (AI-GIS), and cloud native layer (Cloud-GIS). The resource management module and each layer in the processing platform are classified into three categories: task-oriented, resource-oriented, and cloud-based. Finally, we propose research agendas for potential future developments.
Huanghuang Liang, Zheng Zhang 0036, Chuang Hu, Yili Gong, Dazhao Cheng
IEEE Trans. Big Data1
2024 Corrections to "DNN Surgery: Accelerating DNN Inference on the Edge through Layer Partitioning"
abstract
In this paper, we reference the previous conference version and complete the grant number mentioned in the acknowledgments of the conference version.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003
IEEE Trans. Cloud Comput.1
2024 Controlling Aluminum Strip Thickness by Clustered Reinforcement Learning With Real-World Dataset
abstract
Consistent thickness in aluminum strips stands as a pivotal indicator of aluminum sheet product quality. Conventional automatic gauge control systems are complex, multivariable, and strongly coupled. However, the rolling process faces uncertainties, preventing the establishment of precise mathematical models. To tackle this, we propose an aluminum strip cold rolling thickness control method grounded in offline reinforcement learning. To facilitate the learning of better control policies, we construct a dataset of aluminum strip cold rolling process control data derived from real-world historical records and expertise for offline policy training, which is named dataset for aluminum strip cold rolling and comprises 8 373 540 Markov decision process tuples. We employ a clustered approach to handle time-varying production conditions. A constrained filtering scheme is introduced to eliminate problematic data after a data-driven ensemble rolling model is established. Evaluation and case study demonstrate that our method effectively reduces aluminum strip thickness deviations without requiring prior knowledge, thus improving control performance.
Ziqi Xiao, Huanghuang Liang, Chuang Hu, Dazhao Cheng
IEEE Trans. Ind. Informatics3
2023 TAPU: A Transmission-Analytics Processing Unit for Accelerating Multifunctions in IoT Gateways
abstract
Internet of Things (IoT) gateways integrate various sensors and compute initial decisions before transmitting data to the cloud for further processing. As the functions they need to support become increasingly complex, gateways must upgrade their hardware. Network functions (NF) and video analytics (VAs) are two typical examples of hardware requirements: NFs need specialized hardware accelerators, while VAs need parallel processing power. However, gateways are typically constrained by factors, such as power, size, and cost, leading to a need to multiplex functions and minimize hardware overprovisioning. This article proposes a novel accelerator, the transmission-analytic processing unit (TAPU), which uses multi-image FPGA to accelerate VAs and NFs for IoT gateways. We preconfigure one image for VAs and one image for NFs, then multiplex the FPGA resources in the time dimension. The TAPU system design requires both hardware and software revisions. In the hardware design, we discuss our considerations on hardware choice and present a new abstraction of hardware functions to overcome the challenge of application development on different multi-image FPGAs. For the software, we develop a fully functional TAPU system to adapt to dynamic network and VAs workloads. Our evaluation shows that TAPU utilization can reach 92%, considerably increasing VAs and network processing throughput over the current approach. We further evaluate TAPU through two case studies that support a campus traffic monitoring system and an office surveillance system, demonstrating excellent performance improvement and low overhead.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Yili Gong, Dazhao Cheng, Xiaobo Zhou 0002, Yu Wang 0003
IEEE Internet Things J.1
2023 An Edge-Side Real-Time Video Analytics System With Dual Computing Resource Control
abstract
Video analytics systems conduct video preprocessing to filter out unnecessary frames and model inference using appropriately selected neural networks for high analytics speed. Video preprocessing is instruction-intensive computing (IIC) executed by CPU, and model inference is data-intensive computing (DIC) executed by GPU. In this paper, we show the analytics accuracy of existing systems can largely vary in fields, caused by thedynamicIIC and DIC workloads of differentcontentsin applications. Unfortunately, cameras havefixedCPU/GPU resources and cannot effectively adapt to workload dynamics. We develop Gemini, a new edge-side real-time video analytics system enhanced by a dual-image FPGA. We take the advantage of negligible image switching time of dual-image FPGAs, pre-configure one CPU image and one GPU image and elastically multiplex the dual CPU-GPU resources intimedimension. Gemini requires both hardware and software revisions. In hardware, we overcome challenges of hardware-dependent application development, low communication efficiency between the microprocessor and FPGA, and high programming complexity by hardware abstraction, asynchronous data transfer mechanism and stub-skeleton middleware. In software, we overcome the challenge of adapting to the dynamic workloads by a bandit learning approach. We implement Gemini and show that Gemini can improve the analytics accuracy to 90.35%.
Chuang Hu, Qianlong Sang, Huanghuang Liang, Dan Wang 0002, Dazhao Cheng, Jin Zhang 0001, Qing Li 0006, Junkun Peng
IEEE Trans. Computers4
2023 DNN Surgery: Accelerating DNN Inference on the Edge Through Layer Partitioning
abstract
Recent advances in deep neural networks have substantially improved the accuracy and speed of various intelligent applications. Nevertheless, one obstacle is that DNN inference imposes a heavy computation burden on end devices, but offloading inference tasks to the cloud causes a large volume of data transmission. Motivated by the fact that the data size of some intermediate DNN layers is significantly smaller than that of raw input data, we designed the DNN surgery, which allows partitioned DNN to be processed at both the edge and cloud while limiting the data transmission. The challenge is twofold: (1) Network dynamics substantially influence the performance of DNN partition, and (2) State-of-the-art DNNs are characterized by a directed acyclic graph rather than a chain, so that partition is incredibly complicated. To solve the issues, We design a Dynamic Adaptive DNN Surgery(DADS) scheme, which optimally partitions the DNN under different network conditions. We also study the partition problem under the cost-constrained system, where the resource of the cloud for inference is limited. Then, a real-world prototype based on the selif-driving car video dataset is implemented, showing that compared with current approaches, DNN surgery can improve latency up to 6.45 times and improve throughput up to 8.31 times. We further evaluate DNN surgery through two case studies where we use DNN surgery to support an indoor intrusion detection application and a campus traffic monitor application, and DNN surgery shows consistently high throughput and low latency.
Huanghuang Liang, Qianlong Sang, Chuang Hu, Dazhao Cheng, Xiaobo Zhou 0002, Dan Wang 0002, Wei Bao 0001, Yu Wang 0003
IEEE Trans. Cloud Comput.1
2022 Spread: Decentralized Model Aggregation for Scalable Federated Learning
abstract
Federated learning (FL) is a new distributed machine learning paradigm that enables machine learning on edge devices. One unique feature of FL is that edge devices belong to individuals; and since they are not “owned” by the FL coordinator, but can be “federated” instead, there can potentially be a huge number of edge devices. In the current distributed ML architecture, the parameter server (PS) architecture, model aggregation is centralized. When facing a large number of edge devices, the centralized model aggregation becomes the bottleneck and fundamentally restricts system scalability.
Chuang Hu, Huanghuang Liang, Boan Liu, Dazhao Cheng, Dan Wang 0002
ICPP2
2018 Chaos Theory in Urban Traffic Flow: Is Crowd Sensed Data Driving the Macro-traffic Behavior to Oscillation or Equilibrium
abstract
Stability theory tells us that a dynamic system will eventually converge to its stable state, in which the system's overall energy is at its minimum. On the other hand, chaos theory states that small perturbations of the system are able to drive itself from previously-stable state to another state. This phenomenon has been observed in many fields like cosmetol- ogy, physics, biology and chemistry. Our research question is whether chaos theory also applies to the transportation domain. Specifically, when we are given imperfect or delayed crowd- sensed data, will we observe the cyclic/oscillatory transition between different traffic states? This paper aims at investigating this chaotic phenomenon (oscillatory traffic behavior in this paper) on urban transportation with imperfect or delayed crowd-sensed information and delivering recommendations for crowdsensing-based traffic applications to avoid the undesirable oscillations.
Huanghuang Liang, Lu Yang 0002, Jiacheng Wei, Hong Cheng 0002
Intelligent Vehicles Symposium1