Sicong Liu 0005

dblp:44/9804-5 · DBLP profile ↗
← Back
33ranked-venue papers
9as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 25 · 8 first-author · 22 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AdaSprite: Resource-efficient Online Co-Adaptation for V2I Systems Under Large-scale Data Drifts
abstract
The rise of vehicle-infrastructure (V2I) collaboration enables safer and broader perception. To process large-scale V2I video streams, vision-language models (VLMs) are promising as they unify multi-view vision into end-to-end task grounding, reducing handcrafted design. We use Vision Mixture-of-Experts (V-MoE) as the distributed visual backbone of VLMs, leveraging sparse expert routing to enable conditional computation across diverse viewpoints under resource constraints. Yet, V-MoEs face a critical challenge: large-scale data shifts over minutes to hours in V2I systems, amplified by agnostic participants and biased features propagating through experts. To maintain accuracy efficiently, we find it beneficial to co-adapt multiple V-MoEs on edge servers, avoiding the latency and privacy risks of cloud offloading and the accuracy sacrifices of on-device methods. However, the resource-constrained edge poses challenges for efficient co-adaptation: i) DRAM fragmentation and imbalance limit expert parallelism, ii) memory-I/O bottlenecks restrict computation reuse, and iii) asynchronous adaptation increases task-switch overhead. Also, prior work rarely explores the upper bound of concurrent tasks under limited edge resources, a critical factor for practical V2I deployment. To address these, we present AdaSprite. By combining cooperative elastic scaling with multi-level multiplexing, AdaSprite optimizes expert lifespans to reduce DRAM fragmentation, exploits predictable activation patterns for efficient I/O reuse, and employs twin-buffer scheduling to leverage sparsity. On a weak edge, AdaSprite supports up to 17 concurrent V2I tasks (vs. up to 6 for baselines), improving SLO attainment by 1.6x and throughput by 2.1x. Also, it allows users to trade accuracy and concurrency for second-level adaptation.
Lehao Wang, Zhiwen Yu 0001, Sicong Liu 0005, Fengmin Wu, Bin Guo 0001
MobiSys3
2026 EcoAIoT: A Survey on Ecological and Sustainable AIoT Systems With Energy Harvesting
abstract
The proliferation of Internet of Things (IoT) devices is enabling transformative applications across domains such as smart cities, healthcare, and environmental monitoring. However, the large-scale deployment of these devices also raises sustainability concerns. The integration of intelligence at the edge through the emergence of the Artificial Intelligence of Things (AIoT) amplifies computational requirements and energy consumption. This increase is especially pronounced in resource-constrained settings, consequently exacerbating their ecological impacts. To address these energy and environmental challenges, recent research has increasingly focused on energy harvesting (EH) techniques, energy-efficient sensing methods, and ultra-low-power computing architectures. This survey provides a comprehensive review of the current state of ecological and sustainable AIoT systems powered by energy harvesting. It systematically categorizes existing approaches based on energy sources, hardware-software co-optimization strategies, and communication protocols. In addition, it critically analyzes key challenges, particularly energy intermittency, system scalability, and the balance between performance and energy efficiency. This survey also highlights promising research directions, such as bio-inspired energy harvesting, neuromorphic computing, and integrated sensing-communication frameworks. By summarizing current advances and highlighting open research challenges, this survey aims to provide a structured roadmap for the development of next-generation sustainable AIoT systems.
Zhimin Jing, Yasan Ding, Bin Guo 0001, Sicong Liu 0005, Yao Jing, Geyang Song, Jingqi Liu, Zhiwen Yu 0001
IEEE Internet Things J.5
2026 Responsive Test-Time Model Adaptation for Mobile Applications via Runtime-Efficient Sparse Updates
abstract
Adapting the on-device deep learning (DL) models to continual and unpredictable domain shifts is crucial for mobile applications like autonomous driving and augmented reality, which require seamless user experiences in ever-changing environments. Test-time adaptation (TTA) offers a promising solution that fine-tunes models withunlabeledreal-time data just before making predictions. However, TTA introduces a significant challenge, its forward-backward-reforward pipeline increases latency, compromising responsiveness in time-sensitive mobile scenarios. In this paper, we present AdaShadow, a responsive test-time adaptation framework designed for non-stationary mobile data distributions and resource dynamics. AdaShadow focuses on selectively updating only adaptation-critical layers, minimizing the latency impact. While this idea has been explored in general on-device training, the unsupervised and real-time nature of TTA presents unique challenges in assessing layer importance, predicting latency, and planning layer updates efficiently. AdaShadow addresses these challenges through abackpropagation-free importance assessorfor rapid identification of critical layers, a unit-basedruntime latency predictorthat accounts for resource constraints, and anonline layer update schedulerto ensure prompt retraining. Furthermore, a memory I/O-aware computation reuse scheme optimizes the reforward pass to further reduce delays. Our evaluations show that AdaShadow reduces adaptation latency by up to 3.7× (achieving ms-level response time) and improves accuracy by up to 25.4%, all while maintaining low memory and energy costs across CNN, and ViT models, outperforming state-of-the-art methods.
Bin Guo 0001, Sicong Liu 0005, Zimu Zhou, Jiaqi Tang 0005, Shiyan Luo, Geyang Song, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.3
2026 AdaCLP: Efficient Federated Learning for Asynchronous Mobile Devices With Temporally Imbalanced Data
abstract
Federated learning (FL) enables mobile devices to collaboratively train deep learning models while maintaining data privacy and minimizing communication overhead. However, traditional FL methods typically assume access to pre-collected datasets, an impractical assumption for real-world mobile devices that continuously collect sensor data streams without retaining them while operating, due to limited memory constraints. Streaming Federated Learning (SFL) partially addresses this limitation by supporting online learning and asynchronous model aggregation directly from live data streams. Yet, a critical challenge in mobile data streams is thetemporal class imbalance, which may biase the streaming FL process toward early-arriving data classes during the critical learning period (CLP), the initial training phase when the model exhibits the highest plasticity. To address this, we propose${\sf AdaCLP}$, a training scheduler that tracks global model plasticity using a staleness-decayed mechanism specifically designed for the dynamics of real-world mobile environments.${\sf AdaCLP}$dynamically adjusts local training hyperparameters in SFL to prolong the CLP, thereby preserving model plasticity in the face of time-varying, imbalanced data distributions. Furthermore, a CLP-aware dynamic voltage and frequency scaling (DVFS) strategy is integrated to reduce the energy cost of prolonged training, aligning with the tight energy budgets of mobile devices. Experimental results demonstrate that${\sf AdaCLP}$can effectively handle either temporally imbalanced or periodic data, improving accuracy by up to 11% while reducing energy consumption by 69.5% compared to state-of-the-art methods, without modifying the standard FL training pipeline. These demonstrate that${\sf AdaCLP}$is a practical and efficient solution for real-world mobile FL deployment, enabling robust on-device adaptation to evolving data streams.
Sicong Liu 0005, Yuan Xu 0018, Zimu Zhou, Weiye Wu, Bin Guo 0001, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.1
2025 SURGEON: Memory-Adaptive Fully Test-Time Adaptation via Dynamic Activation Sparsity
abstract
Despite the growing integration of deep models into mobile terminals, the accuracy of these models declines significantly due to various deployment interferences. Test-time adaptation (TTA) has emerged to improve the performance of deep models by adapting them to unlabeled target data online. Yet, the significant memory cost, particularly in resource-constrained terminals, impedes the effective deployment of most backward-propagation-based TTA methods. To tackle memory constraints, we introduce Surgeon, a method that substantially reduces memory cost while preserving comparable accuracy improvements during fully test-time adaptation (FTTA) without relying on specific network architectures or modifications to the original training procedure. Specifically, we propose a novel dynamic activation sparsity strategy that directly prunes activations at layer-specific dynamic ratios during adaptation, allowing for flexible control of learning ability and memory cost in a data-sensitive manner. Among this, two metrics, Gradient Importance and Layer Activation Memory, are considered to determine the layer-wise pruning ratios, reflecting accuracy contribution and memory efficiency, respectively. Experimentally, our method surpasses the baselines by not only reducing memory usage but also achieving superior accuracy, delivering SOTA performance across diverse datasets, architectures, and tasks.
Jiaqi Tang 0005, Bin Guo 0001, Fan Dang 0001, Sicong Liu 0005, Zhui Zhu, Ying-Cong Chen, Zhiwen Yu 0001, Yunhao Liu 0001
CVPR5
2025 DeepSwarm: towards swarm deep learning with bi-directional optimization of data acquisition and processing
Sicong Liu 0005, Bin Guo 0001, Lehao Wang, Zimu Zhou, Zhiwen Yu 0001
Frontiers Comput. Sci.1
2025 CrowdMesh: A Dynamic Model Parallel Training System on Mobile Devices
abstract
With the rapid development of artificial intelligence (AI), integrating deep neural networks (DNNs) into mobile and embedded devices has become an important trend. This integration significantly enhances the ability of these devices to collect and analyze perceptual data. Traditionally, the integration paradigm relies on cloud based training and deployment on mobile devices. However, the dynamic characteristics and privacy issues related to real-world perceptual data require training on the device. Despite its advantages, the limited computing resources of mobile devices constitute a key bottleneck that hinders the efficiency of model training. To address this issue, parallel distributed training across mobile device clusters has become a feasible paradigm. However, the inherent mobility of these devices not only increases the possibility of training interruptions, but also exacerbates the inefficiency caused by data imbalance. These challenges make traditional cloud based model parallelization methods unsuitable for mobile environments. To overcome these limitations, this paper proposes a novel model parallel system CrowdMesh designed for mobile device clusters. CrowdMesh consists of three key modules: (1) alliance-game-based dynamic cluster startup and construction,(2) computation-cost-based DL model parallel, (3) device-mobility-aware parameter propagation module. These modules work together to address training interruptions and restarts caused by device mobility, as well as efficiency issues caused by data imbalance.The experimental results show that CrowdMesh performs better than existing cloud and device based model parallelization baselines in various training tasks and deep learning models, and can reduce training latency by more than 18%.
Bin Guo 0001, Sicong Liu 0005, Zhiwen Yu 0001, Daqing Zhang 0001
IEEE Internet Things J.3
2025 AdaScale: Dynamic Context-Aware DNN Scaling via Automated Adaptation Loop on Mobile Devices
abstract
Deep learning is reshaping mobile applications, with a growing trend of deploying deep neural networks (DNNs) directly to mobile and embedded devices to address real-time performance and privacy. To accommodate local resource limitations, techniques like weight compression, convolution decomposition, and specialized layer architectures have been developed. However, the dynamic and diverse deployment contexts of mobile devices pose significant challenges. Adapting deep models to meet varied device-specific requirements for latency, accuracy, memory, and energy is labor-intensive. Additionally, changing processor states, fluctuating memory availability, and competing processes frequently necessitate model recompression to preserve user experience. To address these issues, we introduce AdaScale, an elastic inference framework that automates the adaptation of deep models to dynamic contexts. AdaScale leverages a self-evolutionary model to streamline network creation, employs diverse compression operator combinations to reduce the search space and improve outcomes, and integrates a resource availability awareness block and performance profilers to establish an automated adaptation loop. Our experiments demonstrate that AdaScale significantly enhances accuracy by 5.09%, reduces training overhead by 66.89%, speeds up inference latency by 1.51 to$6.2\times $, and lowers energy costs by$4.69\times $.
Yuzhan Wang, Sicong Liu 0005, Bin Guo 0001, Boqi Zhang, Yasan Ding, Hao Luo 0022, Zhiwen Yu 0001
IEEE Internet Things J.2
2025 ClassTer: Mobile Shift-Robust Personalized Federated Learning via Class-Wise Clustering
abstract
The rise of mobile devices with abundant sensor data and computing power has driven the trend of federated learning (FL) on them. Personalized FL (PFL) aims to train tailored models for each device, addressing data heterogeneity from diverse user behaviors and preferences. However, due to dynamic mobile environments, PFL faces challenges intest-time data shifts, i.e., variations between training and testing. While this issue is well studied in generic deep learning through model generalization or adaptation, this issue remains less explored in PFL, where models often overfit local data. To address this, we introduce${\sf ClassTer}$, a shift-robust PFL framework. We observe that class-wise clustering of clients in cluster-based PFL (CFL) can avoid class-specific biases by decoupling the training of classes. Thus, we propose a paradigm shift from traditional client-wise clustering toclass-wise clustering, which allowseffective aggregationof cluster models into a generalized one via knowledge distillation. Additionally, we extend ClassTer toasynchronousmobile clients to optimize wall clock time by leveraging critical learning periods and both intra- and inter-device scheduling. Experiments show that compared to status quo approaches,${\sf ClassTer}$achieves a reduction of up to 91% in convergence time, and an improvement of up to 50.45% in accuracy.
Sicong Liu 0005, Zimu Zhou, Yuan Xu 0018, Bin Guo 0001, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.2
2025 CrowdHMTware: A Cross-Level Co-Adaptation Middleware for Context-Aware Mobile DL Deployment
abstract
There are many deep learning (DL) powered mobile and wearable applications today continuously and unobtrusively sensing the ambient surroundings to enhance all aspects of human lives. To enable robust and private mobile sensing, DL models are often deployed locally on resource-constrained mobile devices using techniques such as model compression or offloading. However, existing methods, either front-end algorithm level (i.e. DL model compression/partitioning) or back-end scheduling level (i.e. operator/resource scheduling), cannot be locally online because they require offline retraining to ensure accuracy or rely on manually pre-defined strategies, struggle withdynamic adaptability. The primary challenge lies in feeding back runtime performance from theback-endlevel to thefront-endlevel optimization decision. Moreover, the adaptive mobile DL model porting middleware withcross-level co-adaptationis less explored, particularly in mobile environments withdiversityanddynamics. In response, we introduce CrowdHMTware, a dynamic context-adaptive DL model deployment middleware for heterogeneous mobile devices. It establishes anautomated adaptation loopbetween cross-level functional components, i.e. elastic inference, scalable offloading, and model-adaptive engine, enhancing scalability and adaptability. Experiments with four typical tasks across 15 platforms and a real-world case study demonstrate that${\sf CrowdHMTware}$can effectively scale DL model, offloading, and engine actions across diverse platforms and tasks. It hides run-time system issues from developers, reducing the required developer expertise.
Sicong Liu 0005, Bin Guo 0001, Shiyan Luo, Yuzhan Wang, Hao Luo 0022, Yuan Xu 0018, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.1
2025 AdaKnife: Flexible DNN Offloading for Inference Acceleration on Heterogeneous Mobile Devices
abstract
The integration of deep neural network (DNN) intelligence into embedded mobile devices is expanding rapidly, supporting a wide range of applications. DNN compression techniques, which adapt models to resource-constrained mobile environments, often force a trade-off between efficiency and accuracy. Distributed DNN inference, leveraging multiple mobile devices, emerges as a promising alternative to enhance inference efficiency without compromising accuracy. However, effectively decoupling DNN models into fine-grained components for optimal parallel acceleration presents significant challenges. Current partitioning methods, including layer-level and operator or channel-level partitioning, provide only partial solutions and struggle with the heterogeneous nature of DNN compilation frameworks, complicating direct model offloading. In response, we introduce AdaKnife, an adaptive framework for accelerated inference across heterogeneous mobile devices. AdaKnife enables on-demand mixed-granularity DNN partitioning via computational graph analysis, facilitates efficient cross-framework model transitions with operator optimization for offloading, and improves the feasibility of parallel partitioning using a greedy operator parallelism algorithm. Our empirical studies show that AdaKnife achieves a 66.5% reduction in latency compared to baselines.
Sicong Liu 0005, Hao Luo 0022, Bin Guo 0001, Zhiwen Yu 0001, Yuzhan Wang, Yasan Ding, Yuan Yao 0004
IEEE Trans. Mob. Comput.1
2025 AdaFlowLite: Scalable and Non-Blocking Inference on Asynchronous Mobile Data
abstract
The rise of mobile devices equipped with numerous sensors, such as LiDAR and cameras, has driven the adoption of multi-modal deep intelligence for distributed sensing tasks, such as smart cabins and driving assistance. However, the arrival time of mobile sensory data vary due to modality size and network dynamics, which can lead to delays (if waiting for slow data) or accuracy decline (if inference proceeds without waiting). Moreover, the diversity and dynamic nature of mobile systems exacerbate this challenge. In response, we present a shift toopportunisticinference for asynchronous distributed multi-modal data, enabling inference as soon as partial data arrives. While existing methods focus on optimizing modality consistency and complementarity, known as modal affinity, they lack acomputationalapproach to control this affinity in open-world mobile environments.${\sf AdaFlowLite}$pioneers the formulation of structured cross-modality affinity in mobile contexts using a hierarchical analysis-based normalized matrix. This approach accommodates the diversity and dynamics of modalities, generalizing across different types and numbers of inputs. Employing an multi-modal lightweight Swin Transformer (MMLST),${\sf AdaFlowLite}$facilitates real-time and flexible data imputation, adapting to various modalities and downstream tasks without retraining. Experiments show that${\sf AdaFlowLite}$significantly reduces inference latency by up to 80.4% and enhances accuracy by up to 62.1%, while achieving nearly a 50% reduction in energy consumption, outperforming status quo approaches. Also, this method can enhance LLM performance to preprocess asynchronous data.
Sicong Liu 0005, Fengmin Wu, Bin Guo 0001, Zimu Zhou, Hongkai Wen 0001, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.1
2025 AdaShift: Anti-Collapse and Real-Time Deep Model Evolution for Mobile Vision Applications
abstract
As computational hardware advance, integrating deep learning (DL) models into mobile devices has become ubiquitous for visual tasks. However, “data distribution shift” in live sensory data can lead to a degradation in the accuracy of mobile DL models. Conventional domain adaptation methods, constrained by their dependence on pre-compiled static datasets for offline adaptation, exhibit fundamental limitations in real-time practicality. While modern online adaptation methodologies enable incremental model evolution, they remain plagued by two critical shortcomings: computational latency from excessive resource demands on mobile devices that compromise temporal responsiveness, and accuracy collapse stemming from error accumulation through unreliable pseudo-labeling processes. To address these challenges, we introduce AdaShift, an innovative cloud-assisted framework enabling real-time online model adaptation for vision-based mobile systems operating under non-stationary data distributions. Specifically, to ensure real-time performance, the adaptation trigger and plug-and-play adaptation mechanisms are proposed to minimize redundant adaptation requests and reduce per-request costs. To prevent accuracy collapse, AdaShift introduces a novel anti-collapse parameter restoration mechanism that explicitly recovers knowledge, ensuring stable accuracy improvements during model evolution. Through extensive experiments across various vision tasks and model architectures, AdaShift demonstrates superior accuracy and 100ms-level adaptation latency, achieving an optimal balance between accuracy and real-time performance compared to baselines.
Bin Guo 0001, Sicong Liu 0005, Zimu Zheng, Zhiwen Yu 0001
IEEE Trans. Mob. Comput.3
2025 AdaEvo: Edge-Assisted Continuous and Timely DNN Model Evolution for Mobile Devices
abstract
Mobile video applications today have attracted significant attention. Deep learning model (e.g., deep neural network, DNN) compression is widely used to enable on-device inference for facilitating robust and private mobile video applications. The compressed DNN, however, is vulnerable to the agnostic data drift of the live video captured from the dynamically changing mobile scenarios. To combat the data drift, mobile ends rely on edge servers to continuously evolve and re-compress the DNN with freshly collected data. We design a framework, AdaEvo, that efficiently supports the resource-limited edge server handling mobile DNN evolution tasks from multiple mobile ends. The key goal of AdaEvo is to maximize the average quality of experience (QoE), i.e., the proportion of high-quality DNN service time to the entire life cycle, for all mobile ends. Specifically, it estimates the DNN accuracy drops at the mobile end without labels and performs a dedicated video frame sampling strategy to control the size of retraining data. In addition, it balances the limited computing and memory resources on the edge server and the competition between asynchronous tasks initiated by different mobile users. With an extensive evaluation of real-world videos from mobile scenarios and across four diverse mobile tasks, experimental results show that AdaEvo enables up to 34% accuracy improvement and 32% average QoE improvement.
Lehao Wang, Zhiwen Yu 0001, Haoyi Yu, Sicong Liu 0005, Yaxiong Xie, Bin Guo 0001, Yunxin Liu 0001
IEEE Trans. Mob. Comput.4
2024 AdaShadow: Responsive Test-time Model Adaptation in Non-stationary Mobile Environments
abstract
On-device adapting to continual, unpredictable domain shifts is essential for mobile applications like autonomous driving and augmented reality to deliver seamless user experiences in evolving environments. Test-time adaptation (TTA) emerges as a promising solution by tuning model parameters with unlabeled live data immediately before prediction. However, TTA's unique forward-backward-reforward pipeline notably increases the latency over standard inference, undermining the responsiveness in time-sensitive mobile applications. This paper presents AdaShadow, a responsive test-time adaptation framework for non-stationary mobile data distribution and resource dynamics via selective updates of adaptation-critical layers. Although the tactic is recognized in generic on-device training, TTA's unsupervised and online context presents unique challenges in estimating layer importance and latency, as well as scheduling the optimal layer update plan. AdaShadow addresses these challenges with a backpropagation-free assessor to rapidly identify critical layers, a unit-based runtime predictor to account for resource dynamics in latency estimation, and an online scheduler for prompt layer update planning. Also, AdaShadow incorporates a memory I/O-aware computation reuse scheme to further reduce latency in the reforwardpass. Results show that AdaShadow achieves the best accuracy-latency balance under continual shifts. At low memory and energy costs, Adashadow provides a 2x to 3.5x speedup (ms-level) over state-of-the-art TTA methods with comparable accuracy and a 14.8% to 25.4% accuracy boost over efficient supervised methods with similar latency.
Sicong Liu 0005, Zimu Zhou, Bin Guo 0001, Jiaqi Tang 0005, Zhiwen Yu 0001
SenSys2
2024 AdaFlow: Opportunistic Inference on Asynchronous Mobile Data with Generalized Affinity Control
abstract
The rise of mobile devices equipped with numerous sensors, such as LiDAR and cameras, has spurred the adoption of multi-modal deep intelligence for distributed sensing tasks, such as smart cabins and driving assistance. However, the arrival times of mobile sensory data vary due to modality size and network dynamics, which can lead to delays (if waiting for slower data) or accuracy decline (if inference proceeds without waiting). Moreover, the diversity and dynamic nature of mobile systems exacerbate this challenge. In response, we present a shift to opportunistic inference for asynchronous distributed multi-modal data, enabling inference as soon as partial data arrives. While existing methods focus on optimizing modality consistency and complementarity, known as modal affinity, they lack a computational approach to control this affinity in open-world mobile environments. AdaFlow pioneers the formulation of structured cross-modality affinity in mobile contexts using a hierarchical analysis-based normalized matrix. This approach accommodates the diversity and dynamics of modalities, generalizing across different types and numbers of inputs. Employing an affinity attention-based conditional GAN (ACGAN), AdaFlow facilitates flexible data imputation, adapting to various modalities and downstream tasks without retraining. Experiments show that AdaFlow significantly reduces inference latency by up to 79.9% and enhances accuracy by up to 61.9%, outperforming status quo approaches. Also, this method can enhance LLM performance to preprocess asynchronous data.
Fengmin Wu, Sicong Liu 0005, Kehao Zhu, Bin Guo 0001, Zhiwen Yu 0001, Hongkai Wen 0001, Xiangrui Xu 0005, Lehao Wang
SenSys2
2024 CrowdLearning: A Decentralized Distributed Training Framework Based on Collectives of Trusted AIoT Devices
abstract
With the rise of Artificial Intelligence of Things (AIoT), integrating deep neural networks (DNNs) into mobile and embedded devices has become a significant trend, enhancing the data collection and analysis capabilities of IoT devices. Traditional integration paradigms rely on cloud-based training and terminal deployment, but they often suffer from delayed model updates, decreased accuracy, and increased communication overhead in dynamic real-world environments. Consequently, on-device training methods have garnered research focus. However, the limited local perception data and computational resources pose bottlenecks to training efficiency. To address these challenges, Federated Learning emerged but faces issues such as slow model convergence and reduced accuracy due to data privacy concerns that restrict sharing data or model details. In contrast, we propose the concept of trusted clusters in the real world (such as personal devices in smart spaces, trusted devices from the same organization/company, etc.), where devices in trusted clusters focus more on computational efficiency and can also share privacy. We propose CrowdLearning, a decentralized distributed training framework based on trusted AIoT device collectives. This framework comprises two collaborative modules: A heterogeneous resource-aware task offloading module aimed at alleviating training latency bottlenecks, and an efficient communication data reallocation module responsible for determining the timing, manner, and recipients of data transmission, thereby enhancing DNN training efficiency and effectiveness. Experimental results demonstrate that in various scenarios, CrowdLearning outperforms existing federated learning and distributed training baselines on devices, reducing training latency by 55.8% and lowering communication costs by 67.1%.
Sicong Liu 0005, Bin Guo 0001, Zhiwen Yu 0001, Daqing Zhang 0001
IEEE Trans. Mob. Comput.2
2024 AdaMEC: Towards a Context-adaptive and Dynamically Combinable DNN Deployment Framework for Mobile Edge Computing
abstract
With the rapid development of deep learning, recent research on intelligent and interactive mobile applications (e.g., health monitoring, speech recognition) has attracted extensive attention. And these applications necessitate the mobile edge computing scheme, i.e., offloading partial computation from mobile devices to edge devices for inference acceleration and transmission load reduction. The current practices have relied on collaborative DNN partition and offloading to satisfy the predefined latency requirements, which is intractable to adapt to the dynamic deployment context at runtime. AdaMEC, a context-adaptive and dynamically combinable DNN deployment framework, is proposed to meet these requirements for mobile edge computing, which consists of three novel techniques. First, once-for-all DNN pre-partition divides DNN at the primitive operator level and stores partitioned modules into executable files, defined as pre-partitioned DNN atoms. Second, context-adaptive DNN atom combination and offloading introduces a graph-based decision algorithm to quickly search the suitable combination of atoms and adaptively make the offloading plan under dynamic deployment contexts. Third, runtime latency predictor provides timely latency feedback for DNN deployment considering both DNN configurations and dynamic contexts. Extensive experiments demonstrate that AdaMEC outperforms state-of-the-art baselines in terms of latency reduction by up to 62.14% and average memory saving by 55.21%.
Sicong Liu 0005, Bin Guo 0001, Yuzhan Wang, Hao Wang 0182, Zhenli Sheng, Zhiwen Yu 0001
ACM Trans. Sens. Networks2
2023 A Novel Framework for Adaptive Quadruped Robot Locomotion Learning in Uncertain Environments
Bin Guo 0001, Kaixing Zhao, Ruonan Xu, Sicong Liu 0005, Sitong Mao, Shunbo Zhou, Qiaobo Xu, Zhiwen Yu 0001
GPC (2)5
2022 FedAux: An Efficient Framework for Hybrid Federated Learning
abstract
As an enabler of sixth-generation communication technology (6G), Federated Learning (FL) triggers a paradigm shift from "connected things" to "connected intelligence". FL implements on-device learning, where massive end devices jointly and locally train a model without private data leakage. However, FL suffers from problems of low accuracy and convergence rate when no data is shared to the central server and the data distribution is non-IID. In recent years, attempts have been made on hybrid FL, where very small amounts of data (e.g., less than 1%) is shared from the participants. With the opportunities brought by shared data, we notice that the server is capable of receiving the data in order to assist the FL process and mitigate the challenge of non-IID. Notably, existing hybrid FL only applies the model-level technologies belonging to the traditional FL and does not make full use of the characteristics of shared data to make targeted improvements. In this paper, we propose FedAux, a novel hybrid FL method at knowledge-level, which utilizes shared data to construct an auxiliary model and then transfer general knowledge to traditional aggregated model or client model for enhancing the accuracy of global model and speeding up the convergence of global model. We also propose two specific knowledge transfer strategies named c-transfer and i-transfer. We conduct extensive analysis and evaluation of our methods against the well-known FL methods, FedAvg and Hybrid-FL protocol. The results indicate that FedAux shows higher accuracy (10.89%) and faster convergence rate compared with other methods.
Hang Gu, Bin Guo 0001, Jiangtao Wang 0001, Wen Sun 0004, Jiaqi Liu 0002, Sicong Liu 0005, Zhiwen Yu 0001
ICC6
2022 Context-Adaptive Online Reinforcement Learning for Multi-view Video Summarization on Mobile Devices
abstract
The huge amount of video data produced by ubiqui tous cameras imposes significant challenges for users to efficiently obtain useful video information. Multi-view video summarization (MVS) aggregates multi-view videos into information-rich video summaries by considering content correlations within each view and between multiple views. Existing MVS methods fail to concentrate on performance across scenarios and usually achieve satisfactory performance on specific training datasets. However, when faced with unseen video scenarios, the quality of the summaries generated by existing methods may degrade. Moreover, they usually only use cameras for data acquisition, which require a large amount of network bandwidth to transfer the data to the server for processing. To bridge this gap, we propose a context-adaptive online reinforcement learning multi-view video summarization framework (COORS) that meets the low response latency performance requirements of context adaptation while ensuring camera hardware compatibility. Specifically, COORS enables retraining in new contexts by extracting contextindependent rewards, while improving model convergence speed based on representation learning and replica playback. Extensive experiments show that COORS has better performance compared to the state-of-the-art baselines.
Jingyi Hao, Sicong Liu 0005, Bin Guo 0001, Yasan Ding, Zhiwen Yu 0001
ICPADS2
2022 CrowdDesigner: information-rich and personalized product description generation
Qiuyun Zhang, Bin Guo 0001, Sicong Liu 0005, Jiaqi Liu 0002, Zhiwen Yu 0001
Frontiers Comput. Sci.3
2022 CrowdHMT: Crowd Intelligence With the Deep Fusion of Human, Machine, and IoT
abstract
Mobile crowd sensing and computing (MCSC) has become a hot research area in recent years. This article presents our vision of the next generation of MCSC, crowd intelligence with the deep fusion of human, machine, and Internet of Things (IoT), namely, CrowdHMT. It aims to build a self-organizing, self-learning, self-adaptive, and continuous-evolving smart space with the deep fusion of Crowdsourced human, machine, and IoT intelligence. This article first characterizes the concept of CrowdHMT. We further investigate its challenges and techniques, and present its main application areas. Finally, we make discussions about the open issues and future research directions of CrowdHMT.
Bin Guo 0001, Yan Liu 0045, Sicong Liu 0005, Zhiwen Yu 0001, Xingshe Zhou 0001
IEEE Internet Things J.3
2022 CrowdIM: Crowd-Inspired Intelligent Manufacturing Space Design
abstract
Crowd-inspired intelligent manufacturing space (CrowdIM) aims to leverage the aggregated power of heterogeneous human–machine–things (HMT) agents for improving the efficiency of intelligent manufacturing. A significant scientific problem in CrowdIM is how to improve individual skills and crowd intelligence through cooperation, complementation, competition, and confrontation among HMT agents. The emergence mechanism of biological crowd intelligence provides an inspiration to address this challenge. This article explores the mapping mechanisms between natural crowd intelligence and CrowdIM, from the aspects, such as collective dynamics, self-adaptive mechanism, crowd intelligence optimization, graph structure mapping model, evolutionary game dynamics, multiagent learning, and so on. We further propose a general model of CrowdIM and expound it through a typical case study.
Bin Guo 0001, Jiaqi Liu 0002, Sicong Liu 0005, Chen Wang 0018, Zhiwen Yu 0001
IEEE Internet Things J.3
2022 CAQ: Toward Context-Aware and Self-Adaptive Deep Model Computation for AIoT Applications
abstract
Artificial Intelligence of Things (AIoT) has recently accepted significant interests. Remarkably, embedded artificial intelligence (e.g., deep learning) on-device transforms IoT devices into intelligent systems that robustly and privately process data. Quantization technique is widely used to compress deep models for narrowing the resource gap between computation demands and platform supply. However, existing quantization schemes induce unsatisfaction for IoT scenarios since they are oblivious to dynamic changes of application context (e.g., battery and hierarchical memory availability) during the long-term operation. Subsequently, they will mismatch the user-desired resource efficiency and application lifetime. Also, to adapt to the dynamic context, we can neither accept the latency for model retraining with existing hand-crafted quantization nor the overhead for quantization bit width researching with prior on-demand quantization. This article presents a context-aware and self-adaptive deep model quantization (CAQ) system for IoT application scenarios. CAQ integrates a novel switchable multigate quantization framework, optimizing the quantized model accuracy and energy efficiency in diverse contexts. Based on the learned model, CAQ can switch among different gating networks in a context-aware manner and then adopt it to automatically capture the representation importance of various layers for optimal quantization bit-width selection. The experimental results show that CAQ achieves up to 50% storage savings with even 2.61% higher accuracy than the state-of-the-art baselines.
Sicong Liu 0005, Yungang Wu, Bin Guo 0001, Yuzhan Wang, Liyao Xiang, Zhetao Li, Zhiwen Yu 0001
IEEE Internet Things J.1
2022 Investigation of the determinants for misinformation correction effectiveness on social media during COVID-19 pandemic
Bin Guo 0001, Yasan Ding, Jiaqi Liu 0002, Chen Qiu 0002, Sicong Liu 0005, Zhiwen Yu 0001
Inf. Process. Manag.6
2021 JointCS: Joint Search for Deep Model Compression and Segmentation on Heterogeneous IoT Devices
abstract
Deep neural networks (DNNs) play an important role in a variety of intelligent applications (e.g. image classification and target recognition), yet at the cost of heavy computation burden, that makes DNNs difficult to deploy on resource-constrained IoT devices. To solve this problem, there are two categories of model computation adjustment methods: model compression and model segmentation. However, model compression mainly reduces resource consumption at the cost of accuracy while model segmentation reduces resource consumption according to the cost of communication latency. In this paper, we propose Joint Search for Model Compression and Segmentation (JointCS) that highlights the following aspects: 1) we integrate both model compression and model segmentation under an automatic and progressive framework, it simplifies model to fit the different IoT resource requirements. JointCS achieves a series slim models that outperform better both in accuracy and latency. 2) we train a network architecture-aware latency predictor to fast measure the latency of the slimed model on heterogeneous IoT devices. 3) we introduce a search algorithm to select the optimal state in progressively joint search. Finally, we evaluate the performance of our proposed method for image classification on CIFAR datasets comparing with the state-of-the-art approach, the inference time of the proposed method has inference speedup of 12.2 % −30.9 % under the same accuracy.
Bin Guo 0001, Sicong Liu 0005, Chen Qiu 0002, Yunji Liang, Zhiwen Yu 0001
ICPADS3
2021 Decentralized Multi-AGV Task Allocation based on Multi-Agent Reinforcement Learning with Information Potential Field Rewards
abstract
Automated Guided Vehicles (AGVs) have been widely used for material handling in flexible shop floors. Each product requires various raw materials to complete the assembly in production process. AGVs are used to realize the automatic handling of raw materials in different locations. Efficient AGVs task allocation strategy can reduce transportation costs and improve distribution efficiency. However, the traditional centralized approaches make high demands on the control center’s computing power and real-time capability. In this paper, we present decentralized solutions to achieve flexible and self-organized AGVs task allocation. In particular, we propose two improved multi-agent reinforcement learning algorithms, MAD-DPG-IPF (Information Potential Field) and BiCNet-IPF, to realize the coordination among AGVs adapting to different scenarios. To address the reward-sparsity issue, we propose a reward shaping strategy based on information potential field, which provides stepwise rewards and implicitly guides the AGVs to different material targets. We conduct experiments under different settings (3 AGVs and 6 AGVs), and the experiment results indicate that, compared with baseline methods, our work obtains up to 47% task response improvement and 22% training iterations reduction.
Bin Guo 0001, Jiangshan Zhang, Jiaqi Liu 0002, Sicong Liu 0005, Zhiwen Yu 0001, Zhetao Li, Liyao Xiang
MASS5
2021 Towards information-rich, logical dialogue systems with knowledge-enhanced neural models
Hao Wang 0182, Bin Guo 0001, Wei Wu 0014, Sicong Liu 0005, Zhiwen Yu 0001
Neurocomputing4
2021 AdaDeep: A Usage-Driven, Automated Deep Model Compression Framework for Enabling Ubiquitous Intelligent Mobiles
abstract
Recent breakthroughs in deep neural networks (DNNs) have fueled a tremendously growing demand for bringing DNN-powered intelligence into mobile platforms. While the potential of deploying DNNs on resource-constrained platforms has been demonstrated by DNN compression techniques, the current practice suffers from two limitations: 1) merely stand-alone compression schemes are investigated even though each compression technique only suit for certain types of DNN layers; and 2) mostly compression techniques are optimized for DNNs’ inference accuracy, without explicitly considering other application-driven system performance (e.g., latency and energy cost) and the varying resource availability across platforms (e.g., storage and processing capability). To this end, we propose AdaDeep, a usage-driven, automated DNN compression framework for systematically exploring the desired trade-off between performance and resource constraints, from a holistic system level. Specifically, in a layer-wise manner, AdaDeep automatically selects the most suitable combination of compression techniques and the corresponding compression hyperparameters for a given DNN. Thorough evaluations on six datasets and across twelve devices demonstrate that${\sf AdaDeep}$can achieve up to$18.6\times$latency reduction,$9.8\times$energy-efficiency improvement, and$37.3\times$storage reduction in DNNs while incurring negligible accuracy loss. Furthermore,${\sf AdaDeep}$also uncovers multiple novel combinations of compression techniques.
Sicong Liu 0005, Junzhao Du, Kaiming Nan, Zimu Zhou, Hui Liu 0006, Zhangyang Wang, Yingyan (Celine) Lin
IEEE Trans. Mob. Comput.1
2018 On-Demand Deep Model Compression for Mobile Devices: A Usage-Driven Model Selection Framework
abstract
Recent research has demonstrated the potential of deploying deep neural networks (DNNs) on resource-constrained mobile platforms by trimming down the network complexity using different compression techniques. The current practice only investigate stand-alone compression schemes even though each compression technique may be well suited only for certain types of DNN layers. Also, these compression techniques are optimized merely for the inference accuracy of DNNs, without explicitly considering other application-driven system performance (e.g. latency and energy cost) and the varying resource availabilities across platforms (e.g. storage and processing capability). In this paper, we explore the desirable tradeoff between performance and resource constraints by user-specified needs, from a holistic system-level viewpoint. Specifically, we develop a usage-driven selection framework, referred to as AdaDeep, to automatically select a combination of compression techniques for a given DNN, that will lead to an optimal balance between user-specified performance goals and resource constraints. With an extensive evaluation on five public datasets and across twelve mobile devices, experimental results show that AdaDeep enables up to 9.8x latency reduction, 4.3x energy efficiency improvement, and 38x storage reduction in DNNs while incurring negligible accuracy loss. AdaDeep also uncovers multiple effective combinations of compression techniques unexplored in existing literature.
Sicong Liu 0005, Yingyan (Celine) Lin, Zimu Zhou, Kaiming Nan, Hui Liu 0006, Junzhao Du
MobiSys1
2016 CrowdBlueNet: Maximizing Crowd Data Collection Using Bluetooth Ad Hoc Networks
Sicong Liu 0005, Junzhao Du, Rui Li 0047, Hui Liu 0006, Kewei Sha
WASA1
2014 Lightweight construction of the information potential field in wireless sensor networks
abstract
The information gradient-based routing protocols have been proved to be economical and effective by adopting the principle of achieving the global objective through local decision, but lightweight methods to construct the information gradient should be fully investigated, especially in a large-scale network with high information dynamics. In this paper, we focus on the construction of the information gradient by balancing convergence conditions and energy consumption. Therefore, two algorithms, Hierarchical Skeleton-based Construction Algorithm (HSCA) and Estimate value Substitution Algorithm (ESA) are proposed to achieve the goal of fastening the convergence in an energy efficient way. Both of the algorithms obey the typical assumptions on WSNs settings and the gossip-styled propagation principle. Comprehensive simulation results show that the proposed algorithms can reduce iteration times to reach a convergence status by 80% and conserve 30–50% energy consumption on average.
Junzhao Du, Sicong Liu 0005, Hui Liu 0006, Kewei Sha
ICCCN2