Jiacheng Liu 0001

dblp:202/3902-1 · DBLP profile ↗
← Back
45ranked-venue papers
11as first author
41since 2021 · last 2026
0000-0003-0378-2311ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 1 first-author · 13 since 2021Computer networks · 12 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache Compression
Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001
AAAI6
2026 AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning Models
abstract
Large reasoning models (LRMs) have demonstrated remarkable capabilities in solving complex problems through extended chain-of-thought reasoning. However, existing approaches face a fundamental trade-off between computational efficiency and reasoning accuracy. Current methods either lack support for user-specified computational budgets or require maintaining multiple independent models, leading to significant resource overhead. In this paper, we present AdaReason, a unified framework that trains a single base model to support arbitrary user-defined computational budgets through dynamic adapter composition. Our approach introduces three key innovations: (1) a length-adaptive step reward function that stabilizes training across diverse budget constraints, (2) a progressive training strategy that gradually tightens computational bounds while maintaining model performance, and (3) a runtime adapter merging mechanism that dynamically interpolates between different computational preferences. Unlike existing methods that suffer from training instability in large context windows, AdaReason achieves stable convergence through careful reward shaping and progressive constraint tightening. Additionally, we provide a rigorous theoretical analysis, establishing a performance bound for our merged model. Experiments on different reasoning benchmarks demonstrate that AdaReason establishes a new state-of-the-art in the performance-efficiency trade-off and enables flexible runtime budget adaptation.
Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001
AAAI5
2026 Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language Models
abstract
Large Language Models (LLMs) have achieved exceptional performance in complex reasoning via Chain-of-Thought (CoT), yet the associated computational costs remain prohibitive. CoT reasoning contains significant untapped efficiency potential across two dimensions: temporal redundancy, where reasoning steps may be unnecessary, and spatial redundancy, where computations can be performed at reduced precision. While current optimization techniques often necessitate resource-intensive fine-tuning or data curation, we introduce ASTRO (Adaptive Spatial and Temporal Redundancy Optimization), a training-free framework that simultaneously addresses both dimensions. ASTRO leverages Dewey’s reflective thinking model to segment reasoning phases, applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. Empirical results across diverse reasoning benchmarks demonstrate that ASTRO achieves up to an 11.3 \times efficiency gain without compromising accuracy, highlighting the advantages of holistic multi-dimensional redundancy management over isolated optimization methods.
Pengyu Cheng, Qiyuan Zhu, Hao Gu 0001, Ruijie Shen, Xiaofeng Hou, Sirui Han, Jiacheng Liu 0001
ACL (1)10
2026 BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
abstract
Hao Gu, Lujun Li, Hao Wang, Lei Wang, Zheyu Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Sirui Han, Yike Guo
ACL (1)7
2026 Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
abstract
Binxing Xu, Hao Gu, Lujun Li, Hao Wang, Bei Liu, Jiacheng Liu, Qiyuan Zhu, Xintong Yang, Chao Li, Sirui Han, Yike Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Binxing Xu, Hao Gu 0001, Lujun Li 0001, Hao Wang 0097, Jiacheng Liu 0001, Qiyuan Zhu, Xintong Yang, Chao Li 0009, Sirui Han, Yike Guo
ACL (1)6
2026 TierServe: Revenue-Maximizing LLM Inference Scheduling Across Multiple Subscription Tiers
Jiacheng Liu 0001, Xiaofeng Hou, Minyi Guo
APPT2
2026 OrbitGuard: Hierarchical Orbit-Aware Runtime for Spaceborne LLM Inference
Xiaofeng Hou, Jiacheng Liu 0001, Xiaozhi Zhu, Chao Li 0009
APPT3
2026 MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert Offloading
abstract
Mixture-of-experts (MoE) architectures enable scalable Large Language Models (LLMs) with reduced computational overhead, yet their deployment on memory-constrained edge devices is hindered by substantial memory demands. Traditional expert-offloading techniques mitigate memory constraints but often significantly increase inference latency. We introduce MoE-APEX, an Adaptive Precision EXpert offloading system that optimizes MoE inference for edge architectures by dynamically managing expert precision. Our core innovation is to replace less critical cache-miss experts with low-precision variants, reducing loading latency while maintaining accuracy. MoE-APEX introduces three innovative techniques that map the natural hierarchy of MoE computation: (1) a token-level dynamic expert loading mechanism, (2) a layer-level adaptive expert prefetching technique, and (3) a sequence-level cost-aware expert caching policy. These innovations enable MoE-APEX to leverage the benefits of mixed-precision expert inference fully. Implemented atop Llama.cpp, MoE-APEX achieves decoding speedups ranging from 1.34x to 9.75x compared to state-of-the-art MoE offloading systems across diverse edge devices, offering a robust solution for efficient MoE deployment in resource-constrained environments.
Jiacheng Liu 0001, Xiaofeng Hou, Yi-Fei Pu, Jing Wang 0055, Pheng-Ann Heng, Chao Li 0009, Minyi Guo
ASPLOS (2)2
2026 LocMore: Locating More Bursty Latency-critical Jobs on Resource-constrained Nodes
Xiaofeng Hou, Xinkai Wang 0003, Jiacheng Liu 0001, Chao Li 0009, Minyi Guo
IWQoS5
2026 StarkServe: A Framework for Elastic Serverless LLM Inference at the Extreme Edge
Xiaofeng Hou, Jiacheng Liu 0001, Xinkai Wang 0003, Chao Li 0009, Minyi Guo
IWQoS3
2026 Noisy Multi-Label Aggregation With Self-Supervised Graph Transformer in Mobile Crowdsourcing
abstract
Aggregating noisy labels from mobile crowdsourcing (MCS) to recover true labels is a fundamental yet challenging problem, especially due to the sparsity and unreliability of crowd-contributed data. While most prior work addresses only single-label scenarios, real-world MCS applications often require robust solutions for both single-label and multi-label tasks, where each instance may be associated with multiple categories. In this paper, we propose ATHENA, a novel approach that leverages self-supervision signals inherent in MCS data for effective label aggregation. Firstly, we propose a graph transformer model that can learn from the MCS topology and features. Then, we propose self-supervision signals inherently included in the dataset to help aggregate the labels. To address the unique challenges of multi-label aggregation, we further extend our approach toATHENA+, introducing a label message passing (LMP) module that explicitly models correlations and dependencies among labels. We conducted extensive experiments on multiple single-label and multi-label classification datasets, comparing the proposed models with state-of-the-art methods. Our results demonstrate that ATHENA and ATHENA+ are highly effective in aggregating labels and obtain much better performance than existing methods.
Jiacheng Liu 0001, Feilong Tang 0001, Hao Liu 0085, Long Chen 0025, Yanmin Zhu 0006, Jiadi Yu, Yichuan Yu, Xiaofeng Hou
IEEE Trans. Mob. Comput.1
2025 Outlier-Aware Model Merging for Efficient Multitask Inference
abstract
Model merging techniques aim to consolidate multiple fine-tuned models into a single unified model, reducing both storage and computational overhead while retaining task-specific performance. However, existing methods face several limitations: monotonous compression techniques that fail to account for task-specific weight distribution characteristics, weight-magnitude-based compression that fails to consider functional importance revealed by activation patterns, and non-adaptive allocation strategies that ignores task-specific layer importance. To overcome these challenges, we propose OA-Merge, a novel Outlier-Aware Model Merging framework that leverages task activation outliers to enable adaptive compression and resource allocation across tasks. OA-Merge comprises three key components: (1) dynamic hybrid decomposition technique that formulates task vectors as tailored combinations of low-rank and sparse components adapted to task-specific statistical distributions, (2) activation-informed compression methodology that incorporates task-specific activation statistics to prioritize functionally important weights, and (3) task-related allocation that optimizes the distribution of compression resources according to layer-specific importance metrics derived from activation outlier analysis. These hybrid outlier-aware strategies adapt dynamically to each task's intrinsic characteristics, avoiding the pitfalls of one-size-fits-all ways. Extensive experiments on both vision models (e.g., ViT) and language models (e.g., RoBERTa, Qwen) demonstrate that OA-Merge outperforms state-of-the-art baselines, achieving average performance gains of 3.2% on vision tasks and 2.8% on language tasks.
Qiyuan Zhu, Lujun Li 0001, Jiacheng Liu 0001, Pengyu Cheng, Sirui Han, Yike Guo
ACM Multimedia4
2025 SpaceExit: Enabling Efficient Adaptive Computing in Space with Early Exits
Jiacheng Liu 0001, Xiaozhi Zhu, Tongqiao Xu, Xiaofeng Hou, Chao Li 0009
USENIX ATC1
2025 MMBypass: Towards efficient multi-modal AI computing with adaptive bypass network
Yi-Fei Pu, Xinfeng Xia, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Jingling Yuan, Chao Li 0009
J. Parallel Distributed Comput.6
2025 Improving Efficiency in Multi-Modal Autonomous Embedded Systems Through Adaptive Gating
abstract
The parallel advancement of AI and IoT technologies has recently boosted the development of multi-modal computing ($M^{2}C$) on pervasive autonomous embedded systems (AES).$M^{2}C$takes advantage of data from different modalities such as images, audio, and text and is able to achieve notable improvements in accuracy. However, achieving these accuracy gains often comes at the cost of increased computational complexity and energy consumption. Furthermore, the presence of numerous advanced sensors in these systems significantly contributes to power consumption, exacerbating the issue of limited power resources. Collectively, these challenges pose difficulties in deploying$M^{2}C$on small embedded devices with scarce energy resources. In this article, we propose anAdaptiveModalityGating technique calledAMGfor in-situ$M^{2}C$applications. The primary objective ofAMGis to conserve energy while preserving the accuracy advantages of$M^{2}C$. To achieve this goal,AMGincorporates two first-of-its-kind designs. Firstly, it introduces a novel semi-gating architecture that enables partial modality sensor power gating. Specifically, we devise the de-centralizedAMG(D-AMG) and centralizedAMG(C-AMG) architecture. The former buffers raw data on sensors while the latter buffers raw data on the computing board, which are suitable for different edge scenarios respectively. Secondly, it facilitates a self-initialization/tuning process on the AES, which is supported by carefully-built analytical model. Extensive evaluations demonstrate the effectiveness ofAMG. It achieves a 1.6x to 3.8x throughput higher than other power management methods and improves the lifespan of AES by 10% to 280% longer within the same energy budget, while satisfying all performance and latency requirements across various scenarios.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xuehan Tang, Kwang-Ting Cheng, Minyi Guo
IEEE Trans. Computers4
2025 BAT: A Versatile Bipartite Attention-Based Approach for Comprehensive Truth Inference in Mobile Crowdsourcing
abstract
The proliferation of smart mobile devices has catalyzed the growth of Mobile CrowdSourcing (MCS) as a distributed problem-solving paradigm. MCS platforms heavily rely on advanced truth inference techniques to extract reliable information from diverse and potentially noisy crowd-contributed data. Existing truth inference models often made simplified assumptions about workers or tasks, employing complex Bayesian models or stringent data aggregation methods. These approaches tend to be task-specific, primarily limited to categorical labeling, making adaptations to other mobile computing scenarios labor-intensive. To address these limitations, we introduce the Bipartite Attention-driven Truth (BAT), a versatile approach tailored for mobile computing environments. BAT utilizes an Attributed Bipartite Graph (ABG) to holistically model the MCS process, with workers and tasks as nodes connected by edges representing answer-specific attributes. The approach employs a bipartite graph neural network with an innovative attention mechanism to assess the importance of different answers. BAT extends beyond categorical tasks to support numerical ones by incorporating novel feature representations and model extensions. Theoretical analyses clarify the link between answer similarity and worker expertise. Extensive experiments using diverse real-world datasets demonstrate BAT's superior performance compared to state-of-the-art categorical and numerical truth inference models, highlighting its effectiveness in mobile computing scenarios.
Jiacheng Liu 0001, Feilong Tang 0001, Hao Liu 0085, Long Chen 0025, Yichuan Yu, Yanmin Zhu 0006, Jiadi Yu, Xiaofeng Hou, Pheng-Ann Heng
IEEE Trans. Mob. Comput.1
2025 An Adaptive and Interpretable Congestion Control Service Based on Multi-Objective Reinforcement Learning
abstract
The need for an adaptive congestion control (CC) service is crucial due to the heterogeneity of systems and the diversity of applications. Traditional CC methods often fail to adaptively balance throughput and delay, struggling to meet the varied demands of different network applications. In this work, we introduceAuto, a novel CC service that employs Multi-Objective Reinforcement Learning (MORL) to transcend these limitations. Unlike conventional approaches,Autooptimizes policies within a single model to cater to all potential preferences for balancing throughput and delay, making it ideal for diverse and heterogeneous network environments. To enhance operational transparency, we developed an interpretation algorithm that translates MORL into a human- readable decision tree, essential for service computing where clarity and interpretability are crucial. Furthermore,Autoallows users to explicitly set flow priorities and target sending rates, meeting varied application demands. Our extensive evaluations show thatAutonot only consistently outperforms existing CC methods in diverse network conditions but also exhibits robustness to stochastic packet loss and rapid network changes. These capabilities establishAutoas a pioneering solution for next-generation congestion control in networking services.
Jiacheng Liu 0001, Xu Li 0012, Feilong Tang 0001, Peng Li 0017, Long Chen 0025, Jiadi Yu, Yanmin Zhu 0006, Pheng-Ann Heng, Laurence T. Yang
IEEE Trans. Serv. Comput.1
2024 Graph Contrastive Learning for Truth Inference
abstract
Crowdsourcing has become a popular paradigm for collecting large-scale labeled datasets by leveraging numerous annotators. However, these annotators often provide noisy labels due to varying expertise. Truth inference aims to infer accurate consensus labels from noisy crowdsourced annotations. Existing approaches rely heavily on hand-engineered assumptions or ground truth data, limiting their applicability. To address this, we propose GOVERN, a graph contrastive learning framework for truth inference without such external supervision. GOVERN employs a novel graph data augmentation strategy to generate views capturing worker coordination patterns. A contrastive objective then encourages invariant representations across views, enabling the discovery of features related to the hidden consensus. Further, a label correction method based on k-nearest neighbors refines noisy pseudo-labels to supervise model training. Comprehensive experiments on 9 real-world datasets demonstrate that GOVERN outperforms state-of-the-art truth inference techniques.
Hao Liu 0085, Jiacheng Liu 0001, Feilong Tang 0001, Peng Li 0017, Long Chen 0025, Jiadi Yu, Yanmin Zhu 0006, Yanqin Yang, Xiaofeng Hou
ICDE2
2024 M2SN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications
abstract
Multi-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines.
Yi-Fei Pu, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Chao Li 0009
ICME5
2024 A Tale of Two Domains: Exploring Efficient Architecture Design for Truly Autonomous Things
abstract
Autonomous Things (AuT) refers to a collection of self-sufficient tiny devices capable of performing intelligent computations. Looking ahead, AuT promises to enable ubiquitous deployment of intelligence on many emerging consumer electronics and mission-critical infrastructures. Nevertheless, there is an important research gap to date: architecting efficient AuT systems requires both energy autonomy (EA) and inference autonomy (IA). In other words, practical AuT application scenarios necessitate tailored architectures with significantly expanded inference performance and more efficient use of energy.We present CHRYSALIS, a novel automated EA/IA co-design methodology for autonomous things. It aims to guide the transition from a traditional EA-only and IA-only design approach to a truly AuT-oriented architecture design. To fully understand the interrelationship between the EA domain and the IA domain, CHRYSALIS first introduces an architectural modeling framework encompassing every key AuT module involving energy harvesting, intermittent execution, and accelerator control. Based on the holistic system model, we design an intelligent architecture generation tool that can help find the ideal design for targeted AuT scenarios adhering to different SWaP (Size, Weight and Power) constraints. To validate our work, we use CHRYSALIS for fast construction and exploration of efficient AuT design and pre-RTL design in representative AuT scenarios. Extensive evaluation shows that CHRYSALIS outperforms state-of-the-art designs and our proposed technique shows 56.4% better performance on average. We believe that the methodology and tools developed in this paper will foster the development of more performant and practical architectures in the upcoming AuT era.
Xiaofeng Hou, Tongqiao Xu, Chao Li 0009, Jiacheng Liu 0001, Yang Hu 0001, Jieru Zhao, Jingwen Leng, Kwang-Ting Cheng, Minyi Guo
ISCA5
2024 MobiShare: Efficient Decentralized Data Sharing for Mobile Devices
abstract
Existing peer-to-peer data-sharing methods suffer from low data delivery efficiency and scalability due to the naive data request/response procedure and the high redundant data transmission rate. It becomes even worse in large-scale mobile networks considering the limited resources of mobile devices. To address this issue, this paper presents MobiShare, an efficient decentralized data-sharing approach for mobile devices, which allows users to not only share the data but also the data generation methods. To achieve MobiShare, we introduce a function block encoding method and a data request method to enhance sharing efficiency, minimizing costs for decentralized data sharing. We propose a credit payment mechanism where congested devices can send data vouchers instead of actual data, containing the expected transmission time. Based on the load and bandwidth of devices, we build the optimized dissemination tree with data vouchers in a decentralized way to improve scalability. Evaluation results show that MobiShare avoids redundant transmission. It greatly shortens transmission completion time and lowers energy consumption with limited network resources.
Long Chen 0025, Feilong Tang 0001, Xu Li 0012, Jiacheng Liu 0001, Yichuan Yu, Yanqin Yang, Wenchao Xu 0002, Hengzhi Wang
IWQoS4
2024 Practical Network Modeling Using Weak Supervision Signals for Human-Centric Networking in Metaverse
abstract
As the metaverse continues to expand, it becomes increasingly critical to have human-centric networks that are both efficient and high-performing to optimize the user experience. Network modeling plays a fundamental role in optimizing and allocating resources efficiently, and configuring networks to satisfy the demands of diverse applications and users. Recently, traditional queuing theory-based approaches to network modeling have given way to machine learning-based methods. These methods rely on vast amounts of data for building precise models. Although high-precision simulators are ubiquitous, data collection is still an expensive and time-consuming process, resulting in a data bottleneck. In this paper, we propose a weakly supervised learning approach to modeling networks for human-centric networking in the metaverse. Specifically, we identify that queuing theory-based labels can be used to design the supervision signal at a very low cost. Therefore, we propose an approach that combines the inaccurate network modeling obtained from queuing theory-based approaches with an efficient and precise network model through only a small amount of simulation data. To make it a reality, we propose a novel neural network model that combines the powerful graph neural network and transformers. Additionally, we propose several additional supervision signals and a training algorithm to build a better network model. Experimental results demonstrate that our approach reduces the burden of data collection while achieving prediction accuracy comparable to results from large amounts of expensive simulation data. Furthermore, our approach exhibits superior generalization ability.
Jiacheng Liu 0001, Feilong Tang 0001, Zhijian Zheng, Hao Liu 0085, Xiaofeng Hou, Long Chen 0025, Ming Gao 0001, Jiadi Yu, Yanmin Zhu 0006
IEEE J. Sel. Areas Commun.1
2024 A2: Towards Accelerator Level Parallelism for Autonomous Micromobility Systems
abstract
Autonomous micromobility systems (AMS) such as low-speed minicabs and robots are thriving. In AMS, multiple Deep Neural Networks execute in parallel on heterogeneous AI accelerators. An emerging paradigm called Accelerator Level Parallelism (ALP) suggests managing accelerators holistically. However, there lacks a specialized and practical solution populating ALP for an AMS, where the varying real-time requirements under different working scenarios bring an opportunity to dynamically tradeoff between latency and efficiency. Furthermore, accelerator heterogeneity introduces enormous configuration space, and the shared-memory architecture results in dynamic bandwidth interference. In this article, we propose A 2 , a novel AMS resource manager optimizing energy and memory space efficiency under variable latency constraints. We gain insight from prior Learn&Control scheme to design an Analyze&Adapt scheme specialized for heterogeneous AI accelerators under shared-memory architecture. It features analyzing the system thoroughly offline to support two-step adaptation online. We build a prototype of A 2 and evaluate it on a commercial edge platform. We show that A 2 achieves 32.8% improvements in power and 13.8% in memory compared with control-based methods. As for timeliness enhancement, A 2 reduces the deadline violation rate by 9.2 percentage points (12.8% → 3.6%) on average compared to directly porting Learn&Control methods.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xinkai Wang 0003, Quan Chen 0002, Minyi Guo
ACM Trans. Archit. Code Optim.4
2024 Time-Varying Resource Graph Based Processing on the Way for Space-Terrestrial Integrated Vehicle Networks
abstract
Desirable information processing in space-terrestrial integrated vehicle networks (STINs) handles data distributed in different satellites while transmitting, where efficient modeling time-varying resources is critical. Existing works are not applicable to STINs, however, because they lack the joint consideration of different movement patterns and fluctuating loads. In this paper, we propose theTime-Varying Resource Graph (TVRG)to model dynamic resources in STINs, by leveraging the advantages of software-defined networking in flexible resource management. Firstly, we propose theSTIN mobility modelto uniformly model different movement patterns in STINs. Then, we propose alayered Resource Modeling and Abstraction (RMA)approach, where evolutions of node resources are modeled as Markov processes, by encoding predictable topologies and influences of fluctuating loads as states. Besides, we propose the low-complexity domain resource abstraction algorithm by defining two mobility-based and load-aware partial orders on resource abilities. Finally, we formulate theTVRG-based Processing on the Way (TPoW)problem for data flows with processing requirements and multiple sources. We propose aMulti-level Processing on the Way (MPoW)approach with a bounded approximation ratio, realizing adaptive matching of resources and demands of processing and transmission. To evaluate the RMA approach, we propose aTVRG-based Routing (TR)algorithm for time-sensitive and bandwidth-intensive data flows, with the multi-level on-demand scheduling ability. Comprehensive simulation results demonstrate that our RMA-TR and MPoW outperform most related schemes by decreasing nearly 40% bandwidth consumption with the shortest end-to-end delay.
Long Chen 0025, Feilong Tang 0001, Jiacheng Liu 0001, Xu Li 0012, Yanmin Zhu 0006, Jiadi Yu, Laurence T. Yang, Zhetao Li, Bin Yao 0002, Yichuan Yu
IEEE Trans. Mob. Comput.3
2024 WASP: Efficient Power Management Enabling Workload-Aware, Self-Powered AIoT Devices
abstract
The wide adoption of edge AI has heightened the demand for various battery-less and maintenance-free smart systems. Nevertheless, emerging Artificial Intelligence of Things (AIoT) are complex workloads showing increased power demand, diversified power usage patterns, and unique sensitivity to power management (PM) approaches. Existing AIoT devices cannot select the most appropriate PM tuning knob, and therefore they often make sub-optimal decisions. In addition, these PM solutions always assume traditional power regulation circuit which incurs non-negligible power loss and control overhead. This can greatly compromise the potential of AIoT efficiency. In this paper, we explore power management optimization for emerging self-powered AIoT devices. We propose WASP, a highly efficient power management scheme for workload-aware, self-powered AIoT devices. The novelty of WASP is two fold. First, it combines offline profiling and light-weight online control to select the most appropriate PM tuning knobs for the given DNN models. Second, it is well tailored to a reconfigurable voltage regulation module that can make the best use of the limited power budget. Our results show that WASP allows AIoT devices to accomplish 65.6% more inference tasks under a stringent power budget without any performance degradation compared with other existing approaches.
Xiaofeng Hou, Xuehan Tang, Jiacheng Liu 0001, Chao Li 0009, Luhong Liang, Kwang-Ting Cheng
IEEE Trans. Parallel Distributed Syst.3
2024 Adaptive Network Management Service Based on Control Relation Graph for Software-Defined LEO Satellite Networks in 6G
abstract
As the most important incremental component in the advent of the 6G era, Low-Earth-Orbit (LEO) satellite networks are becoming increasingly instrumental, and their integration with Software-Defined Networking (SDN) is progressively recognized as a potent strategy for evolving toward truly service-centric networks, where networks are flexiblely reconstructed based on the service demands. Within such networks, the SDN controllers are responsible for network management by making service-aware resource orchestration. Hence, the placement and assignment of controllers emerge as one of the most critical aspects of the network management service, which becomes particularly challenging when confronted with the unique complexities posed by LEO satellite networks, characterized by their highly dynamic topology and unpredictable load fluctuations. In this paper, for the first time, we tackle the issue of controller placement and assignment with a focus on delivering network management services. Firstly, we formulate theadaptive controller placement and assignmentproblem. Then, we propose thecontrol relation graph (CRG)to capture the control overhead. Next, we present theCRG-based controller placement and assignmentalgorithm and thesliding window based traffic prediction method. Thelookahead-based improvementalgorithm is designed to further decrease management costs. Finally, we conduct a series of theoretical analyses including time complexities. Extensive emulation results demonstrate that our algorithms outperform related schemes in terms of response time and load balancing.
Long Chen 0025, Feilong Tang 0001, Xu Li 0012, Jiacheng Liu 0001, Yanmin Zhu 0006, Jiadi Yu
IEEE Trans. Serv. Comput.4
2023 Label Aggregation with Self-Supervision Enhanced Graph Transformer
abstract
Aggregating noisy labels produced by the crowd of workers to generate true labels is a challenging problem in crowdsourcing. The key behind label aggregation is to effectively utilize the hidden information (e.g., characteristics of workers and questions which are often missing) in the labeling process. Existing methods mainly generated aggregation models based on the complicated Bayesian model or some strong assumptions. Recently, deep learning-based methods attempt to automate label aggregation but need various labels. These all make them hard to deploy to real-world applications. In fact, abundant information in the process of crowdsourcing itself can be extremely helpful to aggregate the labels. In this paper, we propose ATHENA (lAbel aggregaTion witH sElf-supervision eNhanced grAph transformer) to aggregate labels by utilizing the self-supervision signals in crowdsourcing. Firstly, we propose a transformer-based graph neural network that can learn from the crowdsourcing topology and features. Then, we use self-supervision signals inherently included in the dataset to help to aggregate the labels. To be specific, we identify the answer-based self-supervision signal that can predict the answer of any user given to different tasks. In our evaluations, we compare the proposed ATHENA with the other 11 representative methods on 10 datasets. Our experimental results demonstrate that ATHENA is highly effective in aggregating labels and obtains much better performance than existing methods.
Jiacheng Liu 0001, Feilong Tang 0001, Xiaofeng Hou
ECAI1
2023 MMExit: Enabling Fast and Efficient Multi-modal DNN Inference with Adaptive Network Exits
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Kwang-Ting Cheng, Li Li 0012, Minyi Guo
Euro-Par2
2023 Architecting Efficient Multi-modal AIoT Systems
abstract
Multi-modal computing (M2C) has recently exhibited impressive accuracy improvements in numerous autonomous artificial intelligence of things (AIoT) systems. However, this accuracy gain is often tethered to an incredible increase in energy consumption. Particularly, various highly-developed modality sensors devour most of the energy budget, which would make the deployment of M2C for real-world AIoT applications a difficult challenge.
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Jia Chen 0032, Luhong Liang, Kwang-Ting Cheng, Minyi Guo
ISCA2
2023 EAGLE: Heterogeneous GNN-based Network Performance Analysis
abstract
Performance analysis is of great importance for management and optimization of space-terrestrial integrated networks (STINs). Traditional approaches to network performance analysis are often based on idealized assumptions that are deviated from the real network environment. This leads to the fact that these models are usually inefficient and restricted in real-world STINs with complicated behavior and even dynamic capacity. In this paper, we propose a network performance analysis approach EAGLE based on heterogeneous graph neural networks. Firstly, we propose a powerful computer network representation model that can preserve all of the information in computer networks. It represents different components of computer networks as a set of heterogeneous nodes and edges, and finally constructs a heterogeneous graph. Then, we obtain the topological representation for the routers in the network through a bandwidth-aware network embedding model. Based on this heterogeneous graph, we propose a heterogeneous GNN model to accurately predict network KPIs because it can completely capture the rich topological and attribute information of computer networks. Experimental results demonstrate that EAGLE can accurately model different networks, and outperforms both traditional methods and the latest neural network-based methods.
Jiacheng Liu 0001, Feilong Tang 0001, Long Chen 0025, Xu Li 0012, Jiadi Yu, Yanmin Zhu 0006, Yichuan Yu, Yanqin Yang
IWQoS1
2023 SMG: A System-Level Modality Gating Facility for Fast and Energy-Efficient Multimodal Computing
abstract
Achieving low-latency and high-efficiency multimodal computing (MMC) is crucial for deploying high-performance autonomous embedded systems (AES) that has limited energy budgets. However, existing methods have mainly focused on optimizing the computing phase and have overlooked the significant energy and latency overhead during the sensing phase. Therefore, we propose SMG, a system-level modality gating facility to optimize this. Our approach introduces a software-defined DSP gating technique that enables MMC tasks to bypass both the sensing and computing phases of unimportant modalities. We also propose a raw data-activated MMC mechanism that comprises a fast modality tester and adaptive modality executor, which adapts to the modality gating architecture and performs energy-efficient MMC. To evaluate SMG, we implement a prototype of SMG by integrating it into existing AES and analyze it with extensive multimodal video recognition workloads. Our experimental results show that SMG outperforms SOTA approaches by adaptively gating some DSP operations, resulting in substantial improvements in both energy consumption and task latency.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Kwang-Ting Cheng, Minyi Guo
RTSS4
2023 INFER: Distilling knowledge from human-generated rules with uncertainty for STINs
Jiacheng Liu 0001, Feilong Tang 0001, Yanmin Zhu 0006, Jiadi Yu, Long Chen 0025, Ming Gao 0001
Inf. Sci.1
2023 Delay-Optimal Cooperation Transmission in Remote Sensing Satellite Networks
abstract
Many remote sensing applications, such as forest fire monitoring, need to send a large volume of data to the ground with low delay. Therefore, the cooperation transmission, which relies on cooperation among satellites to achieve continuous transmission, emerges as an indispensable technique. Most existing work cannot minimize the delay through dynamic cooperation transmission. In this paper, we investigate how to minimize the delay in remote sensing satellite networks based on cooperation transmission, where cooperation hotspots refer to the satellites with ground-satellite links to the Earth Stations (ESs). First, we propose the cooperation capability model to quantify capabilities of cooperation hotspots. Then, we formulate the satellite cooperation transmission problem and prove its NP-hardness. To solve the problem, we propose the delay-minimized cooperation transmission scheme. Both CCT and DCT algorithms adapt well to the dynamic topology and time-varying available resources. Finally, we formally analyze the approximation ratios and the time complexities of both algorithms. We also prove that the DCT always setups loop-free paths. NS2-based simulation results demonstrate that our schemes have good scalability, and both CCT and DCT algorithms reduce the end-to-end delay on average by more than 21.77%, and significantly improve throughput, packet loss rate and flow completion time.
Long Chen 0025, Feilong Tang 0001, Xu Li 0012, Jiacheng Liu 0001, Yanqin Yang, Jiadi Yu, Yanmin Zhu 0006
IEEE Trans. Mob. Comput.4
2023 Optimized Controller Provisioning in Software-Defined LEO Satellite Networks
abstract
The controller provisioning, which adjusts the number, locations, and members of satellite controllers adaptive to the dynamic network load and topology, fundamentally impacts the performance of software-defined satellite networks (SDSNs). An ideal provisioning strategy should achieve a low total control overhead throughout the entire satellite operation period, which is extremely challenging since the network loadcan only be predicted in a short time scale. Existing methods can hardly achieve this goal for they greedily configure controllers in each time slot, where switches have to frequently migrate from one controller to another. In this paper, we focus on achievingglobally optimized strategieswith onlycurrent network load information. We first propose a comprehensive control overhead model and formulate theControllerProvisioningProblem (CPP)in SDSNs as a non-convex integer programming problem. To solve the problem, we propose an approximate algorithm named AROA by introducing a regularization framework and based on randomized rounding. We theoretically derive its competitive ratio. To produce strategies in time for future large satellite constellations, we further propose a more efficient heuristic algorithm HROA. Evaluations on our built simulation system show that our proposed methods significantly outperform related schemes in control overhead, latency, and scalability.
Xu Li 0012, Feilong Tang 0001, Luoyi Fu, Jiadi Yu, Long Chen 0025, Jiacheng Liu 0001, Yanmin Zhu 0006, Laurence T. Yang
IEEE Trans. Mob. Comput.6
2022 Load-Adaptive and Energy-Efficient Topology Control in LEO Mega-Constellation Networks
abstract
The Low-Earth-Orbit (LEO) mega-constellation networks, by providing low-latency and high-speed communications, are becoming indispensable infrastructures for the future six-generation (6G) architecture. Consequently, the topology, with thousands of satellites equipped with batteries of limited life, has to be adaptively controlled with high energy efficiency. However, existing work lacks the joint consideration of energy efficiency and load adaptation. In this paper, we first propose the line-of-sight condition to determine the candidate ISL set. Next, we model the energy consumption of the LEO mega-constellation networks. Along this direction, we formulate the Load-Adaptive and Energy-Efficient (LAEE) topology control problem in LEO mega-constellation networks and prove its NP-hardness. Finally, we propose the Amortized Energy based Topology Control (AETC) algorithm to solve the LAEE problem, with good adaptation to the fluctuating load and guarantees connectivities between any two satellites. Extensive simulation results demonstrate that the AETC algorithm outperforms related schemes in terms of energy consumption and results in good topology stability.
Long Chen 0025, Feilong Tang 0001, Linghe Kong, Rui Li 0098, Zhi Hou, Jiacheng Liu 0001, Xu Li 0012, Song Guo 0001
GLOBECOM6
2022 WeAnimate: Motion-coherent animation generation from video data
Huanghao Yin, Jiacheng Liu 0001, Guoqiang Li 0001
Multim. Tools Appl.2
2022 Processing-While-Transmitting: Cost-Minimized Transmission in SDN-Based STINs
abstract
Existing Space-Terrestrial Integrated Network (STIN) applications collect all data from multiple satellites and terrestrial nodes to the specific analyze center on the earth for processing, which wastes lots of network resources. To save these resources, we propose a novelprocessing-while-transmittingpattern in the SDN-based STIN architecture. Through a logically centralized control plane, it cooperatively processes a complex task on appropriate nodes during data transmission. Here, the key point is to jointly determine the transmission path and place subtasks adaptive to data distributions, heterogeneous link costs, task characteristics, the dynamic topology, and network resources. In this paper, we firstly formulate theTransmission-cost-minimized joint Routing and Tasks placement Problem (TRTP)in time-varying STINs. We prove it is NP-hard and has no Polynomial-Time Approximation Scheme (PTAS). To solve the problem, we propose theJoint Routing and Task Placement (JRTP)algorithm. It first converts the time-varying STIN to a stable graph to cope with the network dynamics, according to the topology and resources during task processing. Then, it jointly decides the routing and task placement through atask-topology graph model, which converts the TRTP problem on the stable graph to the classic shortest path problem. We prove that the performance of JRTP is bounded in cases when transmission resources are sufficient and further improve it through the idea of reinforcement. The experimental results show that our processing pattern can significantly decrease the transmission cost and delay, and our algorithms outperform most related ones.
Xu Li 0012, Feilong Tang 0001, Yanmin Zhu 0006, Luoyi Fu, Jiadi Yu, Long Chen 0025, Jiacheng Liu 0001
IEEE/ACM Trans. Netw.7
2021 Truth Inference with Bipartite Attention Graph Neural Network from a Comprehensive View
abstract
As crowdsourcing has cast a new solution to numerous tasks, truth inference, which deduces the accurate answer from massive noise labels (answers), has become quite an essential issue. However, existing proposals of truth inference only excel at limited tasks since they excessively depend on modeling either workers or labels with simple assumptions. In this paper, we propose BAT (Bipartite Attention-driven Truth) to flexibly infer the truth in various scenarios. The key behind BAT is to explore a comprehensive approach from the whole topology of crowdsourcing itself rather than any individual component. Specifically, BAT firstly characterizes the crowdsourcing as an attributed bipartite graph (ABG). Then it deploys a bipartite graph neural network (bi-GNN). The bi-GNN relies on a bipartite attention mechanism for exploiting the importance of different answers to compute the correct one. For verifying BAT, we compare BAT with other eight existing truth inference methods on real-world datasets from different domains (image, text, audio). The results show that BAT performs best on different crowdsourcing tasks.
Jiacheng Liu 0001, Feilong Tang 0001, Jielong Huang
ICME1
2021 AlphaR: Learning-Powered Resource Management for Irregular, Dynamic Microservice Graph
abstract
The microservice architecture is a hot trend which proposes to transform the traditional monolith application into massive dynamic and irregular small services. To boost the overall throughput and ensure the guaranteed latency, it is desirable to process massive service requests in parallel with efficient resource sharing in data centers. However, the disaggregation nature of microservice unavoidably upscales the design space of resource management and increases its complexity. In this paper, we propose AlphaR, a learning-powered resource management system tailored to the microservice environment. The basic idea of AlphaR is to generate microservice-specific resource management policies for improving efficiency. Specifically, we take the first step to use bipartite graph as a convenient abstraction for application built with microservices. Based on this, we devise a bipartite feature inference approach named Bi-GNN to extract the temporal characteristics of microservices. Furthermore, we implement a policy network to select appropriate resource allocation choices for maximizing the performance in resource-constrained data centers. AlphaR can improve the mean and p95 response time by up to 80% and 77.5% respectively compared with conventional schemes.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Lu Zhang 0049, Shaolei Ren, Jingwen Leng, Quan Chen 0002, Minyi Guo
IPDPS3
2021 AUTO: Adaptive Congestion Control Based on Multi-Objective Reinforcement Learning for the Satellite-Ground Integrated Network
Xu Li 0012, Feilong Tang 0001, Jiacheng Liu 0001, Laurence T. Yang, Luoyi Fu, Long Chen 0025
USENIX ATC3
2021 Exploiting predicted answer in label aggregation to make better use of the crowd wisdom
Jiacheng Liu 0001, Feilong Tang 0001, Long Chen 0025, Yanmin Zhu 0006
Inf. Sci.1
2020 Fine-Grained Machine Teaching with Attention Modeling
abstract
The state-of-the-art machine teaching techniques overestimate the ability of learners in grasping a complex concept. On one side, since a complicated concept always contains multiple fine-grained concepts, students can only grasp parts of them during a practical teaching process. On the other side, because a single teaching sample contains unequal information in terms of various fine-grained concepts, learners accept them at different levels. Thus, with more and more complicated dataset, it is challenging for us to rethink the machine teaching frameworks. In this work, we propose a new machine teaching framework called Attentive Machine Teaching (AMT). Specifically, we argue that a complicated concept always consists of multiple features, which we call fine-grained concepts. We define attention to represent the learning level of a learner in studying a fine-grained concept. Afterwards, we propose AMT, an adaptive teaching framework to construct the personalized optimal teaching dataset for learners. During each iteration, we estimate the workers' ability with Graph Neural Network (GNN) and select the best sample using a pool-based searching approach. For corroborating our theoretical findings, we conduct extensive experiments with both synthetic datasets and real datasets. Our experimental results verify the effectiveness of AMT algorithms.
Jiacheng Liu 0001, Xiaofeng Hou, Feilong Tang 0001
AAAI1
2020 ANT-man: towards agile power management in the microservice era
abstract
The emerging trend of decomposing cloud applications into microservices has raised new questions about managing the performance/power trade-off of a datacenter at microsecondscale. We introduce ANT-Man, an Auto, Native and Transparent power Management framework that can exploit fine-grained microservice variability for system efficiency. To achieve this, ANT-Man abstracts away two major sources of latency overhead in traditional hierarchical power management frameworks. First, ANT-Man proposes an auto power budgeting scheme for reducing the power coordination latency at the datacenter level. It can proactively determine the power budget tailored to each individual microservice. Second, ANT-Man proposes a native and transparent power control scheme to overcome the power configuration latency for each microservice. It enables super-fast power budget enforcement with nanosecond-scale performance scaling. Extensive experiments on our prototyped system show that ANT-Man could slash power consumption by $ 7.8\sim 43.5\%$ and in the meantime reduce the $95^{\text{th}}$ tail latency by $ 9.7\sim 12.5\%$ compared to existing techniques.
Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Lu Zhang 0049, Yang Hu 0001, Minyi Guo
SC3
2019 Unleashing the Scalability Potential of Power-Constrained Data Center in the Microservice Era
abstract
Recent scale-out cloud services have undergone a shift from monolithic applications to microservices by putting each functionality into lightweight software containers. Although traditional data center power optimization frameworks excel at per-server or per-rack management, they can hardly make informed decisions when facing microservices that have different QoS requirements on a per-service basis. In a power-constrained data center, blindly budgeting power usage could lead to a power unbalance issue: microservices on the critical path may not receive adequate power budget. This unavoidably hinders the growth of cloud productivity.
Xiaofeng Hou, Jiacheng Liu 0001, Chao Li 0009, Minyi Guo
ICPP2
2018 Feedback Based High-Quality Task Assignment in Collaborative Crowdsourcing
abstract
Collaborative Crowdsourcing focuses on tasks that need to be finished by a group of workers, usually is required to get information about worker affinities. However, there are few works study the problem how to get the accurate estimation of worker affinity along with worker skill. Under this situation, the task assignment problem in a cold-start collaborative crowdsourcing platform will be even harder. In this paper, we design a novel collaborative crowdsourcing framework that allows task requesters as well as workers to provide feedback to the task executors or co-workers. Using the information from feedbacks, we build the model to estimate worker affinity and worker skill more accurately and based on which we propose a measurement to measure the matching degree between task and a group of workers. Next, we propose a heuristic task assignment called FCC-SA which averagely distribute skillful workers to all collaborative tasks. Finally, we hire crowd workers to conduct real experiments to test our model and algorithm. Experimental results demonstrate that our FCC-SA algorithm significantly outperforms representative related proposals, and our novel measurement also has better precision than traditional measure method.
Liang Qiao 0001, Feilong Tang 0001, Jiacheng Liu 0001
AINA3