EDBT 2026 Demo / reviewers in the wild / expert
Xiaofeng Hou
dblp:184/8262
· DBLP profile ↗
62ranked-venue papers
13as first author
56since 2021 · last 2026
0000-0003-4372-7851ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 11 first-author · 36 since 2021Computer networks · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache Compression
Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001 |
AAAI | 5 |
| 2026 | AdaReason: Progressive Training of Multi-LoRA Adapters for Budget-Adaptive Language Reasoning ModelsabstractLarge reasoning models (LRMs) have demonstrated remarkable capabilities in solving complex problems through extended chain-of-thought reasoning. However, existing approaches face a fundamental trade-off between computational efficiency and reasoning accuracy. Current methods either lack support for user-specified computational budgets or require maintaining multiple independent models, leading to significant resource overhead. In this paper, we present AdaReason, a unified framework that trains a single base model to support arbitrary user-defined computational budgets through dynamic adapter composition. Our approach introduces three key innovations: (1) a length-adaptive step reward function that stabilizes training across diverse budget constraints, (2) a progressive training strategy that gradually tightens computational bounds while maintaining model performance, and (3) a runtime adapter merging mechanism that dynamically interpolates between different computational preferences. Unlike existing methods that suffer from training instability in large context windows, AdaReason achieves stable convergence through careful reward shaping and progressive constraint tightening. Additionally, we provide a rigorous theoretical analysis, establishing a performance bound for our merged model. Experiments on different reasoning benchmarks demonstrate that AdaReason establishes a new state-of-the-art in the performance-efficiency trade-off and enables flexible runtime budget adaptation. Pengyu Cheng, Xiaofeng Hou, Jiacheng Liu 0001 |
AAAI | 4 |
| 2026 | Adaptive Spatial and Temporal Redundancy Optimization for Efficient Reasoning in Large Language ModelsabstractLarge Language Models (LLMs) have achieved exceptional performance in complex reasoning via Chain-of-Thought (CoT), yet the associated computational costs remain prohibitive. CoT reasoning contains significant untapped efficiency potential across two dimensions: temporal redundancy, where reasoning steps may be unnecessary, and spatial redundancy, where computations can be performed at reduced precision. While current optimization techniques often necessitate resource-intensive fine-tuning or data curation, we introduce ASTRO (Adaptive Spatial and Temporal Redundancy Optimization), a training-free framework that simultaneously addresses both dimensions. ASTRO leverages Dewey’s reflective thinking model to segment reasoning phases, applying a progressive precision reduction strategy coupled with an entropy-based confidence mechanism for adaptive termination. Empirical results across diverse reasoning benchmarks demonstrate that ASTRO achieves up to an 11.3 \times efficiency gain without compromising accuracy, highlighting the advantages of holistic multi-dimensional redundancy management over isolated optimization methods. Pengyu Cheng, Qiyuan Zhu, Hao Gu 0001, Ruijie Shen, Xiaofeng Hou, Sirui Han, Jiacheng Liu 0001 |
ACL (1) | 8 |
| 2026 | TierServe: Revenue-Maximizing LLM Inference Scheduling Across Multiple Subscription Tiers
Jiacheng Liu 0001, Xiaofeng Hou, Minyi Guo |
APPT | 3 |
| 2026 | OrbitGuard: Hierarchical Orbit-Aware Runtime for Spaceborne LLM Inference
Xiaofeng Hou, Jiacheng Liu 0001, Xiaozhi Zhu, Chao Li 0009 |
APPT | 2 |
| 2026 | MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert OffloadingabstractMixture-of-experts (MoE) architectures enable scalable Large Language Models (LLMs) with reduced computational overhead, yet their deployment on memory-constrained edge devices is hindered by substantial memory demands. Traditional expert-offloading techniques mitigate memory constraints but often significantly increase inference latency. We introduce MoE-APEX, an Adaptive Precision EXpert offloading system that optimizes MoE inference for edge architectures by dynamically managing expert precision. Our core innovation is to replace less critical cache-miss experts with low-precision variants, reducing loading latency while maintaining accuracy. MoE-APEX introduces three innovative techniques that map the natural hierarchy of MoE computation: (1) a token-level dynamic expert loading mechanism, (2) a layer-level adaptive expert prefetching technique, and (3) a sequence-level cost-aware expert caching policy. These innovations enable MoE-APEX to leverage the benefits of mixed-precision expert inference fully. Implemented atop Llama.cpp, MoE-APEX achieves decoding speedups ranging from 1.34x to 9.75x compared to state-of-the-art MoE offloading systems across diverse edge devices, offering a robust solution for efficient MoE deployment in resource-constrained environments. Jiacheng Liu 0001, Xiaofeng Hou, Yi-Fei Pu, Jing Wang 0055, Pheng-Ann Heng, Chao Li 0009, Minyi Guo |
ASPLOS (2) | 3 |
| 2026 | AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingabstractGenerative AI, especially LLM, is driving a fundamental shift in software paradigms, prompting cloud providers to build more efficient serving infrastructures. To meet the computational demands of emerging software, modern CPU processors are integrating Accelerator Units (AU) in the pipeline to accelerate key operations, such as Intel AMX for matrix multiplication. Current practices that dedicate AU-enabled CPU exclusively to LLM serving lead to significant resource waste and inferior efficiency. To this end, sharing AU-enabled CPU with general workloads is necessary to harvest redundant resources and improve platform performance-per-watt. However, perfectly sharing AU can be challenging since they introduce three-dimensional variations: variable usage patterns, compulsory frequency interferences, and dissimilar resource bounds. Existing resource managers are oblivious to complex Accelerator Unit Variations (AUV), resulting in performance and efficiency degradations of up to 50 % in shared environments. Therefore, this paper introduces AUM, a novel AU-aware resource manager designed to handle AUV and maximize the efficiency of shared processors. AUM has two cooperative components with three stages for three-dimensional AUV. The background profiler characterizes the usage, frequency, and resource information into a discrete model, guiding the runtime controller to analyze usage-aware requirements, select frequency-aware divisions, and make bound-aware resource decisions. Through extensive evaluations on production AU-enabled CPUs, we show that AUM improves CPU efficiency by$4.7-8.8 \%$while maintaining high-performance AU applications by reducing SLO violations by$\mathbf{7 - 1 1 \%}$compared with state-of-the-art resource managers. Xinkai Wang 0003, Chao Li 0009, Yiming Zhuansun, Jinyang Guo 0001, Xiaofeng Hou, Jing Wang 0055, Weigao Chen, Liping Zhang 0013, Minyi Guo |
HPCA | 5 |
| 2026 | CODO: An Automated Compiler for Comprehensive Dataflow Optimization
Weichuang Zhang, Yiquan Wang, Xinzhou Zhang, Chi Zhang 0005, Xiaofeng Hou, Chao Li 0009, Jieru Zhao, Minyi Guo |
ISCA | 6 |
| 2026 | LocMore: Locating More Bursty Latency-critical Jobs on Resource-constrained Nodes
Xiaofeng Hou, Xinkai Wang 0003, Jiacheng Liu 0001, Chao Li 0009, Minyi Guo |
IWQoS | 2 |
| 2026 | OpScope: Exploiting Operation-Driven Visual Scope for QoS-Stable Cloud Gaming
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo |
IWQoS | 4 |
| 2026 | StarkServe: A Framework for Elastic Serverless LLM Inference at the Extreme Edge
Xiaofeng Hou, Jiacheng Liu 0001, Xinkai Wang 0003, Chao Li 0009, Minyi Guo |
IWQoS | 2 |
| 2026 | DGS: A GPU-based Adaptive Graph Sampling FrameworkabstractGraph sampling plays a critical role in graph learning applications, notably within Graph Neural Networks (GNNs). Typically, the performance of GPU-based graph sampling is determined by the efficiency of sampling kernels. Different sampling methods excel under different conditions, and no single method consistently outperforms others in all scenarios. As sampling applications become increasingly complex, graph-related sparse operations can dominate the computational workload, with performance heavily influenced by storage formats. In this article, we propose DGS, a GPU-based graph sampling framework that can detach the kernel implementation from computation logic. In addition to sampling kernels, DGS jointly optimizes sparse graph kernels. It can adaptively switch between different execution strategies based on various inputs. Experiments show that DGS outperforms current state-of-the-art GPU sampling frameworks, achieving speedups ranging from 1.1× to 92.0×. This adaptability and performance improvement establish DGS as a highly effective and efficient solution for diverse graph sampling scenarios. Junyi Mei, Shixuan Sun, Chao Li 0009, Xinkai Wang 0003, Xiaofeng Hou, Minyi Guo, Yongchao Liu 0004, Chuntao Hong |
ACM Trans. Archit. Code Optim. | 6 |
| 2026 | Enabling Learning-Based Efficiency Optimizer With Shadow Cycles in Resource-Constrained Autonomous Embedded SystemsabstractThe emerging trend of autonomous embedded systems (AES) is promising to minimize human intervention in critical tasks. In the pursuit of maximal per-watt performance, the complex hardware and software of AES require intelligent energy efficiency optimizers (EO), and the stochastic runtime variances require continuous EO. However, deploying the desirable ondevice EO causes severe performance slowdown due to contention on limited computing power with the AES pipeline. We find that there are ignored and underutilized heterogeneous resources within AES for costly EO, which results from unbalanced accelerator behaviors and misaligned parallel inference executions. We experimentally and theoretically analyze theShadow Cycleswithin the realistic autonomous Bird’s Eye View pipeline on commercial embedded platforms, categorizing them into vertical and horizontal types with distinct properties.In this paper, we introduceSHEEO+, a continuous and intelligent energy efficiency optimizer that utilizes ignored heterogeneous shadow cycles. It achieves continuous and lightweight AES monitoring with the observation module, as well as intelligent and efficient AES power management with the optimization module. On the one hand,SHEEO+ observes both the internal runtime status and external environment variance with portable interfaces to capture shadow cycles and real-time states. On the other hand,SHEEO+ optimizes power configurations per iteration based on deep reinforcement learning (DRL) methods. It tailors DRL for two types of shadow cycles and invocates optimization processes based on resource availability. To extensively evaluateSHEEO+, we implement a prototype and deploy it on realistic edge platforms. The evaluation results show thatSHEEO+ utilizes up to 74.2% shadow cycles and achieves up to 18.6% energy efficiency improvements compared to state-of-the-art energy efficiency optimizers with negligible deployment overheads. Xinkai Wang 0003, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Minyi Guo, Yaqian Zhao |
IEEE Trans. Computers | 6 |
| 2026 | Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action RecognitionabstractFew-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios. Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Noisy Multi-Label Aggregation With Self-Supervised Graph Transformer in Mobile CrowdsourcingabstractAggregating noisy labels from mobile crowdsourcing (MCS) to recover true labels is a fundamental yet challenging problem, especially due to the sparsity and unreliability of crowd-contributed data. While most prior work addresses only single-label scenarios, real-world MCS applications often require robust solutions for both single-label and multi-label tasks, where each instance may be associated with multiple categories. In this paper, we propose ATHENA, a novel approach that leverages self-supervision signals inherent in MCS data for effective label aggregation. Firstly, we propose a graph transformer model that can learn from the MCS topology and features. Then, we propose self-supervision signals inherently included in the dataset to help aggregate the labels. To address the unique challenges of multi-label aggregation, we further extend our approach toATHENA+, introducing a label message passing (LMP) module that explicitly models correlations and dependencies among labels. We conducted extensive experiments on multiple single-label and multi-label classification datasets, comparing the proposed models with state-of-the-art methods. Our results demonstrate that ATHENA and ATHENA+ are highly effective in aggregating labels and obtain much better performance than existing methods. Jiacheng Liu 0001, Feilong Tang 0001, Hao Liu 0085, Long Chen 0025, Yanmin Zhu 0006, Jiadi Yu, Yichuan Yu, Xiaofeng Hou |
IEEE Trans. Mob. Comput. | 8 |
| 2025 | AsymServe: Demystifying and Optimizing LLM Serving Efficiency on CPU Acceleration Units
Xinkai Wang 0003, Yiming Zhuansun, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo |
APPT | 5 |
| 2025 | EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersabstractThe complexity of online applications is rapidly increasing, bringing more sophisticated performance anomalies in today's cloud datacenter. To fully understand application behaviors, we should obtain both inter-service communication data via RPC-level tracing and intra-service execution traces via application-level tracing to precisely reason about event causality. However, the average time overhead of existing intra-service tracing schemes on the traced applications is generally about 5-10%, possibly reaching 18% in the worst case. To realize practical intra-service tracing in shared and stressed datacenters, one must achieve extreme tracing efficiency with an overhead at the per-mille level. Xinkai Wang 0003, Xiaofeng Hou, Chao Li 0009, Yuancheng Li 0001, Du Liu, Guoyao Xu, Liping Zhang 0013, Yuemin Wu, Xiaopeng Yuan, Quan Chen 0002, Minyi Guo |
ASPLOS (2) | 2 |
| 2025 | Repurpose Accel-Sim for Next Generation NVIDIA Jetson GPU Architectural DesignabstractThe growing adoption of NVIDIA Jetson devices in edge-AI applications highlights the need for accurate architecture simulation tools on their integrated GPUs. Existing cycle-accurate GPU simulators primarily target traditional discrete GPUs and exhibit significant inaccuracies when applied to Jetson integrated GPUs. While Accel-Sim serves as the most widely used academic simulator for NVIDIA GPU research, its lack of support for the latest Jetson integrated GPUs severely hinders architectural exploration for next generation edge-AI devices.We propose Accel-Sim-J, which bridges the gap by repurposing Accel-Sim simulation framework to NVIDIA Jetson GPUs. We refine three major Accel-Sim framework components by applying tuner modifications, GPGPU-Sim performance model enhancements, and correlator adjustments. These improvements enable precise Jetson GPU simulation support, reducing simulation cycle errors from 29.0% to 22.7% on the Rodinia benchmark and from 26.1% to 16.1% on a transformer block. Furthermore, our enhanced architectural support for Ampere GPUs achieves a considerable reduction in simulation error (from 140.1% to 50.2%) for GEMM kernels.Based on Accel-Sim-J, we conduct a case study investigating the architecture design difference between an edge GPU and a traditional one. Specifically, we compare the optimal Compute-to-Cache (C2C) ratio by changing the L2 cache size of Jetson AGX Orin and RTX 3090. We conclude that Jetson GPUs demonstrate a higher optimal C2C ratio than discrete GPUs for the same workloads. We suggest that designers reduce the on-chip area proportion of the L2 cache in the next generation Jetson GPU design for better performance and efficiency. Chao Li 0009, Xiaofeng Hou, Yaqian Zhao, Jingwen Leng, Li Li 0012, Minyi Guo |
ISLPED | 4 |
| 2025 | TriCooling-Sim: Efficient Thermal Simulation for High-Density Micro AI Data Centers
Jinyang Guo 0001, Xinkai Wang 0003, Jing Wang 0055, Xiaofeng Hou, Chao Li 0009, Minyi Guo |
NPC (2) | 4 |
| 2025 | CGO: Cloud Game Orchestration via Resource Preception and CODEC Optimization
Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo |
NPC (2) | 4 |
| 2025 | SpaceExit: Enabling Efficient Adaptive Computing in Space with Early Exits
Jiacheng Liu 0001, Xiaozhi Zhu, Tongqiao Xu, Xiaofeng Hou, Chao Li 0009 |
USENIX ATC | 4 |
| 2025 | Power synchronization: taming massive diversified serverless functions under power constraints
Du Liu, Lu Zhang 0049, Yechen Xu, Xinkai Wang 0003, Yi-Fei Pu, Xiaofeng Hou, Chao Li 0009, Minyi Guo |
Sci. China Inf. Sci. | 7 |
| 2025 | FLAPS: fluctuation-aware power auction strategy for reducing the power overload probability
Xiaoqing Cai, Han Zhao 0005, Xiaofeng Hou, Weihao Cui, Quan Chen 0002, Chao Li 0009, Minyi Guo |
Frontiers Comput. Sci. | 3 |
| 2025 | MMBypass: Towards efficient multi-modal AI computing with adaptive bypass network
Yi-Fei Pu, Xinfeng Xia, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Jingling Yuan, Chao Li 0009 |
J. Parallel Distributed Comput. | 3 |
| 2025 | Enhancing High-Throughput GPU Random Walks Through Multi-Task Concurrency OrchestrationabstractRandom walk is a powerful tool for large-scale graph learning, but its high computational demand presents a challenge. While GPUs can accelerate random walk tasks, current frameworks fail to fully utilize GPU parallelism due to memory-to-compute bandwidth imbalance. In this article, CoWalker, an efficient GPU framework, is proposed to facilitate concurrent execution of random walks for high overall throughput. CoWalker features three novel designs. First, it incorporates a multi-level execution model that effectively orchestrates diverse walk tasks and reduces GPU stalls based on multiple graph characteristics. Second, it collaboratively manages graph data and streaming multiprocessors to minimize memory access interference and maximize core utilization under concurrent tasks. Finally, a multi-dimensional scheduler selects compatible random walk task combinations based on memory footprints to achieve maximum throughput. CoWalker significantly improves throughput over state-of-the-art baselines by mitigating concurrency overheads and effectively harnessing GPU parallelism. Our extensive evaluations on real-world workloads demonstrate that CoWalker achieves 2.75× higher overall system throughput compared with commercial tools and 1.56× over the SOTA academic system. Chao Li 0009, Xiaofeng Hou, Junyi Mei, Jing Wang 0055, Pengyu Wang 0003, Shixuan Sun, Minyi Guo, Baoping Hao |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Improving Efficiency in Multi-Modal Autonomous Embedded Systems Through Adaptive GatingabstractThe parallel advancement of AI and IoT technologies has recently boosted the development of multi-modal computing ($M^{2}C$) on pervasive autonomous embedded systems (AES).$M^{2}C$takes advantage of data from different modalities such as images, audio, and text and is able to achieve notable improvements in accuracy. However, achieving these accuracy gains often comes at the cost of increased computational complexity and energy consumption. Furthermore, the presence of numerous advanced sensors in these systems significantly contributes to power consumption, exacerbating the issue of limited power resources. Collectively, these challenges pose difficulties in deploying$M^{2}C$on small embedded devices with scarce energy resources. In this article, we propose anAdaptiveModalityGating technique calledAMGfor in-situ$M^{2}C$applications. The primary objective ofAMGis to conserve energy while preserving the accuracy advantages of$M^{2}C$. To achieve this goal,AMGincorporates two first-of-its-kind designs. Firstly, it introduces a novel semi-gating architecture that enables partial modality sensor power gating. Specifically, we devise the de-centralizedAMG(D-AMG) and centralizedAMG(C-AMG) architecture. The former buffers raw data on sensors while the latter buffers raw data on the computing board, which are suitable for different edge scenarios respectively. Secondly, it facilitates a self-initialization/tuning process on the AES, which is supported by carefully-built analytical model. Extensive evaluations demonstrate the effectiveness ofAMG. It achieves a 1.6x to 3.8x throughput higher than other power management methods and improves the lifespan of AES by 10% to 280% longer within the same energy budget, while satisfying all performance and latency requirements across various scenarios. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xuehan Tang, Kwang-Ting Cheng, Minyi Guo |
IEEE Trans. Computers | 1 |
| 2025 | BAT: A Versatile Bipartite Attention-Based Approach for Comprehensive Truth Inference in Mobile CrowdsourcingabstractThe proliferation of smart mobile devices has catalyzed the growth of Mobile CrowdSourcing (MCS) as a distributed problem-solving paradigm. MCS platforms heavily rely on advanced truth inference techniques to extract reliable information from diverse and potentially noisy crowd-contributed data. Existing truth inference models often made simplified assumptions about workers or tasks, employing complex Bayesian models or stringent data aggregation methods. These approaches tend to be task-specific, primarily limited to categorical labeling, making adaptations to other mobile computing scenarios labor-intensive. To address these limitations, we introduce the Bipartite Attention-driven Truth (BAT), a versatile approach tailored for mobile computing environments. BAT utilizes an Attributed Bipartite Graph (ABG) to holistically model the MCS process, with workers and tasks as nodes connected by edges representing answer-specific attributes. The approach employs a bipartite graph neural network with an innovative attention mechanism to assess the importance of different answers. BAT extends beyond categorical tasks to support numerical ones by incorporating novel feature representations and model extensions. Theoretical analyses clarify the link between answer similarity and worker expertise. Extensive experiments using diverse real-world datasets demonstrate BAT's superior performance compared to state-of-the-art categorical and numerical truth inference models, highlighting its effectiveness in mobile computing scenarios. Jiacheng Liu 0001, Feilong Tang 0001, Hao Liu 0085, Long Chen 0025, Yichuan Yu, Yanmin Zhu 0006, Jiadi Yu, Xiaofeng Hou, Pheng-Ann Heng |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | Sub-model Parallelism: A Scale-out Deployment Method for Large Multi-modal DNNsabstractWe have witnessed an increasing usage of multi-modal DNNs with multi-task heads on edge computing scenarios. These networks typically process inputs of different modalities first, then extract features for unified fusion, and finally input the fused features into multi-task heads. Such networks are often used to determine pose and navigate movement direction via multi-modal data obtained from diverse sensory equipment, therefore necessitating low inference latency. An edge device cluster with high-speed interconnection can be employed to support such DNN workload for scaled-out performance.For accelerating model inference on edge devices, previous researchers have proposed methods including model pruning, quantization, etc. However, these methods failed to take advantage of the structural features of multi-modal DNNs with multi-task heads and may impair the model’s prediction accuracy.Based on the intrinsic structure of multi-modal DNNs with multi-task heads, we propose Sub-model Parallelism to achieve scalable execution speedup. Sub-model Parallelism is a scale-out deployment method that first assigns preprocessing tasks of different modalities to different edge devices, then delivers them to a device for modality feature fusion, and finally distributes the fused features to other devices responsible for different task head computations. We run experiments on BEVFusion network and achieve an approximately 30% reduction in latency using two Jetson Orin devices connected by Remote Direct Memory Access (RDMA). Furthermore, we conduct a series of simulation experiments to cover scale-out scenarios and also achieve a good level of latency reduction. We hope that our proposed method can provide valuable experience for the optimized scale-out deployment of large multi-modal DNNs with multi-task heads on multiple edge devices. Xiaofeng Hou, Xiaozhi Zhu, Xinfeng Xia, Mingxi Chen, Chao Li 0009 |
CCGrid | 3 |
| 2024 | Improving the Efficiency of Serverless Computing via Core-Level Power ManagementabstractServerless computing has recently become a significant application paradigm in data centers. However, existing power management methods focus on optimizations at the coarse-grained server level, making them unable to handle the characteristics of these short-lived, dynamic serverless functions. In this context, the unawareness of function-level characteristics by the existing power management systems can severely degrade the energy efficiency of the data centers. To address this challenge, we design a function-level power management system. Instead of relying on server-level schedulers, we propose a novel core-level scheduling policy for serverless functions that can efficiently allocate functions to the most suitable CPU core. Additionally, we propose a power management mechanism for serverless computing that can reduce system power consumption with functions’ QoS guaranteed. Our evaluation shows that our system achieves a maximum power saving of 8.5% and an average power saving of 8% across the majority of loads without incurring any loss in tail latency, as compared to the conventional server-level scheduling system. Du Liu, Jing Wang 0055, Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Xiaoxiang Shi, Minyi Guo |
CCGrid | 6 |
| 2024 | SHEEO: Continuous Energy Efficiency Optimization in Autonomous Embedded SystemsabstractThe emerging trend of autonomous embedded systems minimizing human intervention has raised new questions about continuously maximizing system energy efficiency faced with stochastic runtime variance, which is costly for resource-constrained autonomous embedded systems. Considering heterogeneous hardware and variable software, we envision opportunities for vertical and horizontal shadow cycles within the AES pipeline for management facilities. This paper introduces SHEEO, a continuous energy efficiency optimizer that exploits underutilized heterogeneous computing resources to pursue variability-aware power management. To achieve this, SHEEO constantly monitors inner and outer variances and customizes reinforcement learning into two phases for stochastic runtime variance. We implement and deploy SHEEO on a commercial edge platform. The evaluation results show that SHEEO harvests up to 88% shadow cycles and improves up to 39% energy efficiency compared to state-of-the-art power management techniques with negligible overheads. Xinkai Wang 0003, Chao Li 0009, Qizheng Lyu, Xiaofeng Hou, Jingwen Leng, Minyi Guo |
ICCD | 5 |
| 2024 | AutoVCoder: A Systematic Framework for Automated Verilog Code Generation using LLMsabstractRecently, the use of large language models (LLMs) for software code generation, e.g., C/C++ and Python, has proven a great success. However, LLMs still suffer from low syntactic and functional correctness when it comes to the generation of register-transfer level (RTL) code, such as Verilog. To address this issue, in this paper, we develop AutoVCoder, a systematic open-source framework that significantly improves the LLMs' correctness of generating Verilog code and enhances the quality of its output at the same time. Our framework integrates three novel techniques, including a high-quality hardware dataset generation approach, a two-round LLM fine-tuning method and a domain-specific retrieval-augmented generation (RAG) mechanism. Experimental results demonstrate that AutoVCoder outperforms both industrial and academic LLMs in Verilog code generation. Code and models are available at https://github.com/sjtu-zhao-lab/AutoVCoder. Mingzhe Gao, Jieru Zhao, Zhe Lin 0007, Wenchao Ding 0001, Xiaofeng Hou, Yu Feng 0007, Chao Li 0009, Minyi Guo |
ICCD | 5 |
| 2024 | Graph Contrastive Learning for Truth InferenceabstractCrowdsourcing has become a popular paradigm for collecting large-scale labeled datasets by leveraging numerous annotators. However, these annotators often provide noisy labels due to varying expertise. Truth inference aims to infer accurate consensus labels from noisy crowdsourced annotations. Existing approaches rely heavily on hand-engineered assumptions or ground truth data, limiting their applicability. To address this, we propose GOVERN, a graph contrastive learning framework for truth inference without such external supervision. GOVERN employs a novel graph data augmentation strategy to generate views capturing worker coordination patterns. A contrastive objective then encourages invariant representations across views, enabling the discovery of features related to the hidden consensus. Further, a label correction method based on k-nearest neighbors refines noisy pseudo-labels to supervise model training. Comprehensive experiments on 9 real-world datasets demonstrate that GOVERN outperforms state-of-the-art truth inference techniques. Hao Liu 0085, Jiacheng Liu 0001, Feilong Tang 0001, Peng Li 0017, Long Chen 0025, Jiadi Yu, Yanmin Zhu 0006, Yanqin Yang, Xiaofeng Hou |
ICDE | 10 |
| 2024 | M2SN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware ApplicationsabstractMulti-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines. Yi-Fei Pu, Xiaofeng Hou, Jiacheng Liu 0001, Jing Wang 0055, Minyi Guo, Chao Li 0009 |
ICME | 3 |
| 2024 | CoCG: Fine-grained Cloud Game Co-location on Heterogeneous PlatformabstractCloud games have received widespread attention and exponential growth recently as a key technology for building metaverse. Unlike general tasks in the cloud, the scene-complex, latency-critical, and interaction-intensive features make it challenging for cloud game co-deployment on heterogeneous platforms. Game-grained resource allocation leads to low resource effectiveness. Although previous work tries to explore individual game partitioning methods, they still face the problem of inefficient game hosting decisions and ultimately QoS violations. In this paper, we propose a fine-grained game characteristic and scheduling strategy to co-locate games together for high resource usage effectiveness. First, we fully explore the relationships between game scenes and resource usage behaviors by breaking the cloud game into stages with multiple frames and clustering them. We adopt machine learning methods to predict game resource consumption in real-time. To further improve multi-game parallelism, we co-locate games in a complementary way and steal time from the loading stage to avoid oversubscribing. The evaluation shows that our work increased the throughput of the cloud game deployments by 23.7% with low overhead compared to previous work. Taolei Wang, Chao Li 0009, Jing Wang 0055, Xiaofeng Hou, Minyi Guo |
IPDPS | 5 |
| 2024 | A Tale of Two Domains: Exploring Efficient Architecture Design for Truly Autonomous ThingsabstractAutonomous Things (AuT) refers to a collection of self-sufficient tiny devices capable of performing intelligent computations. Looking ahead, AuT promises to enable ubiquitous deployment of intelligence on many emerging consumer electronics and mission-critical infrastructures. Nevertheless, there is an important research gap to date: architecting efficient AuT systems requires both energy autonomy (EA) and inference autonomy (IA). In other words, practical AuT application scenarios necessitate tailored architectures with significantly expanded inference performance and more efficient use of energy.We present CHRYSALIS, a novel automated EA/IA co-design methodology for autonomous things. It aims to guide the transition from a traditional EA-only and IA-only design approach to a truly AuT-oriented architecture design. To fully understand the interrelationship between the EA domain and the IA domain, CHRYSALIS first introduces an architectural modeling framework encompassing every key AuT module involving energy harvesting, intermittent execution, and accelerator control. Based on the holistic system model, we design an intelligent architecture generation tool that can help find the ideal design for targeted AuT scenarios adhering to different SWaP (Size, Weight and Power) constraints. To validate our work, we use CHRYSALIS for fast construction and exploration of efficient AuT design and pre-RTL design in representative AuT scenarios. Extensive evaluation shows that CHRYSALIS outperforms state-of-the-art designs and our proposed technique shows 56.4% better performance on average. We believe that the methodology and tools developed in this paper will foster the development of more performant and practical architectures in the upcoming AuT era. Xiaofeng Hou, Tongqiao Xu, Chao Li 0009, Jiacheng Liu 0001, Yang Hu 0001, Jieru Zhao, Jingwen Leng, Kwang-Ting Cheng, Minyi Guo |
ISCA | 1 |
| 2024 | CPM: A Cross-layer Power Management Facility to Enable QoS-Aware AIoT SystemsabstractWith the rapid progress and widespread adoption of AI technology, integrating powerful DNN models into AIoT devices in close proximity to users has become increasingly appealing. However, a significant challenge is that it is not easy to achieve the stringent Quality of Service (QoS) standards, especially in terms of real-time latency, demanded by the computationally intensive DNN workloads in energy-limited AIoT environments. To address this challenge, prior research has focused on per-layer power management techniques, which aggressively exploit the unique energy and performance relationships exhibited by each layer of the DNN at an exceedingly fine-grained control granularity. In this study, we identify the limitations of the existing per-layer DVFS mechanisms. They severely overlook the significant DVFS overhead caused by the excessively fine-grained control which can introduce complexity to power management in practical scenarios, consequently deteriorating QoS. To mitigate these challenges, we propose CPM, a Cross-layer Power Mangement facility which automatically modularizes different DNN layers and performs the best DVFS policy, thereby enhancing QoS by ensuring lower latency in real-time AIoT systems. Additionally, we integrate CPM into mainstream commercial AIoT boards and systems to validate its effusiveness. The results show that CPM can reduce the execution latency by up to 45.76% while improving the energy efficiency by up to 31.58% of real AIoT systems compared to SOTA per-layer power management methods. Xiaofeng Hou, Tongqiao Xu, Chao Li 0009, Minyi Guo |
IWQoS | 1 |
| 2024 | Jigsaw: Taming BEV-centric Perception on Dual-SoC for Autonomous DrivingabstractReal-time perception is important for autonomous driving. We observe an emerging trend using one large and critical fusion-based Bird’s-Eye-View (BEV) Deep Neural Network (DNN) model to perform core perception tasks. It collaborates with a few auxiliary Perspective-View (PV) models, forming a BEV-centric paradigm. Organizing the BEV and PV models respecting their distinct real-time requirements becomes challenging, especially on the state-of-the-practice GPU-integrated dual System-on-Chip (SoC) platform. It remains unclear how to appropriately allocate the separated GPU resource to BEV and PV models, satisfying their distinct real-time requirements with latency predictability. No public solution has been proposed for this emerging software-hardware combination.This paper explores parallelism and a timeslot-filling mechanism to organize tasks. We propose Jigsaw, a specialized execution timeline management framework for BEV-centric perception on dual-SoC. First, it exploits component parallelism to carefully place BEV model components and reduce BEV model latency. Second, we recognize two types of idle GPU timeslots left by a parallelized BEV model. The stable timeslot can offer hard real-time guarantee for PV models, while the unstable timeslot could only provide soft real-time capability. Therefore, Jigsaw schedules PV models by timeslot filling to ensure latency predictability of BEV model and deadline satisfaction of PV models. The framework is implemented in compliance with the practical computing stack in modern autonomous vehicles. It is evaluated on a dual-SoC prototype connected via a PCIe bus. Results show that it achieves $1.52-1.63 \times$ speedup for the BEV model compared to no parallelism. It also ensures deadline satisfaction for PV models without interference in BEV model latency predictability. Chao Li 0009, Xiaofeng Hou, Xinkai Wang 0003, Guangjun Bao, Bingchuan Sun, Shibo Rui, Minyi Guo |
RTSS | 3 |
| 2024 | Boosting Data Center Performance via Intelligently Managed Multi-backend Disaggregated MemoryabstractExisting disaggregated memory (DM) systems face a problem of underutilized far memory bandwidth, which greatly limits the data throughput when processing data-intensive applications. Specifically, prior works all target runtime design for a single PCIe-based secondary memory device (i.e., single-backend far memory) with low data bandwidth and high system overhead. In this work, we take the first step to realize a well-crafted, multi-backend DM system with scale-out far memory paths. We propose xDM, a novel DM management scheme that can dynamically build and implicitly select appropriate far memory access paths. As part of xDM, we devise a smart far memory configuration strategy that can further optimize bandwidth usage effectiveness by tuning a wide set of key parameters based on synthesized information of application page data. Our design shows up to $3.9 \times$ data swap performance speedup, $2.8 \times$ data throughput increase, and $5.1 \times$ data center task throughput improvement compared with state-of-the-art works. Jing Wang 0055, Hanzhang Yang, Chao Li 0009, Yiming Zhuansun, Wang Yuan, Xiaofeng Hou, Minyi Guo, Yang Hu 0001, Yaqian Zhao |
SC | 7 |
| 2024 | Practical Network Modeling Using Weak Supervision Signals for Human-Centric Networking in MetaverseabstractAs the metaverse continues to expand, it becomes increasingly critical to have human-centric networks that are both efficient and high-performing to optimize the user experience. Network modeling plays a fundamental role in optimizing and allocating resources efficiently, and configuring networks to satisfy the demands of diverse applications and users. Recently, traditional queuing theory-based approaches to network modeling have given way to machine learning-based methods. These methods rely on vast amounts of data for building precise models. Although high-precision simulators are ubiquitous, data collection is still an expensive and time-consuming process, resulting in a data bottleneck. In this paper, we propose a weakly supervised learning approach to modeling networks for human-centric networking in the metaverse. Specifically, we identify that queuing theory-based labels can be used to design the supervision signal at a very low cost. Therefore, we propose an approach that combines the inaccurate network modeling obtained from queuing theory-based approaches with an efficient and precise network model through only a small amount of simulation data. To make it a reality, we propose a novel neural network model that combines the powerful graph neural network and transformers. Additionally, we propose several additional supervision signals and a training algorithm to build a better network model. Experimental results demonstrate that our approach reduces the burden of data collection while achieving prediction accuracy comparable to results from large amounts of expensive simulation data. Furthermore, our approach exhibits superior generalization ability. Jiacheng Liu 0001, Feilong Tang 0001, Zhijian Zheng, Hao Liu 0085, Xiaofeng Hou, Long Chen 0025, Ming Gao 0001, Jiadi Yu, Yanmin Zhu 0006 |
IEEE J. Sel. Areas Commun. | 5 |
| 2024 | FlowWalker: A Memory-efficient and High-performance GPU-based Dynamic Graph Random Walk FrameworkabstractDynamic graph random walk (DGRW) emerges as a practical tool for capturing structural relations within a graph. Effectively executing DGRW on GPU presents certain challenges. First, existing sampling methods demand a pre-processing buffer, causing substantial space complexity. Moreover, the power-law distribution of graph vertex degrees introduces workload imbalance issues, rendering DGRW embarrassed to parallelize. In this paper, we propose FlowWalker, a GPU-based dynamic graph random walk framework. FlowWalker implements an efficient parallel sampling method to fully exploit the GPU parallelism and reduce space complexity. Moreover, it employs a sampler-centric paradigm alongside a dynamic scheduling strategy to handle the huge amounts of walking queries. FlowWalker stands as a memory-efficient framework that requires no auxiliary data structures in GPU global memory. We examine the performance of FlowWalker extensively on ten datasets, and experiment results show that FlowWalker achieves up to 752.2×, 72.1×, and 16.4× speedup compared with existing CPU, GPU, and FPGA random walk frameworks, respectively. Case study shows that FlowWalker diminishes random walk time from 35% to 3% in a pipeline of ByteDance friend recommendation GNN training. Junyi Mei, Shixuan Sun, Chao Li 0009, Cheng Chen 0008, Jing Wang 0055, Cheng Zhao 0001, Xiaofeng Hou, Minyi Guo, Bingsheng He, Xiaoliang Cong |
Proc. VLDB Endow. | 9 |
| 2024 | Potamoi: Accelerating Neural Rendering via a Unified Streaming ArchitectureabstractNeural Radiance Field (NeRF) has emerged as a promising alternative for photorealistic rendering. Despite recent algorithmic advancements, achieving real-time performance on today’s resource-constrained devices remains challenging. In this article, we identify the primary bottlenecks in current NeRF algorithms and introduce a unified algorithm-architecture co-design, Potamoi , designed to accommodate various NeRF algorithms. Specifically, we introduce a runtime system featuring a plug-and-play algorithm, SpaRW , which significantly reduces the per-frame computational workload and alleviates compute inefficiencies. Furthermore, our unified streaming pipeline coupled with customized hardware support effectively tames both SRAM and DRAM inefficiencies by minimizing repetitive DRAM access and completely eliminating SRAM bank conflicts. When evaluated against a baseline utilizing a dedicated DNN accelerator, our framework demonstrates a speedup and energy reduction of 53.1× and 67.7×, respectively, all while maintaining high visual quality with less than a 1.0 dB reduction in peak signal-to-noise ratio. Yu Feng 0007, Weikai Lin, Zihan Liu 0002, Jingwen Leng, Minyi Guo, Han Zhao 0005, Xiaofeng Hou, Jieru Zhao, Yuhao Zhu 0001 |
ACM Trans. Archit. Code Optim. | 7 |
| 2024 | A2: Towards Accelerator Level Parallelism for Autonomous Micromobility SystemsabstractAutonomous micromobility systems (AMS) such as low-speed minicabs and robots are thriving. In AMS, multiple Deep Neural Networks execute in parallel on heterogeneous AI accelerators. An emerging paradigm called Accelerator Level Parallelism (ALP) suggests managing accelerators holistically. However, there lacks a specialized and practical solution populating ALP for an AMS, where the varying real-time requirements under different working scenarios bring an opportunity to dynamically tradeoff between latency and efficiency. Furthermore, accelerator heterogeneity introduces enormous configuration space, and the shared-memory architecture results in dynamic bandwidth interference. In this article, we propose A 2 , a novel AMS resource manager optimizing energy and memory space efficiency under variable latency constraints. We gain insight from prior Learn&Control scheme to design an Analyze&Adapt scheme specialized for heterogeneous AI accelerators under shared-memory architecture. It features analyzing the system thoroughly offline to support two-step adaptation online. We build a prototype of A 2 and evaluate it on a commercial edge platform. We show that A 2 achieves 32.8% improvements in power and 13.8% in memory compared with control-based methods. As for timeliness enhancement, A 2 reduces the deadline violation rate by 9.2 percentage points (12.8% → 3.6%) on average compared to directly porting Learn&Control methods. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Xinkai Wang 0003, Quan Chen 0002, Minyi Guo |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | WASP: Efficient Power Management Enabling Workload-Aware, Self-Powered AIoT DevicesabstractThe wide adoption of edge AI has heightened the demand for various battery-less and maintenance-free smart systems. Nevertheless, emerging Artificial Intelligence of Things (AIoT) are complex workloads showing increased power demand, diversified power usage patterns, and unique sensitivity to power management (PM) approaches. Existing AIoT devices cannot select the most appropriate PM tuning knob, and therefore they often make sub-optimal decisions. In addition, these PM solutions always assume traditional power regulation circuit which incurs non-negligible power loss and control overhead. This can greatly compromise the potential of AIoT efficiency. In this paper, we explore power management optimization for emerging self-powered AIoT devices. We propose WASP, a highly efficient power management scheme for workload-aware, self-powered AIoT devices. The novelty of WASP is two fold. First, it combines offline profiling and light-weight online control to select the most appropriate PM tuning knobs for the given DNN models. Second, it is well tailored to a reconfigurable voltage regulation module that can make the best use of the limited power budget. Our results show that WASP allows AIoT devices to accomplish 65.6% more inference tasks under a stringent power budget without any performance degradation compared with other existing approaches. Xiaofeng Hou, Xuehan Tang, Jiacheng Liu 0001, Chao Li 0009, Luhong Liang, Kwang-Ting Cheng |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | Not All Resources are Visible: Exploiting Fragmented Shadow Resources in Shared-State Scheduler ArchitectureabstractWith the rapid development of cloud computing, the increasing scale of clusters and task parallelism put forward higher requirements on the scheduling capability at scale. To this end, the shared-state scheduler architecture has emerged as the popular solution for large-scale scheduling due to its high scalability and utilization. In such an architecture, a central resource state view periodically updates the global cluster status to distributed schedulers for parallel scheduling. However, the schedulers obtain broader resource views at the cost of intermittently stale states, rendering resources released invisible to schedulers until the next view update. These fleeting resource fragments are referred to as shadow resources in this paper. Current shared-state solutions overlook or fail to systematically utilize the shadow resources, leaving a void in fully exploiting these invisible resources. Xinkai Wang 0003, Yuancheng Li 0001, Chao Li 0009, Xiaofeng Hou, Jing Wang 0055, Quan Chen 0002, Jingwen Leng, Minyi Guo, Leibo Wang |
SoCC | 5 |
| 2023 | Label Aggregation with Self-Supervision Enhanced Graph TransformerabstractAggregating noisy labels produced by the crowd of workers to generate true labels is a challenging problem in crowdsourcing. The key behind label aggregation is to effectively utilize the hidden information (e.g., characteristics of workers and questions which are often missing) in the labeling process. Existing methods mainly generated aggregation models based on the complicated Bayesian model or some strong assumptions. Recently, deep learning-based methods attempt to automate label aggregation but need various labels. These all make them hard to deploy to real-world applications. In fact, abundant information in the process of crowdsourcing itself can be extremely helpful to aggregate the labels. In this paper, we propose ATHENA (lAbel aggregaTion witH sElf-supervision eNhanced grAph transformer) to aggregate labels by utilizing the self-supervision signals in crowdsourcing. Firstly, we propose a transformer-based graph neural network that can learn from the crowdsourcing topology and features. Then, we use self-supervision signals inherently included in the dataset to help to aggregate the labels. To be specific, we identify the answer-based self-supervision signal that can predict the answer of any user given to different tasks. In our evaluations, we compare the proposed ATHENA with the other 11 representative methods on 10 datasets. Our experimental results demonstrate that ATHENA is highly effective in aggregating labels and obtains much better performance than existing methods. Jiacheng Liu 0001, Feilong Tang 0001, Xiaofeng Hou |
ECAI | 3 |
| 2023 | MMExit: Enabling Fast and Efficient Multi-modal DNN Inference with Adaptive Network Exits
Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Kwang-Ting Cheng, Li Li 0012, Minyi Guo |
Euro-Par | 1 |
| 2023 | Architecting Efficient Multi-modal AIoT SystemsabstractMulti-modal computing (M2C) has recently exhibited impressive accuracy improvements in numerous autonomous artificial intelligence of things (AIoT) systems. However, this accuracy gain is often tethered to an incredible increase in energy consumption. Particularly, various highly-developed modality sensors devour most of the energy budget, which would make the deployment of M2C for real-world AIoT applications a difficult challenge. Xiaofeng Hou, Jiacheng Liu 0001, Xuehan Tang, Chao Li 0009, Jia Chen 0032, Luhong Liang, Kwang-Ting Cheng, Minyi Guo |
ISCA | 1 |
| 2023 | High-Throughput GPU Random Walk with Fine-Tuned Concurrent Query ProcessingabstractRandom walk serves as a powerful tool in dealing with large-scale graphs, reducing data size while preserving structural information. Unfortunately, existing system frameworks all focus on the execution of a single walker task in serial. We propose CoWalker, a high-throughput GPU random walk framework tailored for concurrent random walk tasks. It introduces a multi-level concurrent execution model to allow concurrent random walk tasks to efficiently share GPU resources with low overhead. Our system prototype confirms that the proposed system could outperform (up to 54%) the state-of-the-art in a wide spectral of scenarios. Chao Li 0009, Pengyu Wang 0003, Xiaofeng Hou, Jing Wang 0055, Shixuan Sun, Minyi Guo, Dongbai Chen, Xiangwen Liu |
PPoPP | 4 |
| 2023 | SMG: A System-Level Modality Gating Facility for Fast and Energy-Efficient Multimodal ComputingabstractAchieving low-latency and high-efficiency multimodal computing (MMC) is crucial for deploying high-performance autonomous embedded systems (AES) that has limited energy budgets. However, existing methods have mainly focused on optimizing the computing phase and have overlooked the significant energy and latency overhead during the sensing phase. Therefore, we propose SMG, a system-level modality gating facility to optimize this. Our approach introduces a software-defined DSP gating technique that enables MMC tasks to bypass both the sensing and computing phases of unimportant modalities. We also propose a raw data-activated MMC mechanism that comprises a fast modality tester and adaptive modality executor, which adapts to the modality gating architecture and performs energy-efficient MMC. To evaluate SMG, we implement a prototype of SMG by integrating it into existing AES and analyze it with extensive multimodal video recognition workloads. Our experimental results show that SMG outperforms SOTA approaches by adaptively gating some DSP operations, resulting in substantial improvements in both energy consumption and task latency. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Kwang-Ting Cheng, Minyi Guo |
RTSS | 1 |
| 2023 | Optimizing GPU-Based Graph Sampling and Random Walk for Efficiency and ScalabilityabstractGraph sampling and random walk algorithms are playing increasingly important roles today because they can significantly reduce graph size while preserving structural information, thus enabling computationally intensive tasks on large-scale graphs. Current frameworks designed for graph sampling and random walk tasks are generally not efficient in terms of memory requirement and throughput. Not to mention that some of them result in biased results. To solve the above problems, we introduce Skywalker+, a high-performance graph sampling and random walk framework on multiple GPUs supporting multiple algorithms. Skywalker+ makes four key contributions: First, it realizes highly paralleled alias method on GPUs. Second, it applies finely adjusted workload-balancing techniques and locality-aware execution modes to present a highly efficient execution engine. Third, it optimizes the GPU memory usage with efficient buffering and data compression schemes. Last, it scales to multi-GPU to further enhance the system throughput. Abundant experiments show that Skywalker+ exhibits significant advantage over the baselines both in performance and utility. Pengyu Wang 0003, Chao Li 0009, Jing Wang 0055, Taolei Wang, Lu Zhang 0049, Xiaofeng Hou, Minyi Guo |
IEEE Trans. Computers | 7 |
| 2022 | Exploring Efficient Microservice Level ParallelismabstractThe microservice architecture has recently become a driving trend in the cloud by disaggregating a monolithic application into many scenario-oriented service blocks (microservices). The decomposition process results in a highly dynamic execution scenario, in which various chained microservices contend for computing resources in different ways. While parallelism has been exploited at both the instruction/thread level and the task/request level, very limited work has been done with the grain-size of a microservice. Current parallel processing solutions are sub-optimal as they neither capture the unique characteristics of microservices nor consider the uncertainty arises in the microservice environment. In this work we introduce microservice level parallelism (MLP), a technique that aims to precisely coalesce and align parallel microservice chains for better system performance and resource utilization. We identify major issues that prevent servers from effectively exploiting MLP and we define metrics that can guide MLP optimization. We propose v-MLP, a volatility-aware MLP that is able to adapt to a highly heterogeneous and dynamic microservice environment. We show that v-MLP can reduce tail latency by up to 50% and improve resource utilization by up to 15 % under various scenarios. Xinkai Wang 0003, Chao Li 0009, Lu Zhang 0049, Xiaofeng Hou, Quan Chen 0002, Minyi Guo |
IPDPS | 4 |
| 2022 | Cloud-Native Server Consolidation for Energy-Efficient FaaS Deployment
Lu Zhang 0049, Yi-Fei Pu, Du Liu, Zeyi Lin, Xiaofeng Hou, Shang Yue, Chao Li 0009, Minyi Guo |
NPC | 6 |
| 2022 | Performance optimization for cloud computing systems in the microservice era: state-of-the-art and research opportunities
Xiaofeng Hou, Lu Zhang 0049, Chao Li 0009, Wenli Zheng, Minyi Guo |
Frontiers Comput. Sci. | 2 |
| 2022 | Tapping into NFV Environment for Opportunistic Serverless Edge Function DeploymentabstractEven with Network Function Virtualization (NFV), many commodity network servers have spare cycles. Despite that they are small and irregularly occur, spare cycles are fit for deploying short-lived serverless computing functions at the network edge. In this work, we perform detailed analyses of the benefits and limitations of co-locating serverless functions on NFV-ready servers. We proposeNEMO, a novel platform that enables efficient serverless edge function deployment in the NFV environment. NEMO can intelligently harvest spare cycles of network functions to warm up the serverless functions and speed up the function invocation in an agile manner. Besides, NEMO can judiciously manage the thread conflict in a resource-limited environment. We build a prototype of NEMO. Our thorough evaluations show that NEMO can harvest up to 41% spare cycles and achieve about 12.5$\sim$25X performance improvement compared with straightforward co-location. Lu Zhang 0049, Weiqi Feng, Chao Li 0009, Xiaofeng Hou, Pengyu Wang 0003, Jing Wang 0055, Minyi Guo |
IEEE Trans. Computers | 4 |
| 2022 | Integrated Power Anomaly Defense: Towards Oversubscription-Safe Data CentersabstractEnergy storage devices (e.g., batteries) are critical components for high-availability data center infrastructure today. Without resilient energy management of these devices, existing power-hungry data centers are largely unguarded targets for cyber criminals. Particularly for some of today's scale-out data centers, power infrastructure oversubscription unavoidably taxes the data center's backup energy resources (i.e., UPS), leaving very little room for dealing with power emergency. As a result, an attacker could manipulate the computing system to generate peak power demand and disrupt power-constrained server racks. This article aims at protecting data centers from malicious loads that seek to drain precious energy backup, overload server racks and compromise workload performance. We term such load as Elusive Power Peak (EPP) and demonstrate its basic three-phase attacking model. To defend against EPP, we propose IPAD, a remediation solution build on integrated software and hardware mechanisms. IPAD not only increases the attacking cost considerably by hiding vulnerable server racks from visible power peaks, but also strengthens the last line of defense against hidden power spikes with fine-grained power control strategy. We show that IPAD can effectively raise the bar of power-related attack, with reasonable design overhead. Xiaofeng Hou, Chao Li 0009, Jinghang Yang, Wenli Zheng, Xiaoyao Liang, Minyi Guo |
IEEE Trans. Cloud Comput. | 1 |
| 2021 | AlphaR: Learning-Powered Resource Management for Irregular, Dynamic Microservice GraphabstractThe microservice architecture is a hot trend which proposes to transform the traditional monolith application into massive dynamic and irregular small services. To boost the overall throughput and ensure the guaranteed latency, it is desirable to process massive service requests in parallel with efficient resource sharing in data centers. However, the disaggregation nature of microservice unavoidably upscales the design space of resource management and increases its complexity. In this paper, we propose AlphaR, a learning-powered resource management system tailored to the microservice environment. The basic idea of AlphaR is to generate microservice-specific resource management policies for improving efficiency. Specifically, we take the first step to use bipartite graph as a convenient abstraction for application built with microservices. Based on this, we devise a bipartite feature inference approach named Bi-GNN to extract the temporal characteristics of microservices. Furthermore, we implement a policy network to select appropriate resource allocation choices for maximizing the performance in resource-constrained data centers. AlphaR can improve the mean and p95 response time by up to 80% and 77.5% respectively compared with conventional schemes. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Lu Zhang 0049, Shaolei Ren, Jingwen Leng, Quan Chen 0002, Minyi Guo |
IPDPS | 1 |
| 2020 | Fine-Grained Machine Teaching with Attention ModelingabstractThe state-of-the-art machine teaching techniques overestimate the ability of learners in grasping a complex concept. On one side, since a complicated concept always contains multiple fine-grained concepts, students can only grasp parts of them during a practical teaching process. On the other side, because a single teaching sample contains unequal information in terms of various fine-grained concepts, learners accept them at different levels. Thus, with more and more complicated dataset, it is challenging for us to rethink the machine teaching frameworks. In this work, we propose a new machine teaching framework called Attentive Machine Teaching (AMT). Specifically, we argue that a complicated concept always consists of multiple features, which we call fine-grained concepts. We define attention to represent the learning level of a learner in studying a fine-grained concept. Afterwards, we propose AMT, an adaptive teaching framework to construct the personalized optimal teaching dataset for learners. During each iteration, we estimate the workers' ability with Graph Neural Network (GNN) and select the best sample using a pool-based searching approach. For corroborating our theoretical findings, we conduct extensive experiments with both synthetic datasets and real datasets. Our experimental results verify the effectiveness of AMT algorithms. Jiacheng Liu 0001, Xiaofeng Hou, Feilong Tang 0001 |
AAAI | 2 |
| 2020 | ANT-man: towards agile power management in the microservice eraabstractThe emerging trend of decomposing cloud applications into microservices has raised new questions about managing the performance/power trade-off of a datacenter at microsecondscale. We introduce ANT-Man, an Auto, Native and Transparent power Management framework that can exploit fine-grained microservice variability for system efficiency. To achieve this, ANT-Man abstracts away two major sources of latency overhead in traditional hierarchical power management frameworks. First, ANT-Man proposes an auto power budgeting scheme for reducing the power coordination latency at the datacenter level. It can proactively determine the power budget tailored to each individual microservice. Second, ANT-Man proposes a native and transparent power control scheme to overcome the power configuration latency for each microservice. It enables super-fast power budget enforcement with nanosecond-scale performance scaling. Extensive experiments on our prototyped system show that ANT-Man could slash power consumption by $ 7.8\sim 43.5\%$ and in the meantime reduce the $95^{\text{th}}$ tail latency by $ 9.7\sim 12.5\%$ compared to existing techniques. Xiaofeng Hou, Chao Li 0009, Jiacheng Liu 0001, Lu Zhang 0049, Yang Hu 0001, Minyi Guo |
SC | 1 |
| 2019 | Unleashing the Scalability Potential of Power-Constrained Data Center in the Microservice EraabstractRecent scale-out cloud services have undergone a shift from monolithic applications to microservices by putting each functionality into lightweight software containers. Although traditional data center power optimization frameworks excel at per-server or per-rack management, they can hardly make informed decisions when facing microservices that have different QoS requirements on a per-service basis. In a power-constrained data center, blindly budgeting power usage could lead to a power unbalance issue: microservices on the critical path may not receive adequate power budget. This unavoidably hinders the growth of cloud productivity. Xiaofeng Hou, Jiacheng Liu 0001, Chao Li 0009, Minyi Guo |
ICPP | 1 |
| 2019 | When Power Oversubscription Meets Traffic Flood Attack: Re-Thinking Data Center Peak Load ManagementabstractThe state-of-the-art techniques on data center peak power management are too optimistic; they overestimate their benefits in a potentially insecure operating environment. Especially in data centers that oversubscribe power infrastructure, it is likely that unexpected traffics can violate power budget before an effective network DoS attack is observed. In this work, we take the first to investigate the joint effect of power throttling and traffic flooding. We characterize a special operating region in which DoS attacks can provoke undesirable power peaks without exhibiting network traffic anomalies. In this region, an attacker can trigger power emergency by sending normal traffics throughout the Internet. We term this new type of threat as DOPE (Denial of Power and Energy). We show that existing technologies are insufficient for eliminating DOPE without negative performance effects on legitimate users. To enhance data center resiliency, we propose a request-aware power management framework called Anti-DOPE. The key feature of Anti-DOPE is bridging the gap between network traffic controlling and server power management. Specifically, it pre-processes of incoming requests to isolate malicious power attacks on the network load balancer side and then post-processes of compute node performance to minimize the collateral damage it may cause. Anti-DOPE is orthogonal to prior power management schemes and requires minute system modification. Using Alibaba container trace we show that Anti-DOPE allows 44% shorter average response time. It also improves the 90th percentile tail latency by 68.1% compared to the other power controlling methods. Xiaofeng Hou, Mingyu Liang, Chao Li 0009, Wenli Zheng, Quan Chen 0002, Minyi Guo |
ICPP | 1 |
| 2018 | Power Grab in Aggressively Provisioned Data Centers: What is the Risk and What Can Be Done About ItabstractAggressively provisioned data centers achieve great cost savings by over-committing the very expensive power distribution infrastructure. However, existing proposals for managing load power demand in such a data center are largely utilization-driven, overlooking power-related interferences among users. An important observation is that some tasks can impact existing power budget management framework and disrupt normal operation by taking away the precious public power capacity. This vulnerability exposes data centers to a new type of risk that we call power grab, which is essentially hostile power resource competition. It could worsen the performance-utilization tradeoff in a power-constrained computing environment. Anticipating a growing case for power-oriented com-petition, we propose CFP, a resilient power capacity management frame-work for improving the fairness and service quality in scale-out data centers. Our solution features a market-based power re-source allocation and billing scheme that involves users in the loop. It allows the data center to bypass the formidable task of identifying malicious users and defend against power grab with reward and punishment incentives. We build a proof-of-concept system and also evaluate our design with realistic Google cluster traces. Compared to prior arts, CFP can increase the average performance-cost ratio by 1.8X. It can boost the total throughput in an APDC by 15% under severe power contention. Our design allows scale-out data centers to safely exploit the benefits that power over-subscription may provide, with minor overhead. Xiaofeng Hou, Luoyao Hao, Chao Li 0009, Quan Chen 0002, Wenli Zheng, Minyi Guo |
ICCD | 1 |
| 2016 | Power Attack Defense: Securing Battery-Backed Data CentersabstractBattery systems are crucial components for mission-critical data centers. Without secure energy backup, existing under-provisioned data centers are largely unguarded targets for cyber criminals. Particularly for today's scale-out servers, power oversubscription unavoidably taxes a data center's backup energy resources, leaving very little room for dealing with emergency. Besides, the emerging trend towards deploying distributed energy storage architecture causes the associated energy backup of each rack to shrink, making servers vulnerable to power anomalies. As a result, an attacker can generate power peaks to easily crash or disrupt a power-constrained system. This study aims at securing data centers from malicious loads that seek to drain their precious energy storage and overload server racks without prior detection. We term such load as Power Virus (PV) and demonstrate its basic two-phase attacking model and characterize its behaviors on real systems. The PV can learn the victim rack's battery characteristics by disguising as benign loads. Once gaining enough information, the PV can be mutated to generate hidden power spikes that have a high chance to overload the system. To defend against PV, we propose power attack defense (PAD), a novel energy management patch built on lightweight software and hardware mechanisms. PAD not only increases the attacking cost considerably by hiding vulnerable racks from visible spikes, it also strengthens the last line of defense against hidden spikes. Using Google cluster traces we show that PAD can effectively raise the bar of a successful power attack: compared to prior arts, it increases the data center survival time by 1.6~11X and provides better performance guarantee. It enables modern data centers to safely exploit the benefits that power oversubscription may provide, with the slightest cost overhead. Chao Li 0009, Zhenhua Wang 0007, Xiaofeng Hou, Haopeng Chen, Xiaoyao Liang, Minyi Guo |
ISCA | 3 |