VLDB 2026 Research / reviewers in the wild / expert
Jianguo Yao 0002
dblp:34/6356-2
· DBLP profile ↗
73ranked-venue papers
13as first author
31since 2021 · last 2026
0000-0002-1142-4496ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 39 · 6 first-author · 19 since 2021Software engineering, systems software and programming languages · 12 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 1 since 2021Computer networks · 7 · 2 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Security and privacy · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | REPA: Reconfigurable PIM for the Joint Acceleration of KV Cache Offloading and ProcessingabstractThe use of KV cache in LLM inference leads to large memory footprint and sub-optimal decoding performance. Prior studies typically address one of these two limitations by either offloading or stage-split inference. In this paper, we explore and reveal the possibility of a joint solution, and propose REPA, a GPU-PIM hybrid system to prototype this idea. We leverage reconfigurable ReRAM PIM to achieve fast KV cache persistence, and balance the requirement of processing speed and memory capacity. To fully unleash the parallelization potential of REPA, we propose optimizations in (1) architecture, (2) data mapping and (3) pipelining: (1) We propose bulk-wise memory instructions and multi-level controllers to enable finer-grained parallelism in the PIM device. (2) We propose locality-aware data mapping to make the best of the aforementioned architectural optimization, and reduce long-range data transfer on chip. (3) We adopt sub-batch pipelining to reduce idleness in batches, and propose transfer overlapping to shadow the KV cache transfer by computation. Experimental results show that REPA exhibits high inference speed, energy efficiency and integratability. It is 1.5--6.5× faster, and 8--10× more efficient than NVIDIA A100. It also outperforms state-of-the-art DRAM PIM systems by up to 1.4× for long context inference. When integrated into existing offloading systems, REPA achieves 1.4--2.0× offloading speed, and 1.2--1.4× end-to-end speedup, showcasing its high potential for fast KV cache offloading and processing. Junlong Yang, Bo Peng 0043, Jianguo Yao 0002 |
ASPLOS (2) | 4 |
| 2026 | Mitigating Cold Starts in Container-Ephemeral Large Language Model Serving
Bo Peng 0043, Jianguo Yao 0002 |
IWQoS | 3 |
| 2026 | SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided SwappingabstractRecent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be transformed to matrix multiplications. However, this transformation introduces undesired sparsity into the kernel matrix, leading to significant redundant computation. Qiqi Gu 0002, Chenpeng Wu, Heng Shi 0005, Jianguo Yao 0002 |
PPoPP | 4 |
| 2026 | Effective Offline LLM and DNN based Matching, Filtering and Ranking for Search Ads Retrieval in E-CommerceabstractFor search Ads retrieval in e-Commerce, when a customer inputs a query, from tens of millions of Ads products, the system needs to quickly retrieve a limited number of candidates that can not only meet the customer's search intension but also have high conversion probability. This constructs challenges for both retrieval efficiency and retrieval quality on the perspective of relevance and engagement. To alleviate these challenges, traditional solutions usually adopt a cascading architecture including the recall, pre-ranking, ranking and re-ranking stages. These methods are helpful, however, due to the hard limit of online customer request processing latency, only a small ratio of relevant candidates can be retrieved and the quality of pre-ranking is limited. Different from the online cascading architecture, we propose an effective offline-based solution: for each query, we first use multi-path recall, such as BERT-based embedding to retrieve a much larger number of candidates. Then, a carefully designed LLM-based relevance model is used to filter out the irrelevant candidates. Finally, we use powerful DNN models to rank the left candidates, and only a small number of best candidates are passed to the online system for real-time usage. In the online system, we keep a single ranking stage and use the most powerful models only. Compared to the online cascading architecture, our proposed method can not only largely reduce the online system's overhead and latency, but also significantly improve the overall quality. Real world A/B test experiments on Coupang search Ads show that this solution can improve the revenue, GMV, and advertiser ROAS by 5.79%, 11.37%, and 6.69%, respectively. The relevance defect ratio can be reduced by 68%. Meanwhile, the customer request processing P99 latency can reduced by 21%, when the online serving cluster's hardware cost is reduced by 52%. Shuping Ji, Jianguo Yao 0002 |
SIGIR | 2 |
| 2026 | Scaling NVMM-based file system on intensive shared file access
Qiqi Gu 0002, Chenpeng Wu, Bingheng Yan, Jianguo Yao 0002 |
J. Syst. Archit. | 6 |
| 2026 | SpaceFusion++: An operator fusion scheduler for neural language model inference
Jianguo Yao 0002, Haibing Guan |
J. Syst. Archit. | 2 |
| 2026 | Zero2M: Optimizing Tenant-Level I/O Management for Future Faster NVMe Storage with FPGAabstractHigh-speed Non-Volatile Memory Express (NVMe) Solid-State Drives (SSDs) are shared by multiple tenants in cloud scenarios to improve resource utilization. Tenant-level I/O management is necessary to achieve reliable QoS control during sharing. Unfortunately, our investigation finds that CPUs inevitably participate in I/O management for existing solutions because SSDs are not tenant-sensitive and have limited internal computing resources. It introduces additional CPU costs and latency overhead when serving future faster SSDs. We propose that the Field Programmable Logic Gate Array (FPGA) is a promising alternative for freeing tenant-level I/O management from CPUs. However, implementing tenant-level I/O management using the FPGA requires addressing the following challenges: (1) System compatibility and tenant identification; (2) Efficient FPGA workflows that will not become a bottleneck; (3) Fast I/O management workflow that introduces the lowest additional CPU costs and latency. This article presents Zero2M, a novel CPU-free system designed to optimize the additional CPU costs and latency overhead in tenant-level I/O management for future faster NVMe SSDs. Zero2M proposes a dedicated FPGA-based NVMe controller to preserve system compatibility and identify tenants using the namespace mechanism in NVMe. It allows I/O management without modifying host software, which existing solutions cannot achieve. The parallelized and pipelined workflows are proposed in the controller to accelerate I/O command processing and prevent the controller from becoming a bottleneck for the I/O management workflow. The read/write speed of the Zero2M controller is 4.65 \(\times\) /4.92 \(\times\) faster than the state-of-the-art hardware-accelerated controller. The I/O management workflow is formulated as a novel parallelized and pipelined accelerator and integrated into the workflow of Zero2M’s controller. It optimizes additional CPU costs and latency overhead for tenant-level I/O management. Experiments present that Zero2M reduces an average of 3.01 \(\times\) CPU usage while maintaining the lowest latency overhead (7.62 \(\times\) lower on average) compared to the state-of-the-art solution. It also removes the CPU dependency for tenant-level I/O management for the first time. Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2025 | Postiz: Extending Post-increment Addressing for Loop Optimization and Code Size ReductionabstractMemory access instructions with auto-addressing modes are prevalent in various Instruction Set Architectures (ISAs), yet their use in compilers remains limited. Existing methods address code optimization in one of two ways: they either focus on reducing code size, but are constrained to basic block-level optimizations and may not fully exploit architectural benefits, or they optimize loop performance, often neglecting the advantages of post-increment instructions and focusing primarily on innermost loops while leaving outer loops unoptimized. To address these shortcomings and meet the needs of real-world Machine Learning (ML) applications, we introduce Postiz, a novel post-increment loop optimization technique. Postiz extends post-increment optimizations beyond traditional limits, incorporating enhancements for inner loops, cross-loop regions, and nested loop structures. Through a profitability analysis, Postiz optimizes code judiciously, leveraging architectural advantages and reducing code size without compromising improvement made by other optimizations. Our experiments show that Postiz is effective, achieving an optimization coverage of 98.04% on MobileNet and BERT benchmarks. In comparison to default LLVM optimization, Postiz generates approximately four times more post-increment instructions. Moreover, it reduces code size by an average of 9.45% across various platforms. These improvements represent significant advancements over current methods, showcasing Postiz’s potential to enhance compiler optimizations in a meaningful way. Enming Fan, Xiaofeng Guan, Heng Shi 0005, Hao Zhou 0009, Jianguo Yao 0002 |
CGO | 6 |
| 2025 | Leopard: Hardware Pass-Through Remote Storage Access with Queue Concurrency for Edge Intelligent WorkstationsabstractEdge intelligent workstations load (store) massive empirical data from (to) remote cloud storage due to limited local storage. However, current remote storage access frameworks are complex. They use expensive computing resources to manipulate multiple concurrent request queues in modern high-speed storage devices for saturating performance. Complex software stacks and limited CPUs on edge intelligent workstations hinder saturating concurrent request queues, thus resulting in up to $75 \%$ performance degradation for remote storage. We propose Leopard, a hardware pass-through remote storage access framework with queue concurrency, which provides lossless remote storage access for edge intelligence. Leopard proposes a custom NVMe controller using SmartNIC’s FPGA core to emulate it as an NVMe device, which eliminates complex remote storage stacks for edge workstations. Operations for remote storage access are implemented as hardware circuits inside the controller to eliminate CPU cycles. Parallelized and pipelined workflows are proposed for hardware circuits to accelerate remote storage access operations. Our evaluation presents that Leopard exhibits $1.09 \times \sim 6.04 \times$ lower remote storage access latency than SOTA solutions for realistic workloads in edge intelligent workstations. Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
DAC | 3 |
| 2025 | Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor CoresabstractThe escalating size of Mixture-of-Experts (MoE) based Large Language Models (LLMs) presents significant computational and memory challenges, necessitating innovative solutions to enhance efficiency without compromising model accuracy. Structured sparsity emerges as a compelling strategy to address these challenges by leveraging the emerging sparse computing hardware. Prior works mainly focus on the sparsity in model parameters, neglecting the inherent sparse patterns in activations. This oversight can lead to additional computational costs associated with activations, potentially resulting in suboptimal performance. Chenpeng Wu, Qiqi Gu 0002, Heng Shi 0005, Jianguo Yao 0002, Haibing Guan |
EuroSys | 4 |
| 2025 | SpaceFusion: Advanced Deep Learning Operator Fusion via Space-Mapping GraphabstractThis work proposes SpaceFusion, an advanced scheduler for efficient deep learning operator fusion. First, we develop a novel abstraction, the Space-Mapping Graph (SMG), to holistically model the spatial information of both inter- and intra-operator dependencies. Subsequently, we introduce the spatial and temporal slicers to decompose the fused spaces defined in SMGs, generating fusion schedules by analyzing and transforming dependencies. Finally, we present auto-scheduling methods that use the slicers to automatically create high-performance fusion schedules tailored to specific hardware resource configurations. End-to-end performance evaluations reveal that SpaceFusion achieves up to 8.79x speedup (3.54x on average) over baseline implementations from Huggingface for Transformer models, and a maximum of 2.21x speedup compared to the state-of-the-art manually-tuned implementations powered by FlashAttention. Jianguo Yao 0002, Haibing Guan |
EuroSys | 2 |
| 2025 | ReHSS: Optimizing Latency for Cloud Hybrid Storage Systems Using in-Network PlacementabstractModern cloud hybrid storage systems have been concentrating on strategic data placement to provide low disk I/O latency for various workloads. However, previous studies primarily focus on exploring adaptive data placement algorithms with high placement accuracy, overlooking the computing latency introduced by these algorithms. It significantly increases the end-to-end latency that determines the quality of service (QoS) for hybrid storage systems. We propose ReHSS, a novel in-network data placement framework that optimizes the end-to-end latency for hybrid storage systems using modern SmartNIC. We first investigate the overhead of adaptive data placement for hybrid storage systems. Then, we propose a comprehensive hardware/software co-optimization solution based on in-network processing that includes algorithm acceleration, data transmission and processing, and computing and communication overlapping. Experimental results present that compared to the SOTA solution, ReHSS optimizes the end-to-end latency for hybrid storage systems by$1.54 \times \sim 15.18 \times$. Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
IWQoS | 3 |
| 2025 | <tt>STRCMP</tt>: Integrating Graph Structural Priors with Language Models for Combinatorial Optimization
Xijun Li, Jiexiang Yang, Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
NeurIPS | 5 |
| 2025 | SHC-DP: Software-hardware collaborative in-network data placement for hybrid storage systems
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
J. Syst. Archit. | 3 |
| 2025 | Efficient Parallel Boolean Expression MatchingabstractBoolean expression matching plays an important role in many applications. However, existing solutions still show efficiency and scalability limitations. For example, existing solutions often exhibit degraded performance when applied to high-dimensional and diverse workloads, and existing algorithms rarely consider supporting concurrent matching and index updating under multicore environments. To overcome these limitations, in this article, we first design the PS-Tree data structure to efficiently index Boolean expressions in one dimension. By dividing predicates into disjoint predicate spaces, PS-Tree achieves high matching performance and good expressiveness. Based on the PS-Tree , we propose a Boolean expression matching algorithm called PSTDynamic . By dynamically adjusting the index and efficiently filtering out a large proportion of unmatching expressions, PSTDynamic achieves high matching performance under high-dimensional and diverse workloads. For multicore environment, we further extend the PSTDynamic algorithm to PSTParallel to achieve scalability with lower matching latency and higher matching throughput. We run experiments on both synthetic and real-world datasets. The experiments verify that our proposed algorithms show high efficiency and parallelism. Moreover, they also achieve fast index construction and a small memory footprint. Comprehensive experiments show that our solutions drastically outperform state-of-the-art methods. Shuping Ji, Jianguo Yao 0002, Wei Wang 0049, Jun Wei 0001, Hans-Arno Jacobsen |
ACM Trans. Database Syst. | 2 |
| 2024 | Boost Linear Algebra Computation Performance via Efficient VNNI UtilizationabstractIntel's Vector Neural Network Instruction (VNNI) provides higher efficiency on calculating dense linear algebra (DLA) computations than conventional SIMD instructions. However, existing auto-vectorizers frequently deliver suboptimal utilization of VNNI by either failing to recognize VNNI's unique computation pattern at the innermost loops/basic blocks, or producing inferior code through constrained and rudimentary peephole optimizations/pattern matching techniques. Auto-tuning frameworks might generate proficient code but are hampered by the necessity for sophisticated pattern templates and extensive search processes. Hao Zhou 0009, Qiukun Han, Heng Shi 0005, Yalin Zhang 0004, Jianguo Yao 0002 |
ASPLOS (3) | 5 |
| 2024 | PresCount: Effective Register Allocation for Bank Conflict ReductionabstractModern processors with large multi-banked register files often rely on hardware solutions to resolve bank conflicts efficiently. However, these hardware-based methods, while flexible, can incur runtime penalties and restrict the exploration of optimized hardware designs. In contrast, compiler-based methods for register bank assignments avoid runtime overhead. However, incorporating bank assignment into the complex register allocation process presents significant challenges, leading existing methods to adopt conservative approaches to avoid potential side effects. This paper introduces the novel register allocation method PresCount, which enhances the coloring strategy for the Register Conflict Graph (RCG) and incorporates a bank pressure tracking mechanism to improve performance. The integrated register bank assigner in PresCount effectively reduces bank conflicts, achieving remarkable reductions of 43.28% and 27.76%, respectively, compared to existing methods on platforms with rich register banks and limited register budgets, as demonstrated by SPECfp and CNN-KERNEL benchmarks. Furthermore, a subgroup splitting technique is introduced to facilitate register allocation under the bank-subgroup register file design, specifically our Domain-Specific Architecture (DSA) for AI computing. This technique demonstrates an impressive 99.85% reduction in bank conflicts for domain-specific kernel functions. By addressing the challenges of bank conflicts in register allocation, the proposed PresCount method showcases significant improvements in performance and efficiency for platforms with different register configurations and domain-specific workloads, allowing for more flexible exploration of optimized hardware designs. Xiaofeng Guan, Hao Zhou 0009, Guoqing Bao, Handong Li, Jianguo Yao 0002 |
CGO | 6 |
| 2024 | UFront: Toward A Unified MLIR Frontend for Deep LearningabstractAutomatic code generation for ML systems has gained popularity with the advent of compiler techniques like Multi-Level Intermediate Representation (Multi-Level IR, or MLIR). State-of-the-art MLIR frontends, including IREE-TF, Torch-MLIR, and ONNX-MLIR, aim to bridge the gap between ML frameworks and low-level hardware architectures through MLIR's progressive lowering pipeline. However, existing MLIR frontends encounter challenges such as inflexible high-level IR conversion, limited higher-level optimization opportunities, and reduced compatibility and efficiency, leading to software fragmentation and restricting their practical applications within the MLIR ecosystem. To address these challenges, we introduce UFront, a unified MLIR frontend employing a two-stage operator-to-operator compilation workflow. Unlike traditional frontends that compile model source code into binaries step by step with different MLIR transform passes, UFront decouples the process into two distinct stages. It first performs instantaneous model tracing, delegates traced computing nodes as standard Deep Neural Network (DNN) operators and transforms models written in different frameworks into unified high-level IR without relying on MLIR passes, enhancing conversion flexibility. Meanwhile, it performs high-level graph optimizations such as constant folding and operator fusion to produce more efficient high-level IR. In the second stage, UFront directly converts high-level IR into standard TOSA IR using proposed lowering patterns, eliminating transform redundancies and ensuring lower-level compatibility with existing ML compiler backends. This two-stage compilation approach enables consistent end-to-end code generation and optimization of various DNN models written in different formats within a single workflow. Extensive experiments on popular DNN models written in various frameworks demonstrate that UFront exhibits higher compatibility, faster end-to-end compilation, and is capable of producing more efficient binary execution compared to SOTA works. Guoqing Bao, Heng Shi 0005, Chengyi Cui, Yalin Zhang 0004, Jianguo Yao 0002 |
ASE | 5 |
| 2024 | Ripple: Large-Scale Service and Configuration Management in the CloudabstractMicroservice architectures backed by container technology have been widely used in many real-world cloud-native applications. By enabling customers to manage their services and configurations in the cloud in a centralized, externalized, and dynamic manner, efficient service and configuration management plays a fundamental role in building cloud-native service-centric applications. The number of containers in cloud data centers continues to increase. For example, in the Alibaba Cloud, the number of containers reached hundreds of thousands by 2023 and is expected to reach several million soon. At this scale, existing service and configuration management solutions have limited efficiency, scalability and robustness. Other related approaches, such as message bus systems and publish/subscribe (pub/sub for short) systems, also do not work well for large-scale service and configuration management in the cloud, as their designs are more general purpose directed. To overcome these limitations, we design a system, called Ripple, that uniquely combines several existing and some novel features such as consistent hashing-based workload distribution, dynamic destination list-based and client-assisted message delivery, incremental update, and adaptive load balancing. Approaches exhibiting these features have not been well investigated in the domain of service and configuration management. We compare our proposed solution with existing academic and industrial approaches. The experiments show that our solution greatly outperforms its counterparts. For example, for the same workload, when Ripple is used, the average message delivery latency and network bandwidth consumption can be reduced by up to 77% and 93%, respectively. Shuping Ji, Wei Wang 0049, Jianguo Yao 0002, Hans-Arno Jacobsen |
Middleware | 5 |
| 2023 | High Performance and Power Efficient Accelerator for Cloud InferenceabstractFacing the growing complexity of Deep Neural Networks (DNNs), high-performance and power-efficient AI accelerators are desired to provide effective and affordable cloud inference services. We introduce our flagship product, i.e., the Cloudblazer i20 accelerator, which integrates the innovated Deep Thinking Unit (DTU 2.0). The design is driven by requests drawn from various AI inference applications and insights learned from our previous products. With careful tradeoffs in hardware-software co-design, Cloudblazer i20 delivers impressive performance and energy efficiency while maintaining acceptable hardware costs and software complexity/flexibility. To tackle computation- and data-intensive workloads, DTU 2.0 integrates powerful vector/matrix engines and a large-capacity multi-level memory hierarchy with high bandwidth. It supports comprehensive data flow and synchronization patterns to fully exploit parallelism in computation/memory access within or among concurrent tasks. Moreover, it enables sparse data compression/decompression, data broadcasting, repeated data transfer, and kernel code prefetching to optimize bandwidth utilization and reduce data access overheads. To utilize the underlying hardware and simplify the development of customized DNNs/operators, the software stack enables automatic optimizations (such as operator fusion and data flow tuning) and provides diverse programming interfaces for developers. Lastly, the energy consumption is optimized through dynamic power integrity and efficiency management, eliminating integrity risks and energy wastes. Based on the performance requirement, developers also can assign their workloads with the entire or partial hardware resources accordingly. Evaluated with 10 representative DNN models widely adopted in various domains, Cloudblazer i20 outperforms Nvidia T4 and A10 GPUs with a geometric mean of 2.22x and 1.16x in performance and 1.04x and 1.17x in energy efficiency, respectively. The improvements demonstrate the effectiveness of Cloudblazer i20’s design that emphasizes performance, efficiency, and flexibility. Jianguo Yao 0002, Hao Zhou 0009, Yalin Zhang 0004, Chuang Feng, Qiaojuan Hu |
HPCA | 1 |
| 2023 | LPNS: Scalable and Latency-Predictable Local Storage Virtualization for Unpredictable NVMe SSDs in Clouds
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
USENIX ATC | 3 |
| 2023 | FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance OptimizationsabstractWith the rapid development of cloud computing, numerous cloud services, containers, and virtual machines have been bringing tremendous demands on high-performance memory resources to modern data centers. Heterogeneous memory, especially the newly released Optane memory, offer appropriate alternatives against DRAM in clouds with the advantages of larger capacity, lower purchase cost, and promising performance. However, cloud services suffer serious implementation inconvenience and performance degradation when using hybrid DRAM and Optane memory. This article proposes FlexHM, a practical system to manage transparent heterogeneous memory resources and flexibly optimize memory access performance for all VMs, containers, and native applications. We present an open-source prototype of FlexHM in Linux with several main contributions. First, FlexHM raises a novel two-level NUMA design to manage DRAM and Optane memory as transparent main memory resources. Second, FlexHM provides flexible and efficient memory management, helping optimize memory access performance or save purchase costs of memory resources for differential cloud services with customized management strategies. Finally, the evaluations show that cloud workloads using 50% Optane slow memory on FlexHM can achieve up to 93% of the performance when using all-DRAM, and FlexHM provides up to 5.8× improvement over the previous heterogeneous memory system solution when workloads use the same ratio of DRAM and Optane memory. Bo Peng 0043, Yaozu Dong, Jianguo Yao 0002, Fengguang Wu, Haibing Guan |
ACM Trans. Archit. Code Optim. | 3 |
| 2023 | An Economy-Oriented GPU Virtualization With Dynamic and Adaptive OversubscriptionabstractGPU is becoming attractive around multiple academic and industrial area because of its massively parallel computing ability. However, there are still some obstacles which the GPU virtualization technologies should overcome to reach their maturity. These obstacles mainly include the problem of resource allocation strategy to guarantee possible higher yield. This shortage has already become an obvious barrier to the practical GPU usage in the cloud for satisfying business and academical requirements. There are many mature pieces of research in the area of oversubscribed cloud computing to enhance economic efficiency. However, the study on GPU oversubscription is almost blank for the just started use of GPU in cloud computing. This paper introduces gOver, an economy-oriented GPU resource oversubscription system based on the GPU virtualization platform. gOver is able to share and modulate GPU resource among workloads in an adaptive and dynamic manner, guaranteeing the QoS level at the same time. We evaluate the proposed gOver strategy with designed experiments with specific workload characteristics. The experimental results show that our dynamic GPU oversubscription solution improves the economic efficiency by 20% over traditional GPU sharing strategy, and outperforms the static oversubscription method by much better stability in QoS control. Jianguo Yao 0002, Qiumin Lu, Run Tian, Keqin Li 0001, Haibing Guan |
IEEE Trans. Computers | 1 |
| 2022 | Artificial Neural Networks for Downbeat Estimation and Varying Tempo Induction in Music Signals
Sarah Nadi, Jianguo Yao 0002 |
ICONIP (6) | 2 |
| 2022 | Learning to Optimize DAG Scheduling in Heterogeneous EnvironmentabstractScheduling job flows efficiently and rapidly on distributed computing clusters is one of huge challenges for daily operation of data centers. In a practical scenario, a single job consists of numerous stages with complex dependency relation represented as a Directed Acyclic Graph (DAG) structure. Nowa-days a data center usually equips with a cluster of heterogeneous computing servers which are different in the hardware/software configuration. From both the cost saving and environmental friendliness, the data centers could benefit a lot from optimizing the job scheduling problems in the heterogeneous environment. Thus the problem has attracted more and more attention from both the industry and academy. In this paper, we propose a task-duplication based learning algorithm, namely LACHESIS22The second of the Three Fates in ancient Greek mythology, who deter-mines destiny., aiming to optimize the problem. In the proposed approach, it first perceives the topological dependencies between jobs using a reinforcement learning framework and a specially designed graph neural network (GNN) to select the most promising task to be executed. Then the task is assigned to a specific executor with the consideration of duplicating all its precedent tasks according to an expert-designed rules. We have conducted extensive experiments over standard workloads to evaluate the proposed solution. The experimental results suggest that LACHESIS can achieve at most 26.7% reduction of makespan and 35.2% improvement of speedup ratio over seven strong baseline algorithms, including the state-of-the-art heuristics methods and a variety of deep reinforcement learning based algorithms. Yunfan Zhou, Xijun Li, Jinhong Luo, Mingxuan Yuan, Jianguo Yao 0002 |
MDM | 6 |
| 2022 | SWVM: a light-weighted virtualization platform based on Sunway CPU architecture
Jianguo Yao 0002, Qiumin Lu, Xingyan Wang, Hanyang Ma, Haibing Guan |
Sci. China Inf. Sci. | 1 |
| 2022 | MDev-NVMe: Mediated Pass-Through NVMe Virtualization Solution With Adaptive PollingabstractThe fast access to data and high parallel processing in high-performance computing instigates an urgent demand on the improvement of the NVMe storage within modern data centers. However, the former NVMe virtualization’s unsatisfactory performance demonstrates that NVMe devices are often underutilized within cloud platforms. An NVMe virtualization mechanism with high performance and device sharing has captured researchers and developers’ attention. This article introduces MDev-NVMe, a new virtualization solution for NVMe storage device with (1) full NVMe storage virtualization for VMs running native NVMe driver, (2) a mediated pass-through mechanism for NVMe management, and (3) adaptive configuration of active polling optimization to simultaneously achieve high throughput, low latency performance, and substantial device scalability. We practically implement the MDev-NVMe as a Linux kernel module. This article subsequently evaluates MDev-NVMe with Intel OPTANE and P3600 SSD by comparing several mainstream NVMe virtualization mechanisms using application-level I/O benchmarks. MDev-NVMe with active polling can demonstrate a 142 percent improvement over native (interrupt-driven) throughput and over 2.5 × theVirtiothroughput with only 70 percent native average latency and 31 percentVirtioaverage latency. Finally, the advantages of MDev-NVMe and the importance of adaptive polling are discussed, offering evidence that MDev-NVMe is a superior virtualization choice for cloud storage. Bo Peng 0043, Jianguo Yao 0002, Yaozu Dong, Haibing Guan |
IEEE Trans. Computers | 2 |
| 2022 | RESERVE: An Energy-Efficient Edge Cloud Architecture for Intelligent Multi-UAVabstractWith the increasing attraction of unmanned aerial vehicles (UAVs) in civil, public, and military applications, multi-UAV systems can perform environmental and disaster monitoring, border surveillance, and search and rescue. It is foreseen that these multi-UAV-based applications will be an important trend for edge computing scenarios. However, due to UAVs’ limited energy supplies as well as their continuous increase in the number of sensors, energy efficiency is a critical issue in multi-UAV systems. We believe that the fusion of edge computing and cloud computing can provide effective support for energy savings. This article presents an energy-efficient edge cloud architecture called RESERVE for intelligent multi-UAV. Under RESERVE, we study the energy-efficient computation offloading decision-making problem in a decentralized manner. The problem is formulated as a three-layer game in which the discretionary approach to reaching Nash Equilibrium is presented. Based on the proposed game, we design decentralized algorithms for two different cases. The algorithms can both achieve Nash Equilibrium. Furthermore, we propose a decentralized computation offloading mechanism and analyze the performance of the game by its efficiency ratio. We conduct simulation experiments and design a framework prototype. Evaluation results demonstrate that the proposed game methods can achieve more than 30 percent extra energy consumption reduction compared with the state-of-the-art decentralized algorithm and less than 10 percent performance loss relative to the centralized solution. The prototype framework we have developed proves the concept we propose. Beiqing Chen, Haihang Zhou, Jianguo Yao 0002, Haibing Guan |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Adaptive live migration of virtual machines under limited network bandwidthabstractLive migration is a crucial feature in existing virtualization platforms. Since memory is dirtied rapidly during the execution of a virtual machine (VM), boosting memory migration speed becomes a significant factor in guaranteeing a high-level success ratio and efficiency. However, the statically-configured migration strategy cannot cope with various workloads running in VMs, resulting in frequently aborted migration processes and low success ratio. This paper proposed a one-for-all migration architecture called Adaptive Live Migration (AdaMig) to address these issues. This QEMU-based solution dynamically switches migration methods and tunes related parameters by monitoring the run-time statistics from the migration process and the physical host. Once AdaMig detects the tendency that migration cannot converge, it will switch to another migration method to synchronize remaining dirty pages. During the whole process, AdaMig also dynamically tunes migration parameters according to current resources available in the physical host and migration efficiency. Experimental results reflect that AdaMig improves the success ratio from 26.7% to 93.3% over various workloads, and migration time is reduced by up to 45.5% in comparison with the original solution in QEMU. Handong Li, Guangrong Xiao, Qiumin Lu, Jianguo Yao 0002 |
VEE | 6 |
| 2021 | A Throughput-Oriented NVMe Storage Virtualization With Workload-Aware ManagementabstractStorage virtualization is an important component of large-scale online services in multi-tenant clouds. It typically shares the physical storage among guest machines and performs transactional operations for high-performance data processing. However, even with the recent mediated pass-through virtualization optimization, the operations of multi-tenant storage I/O meet the bottleneck, and thus degrade the throughput performance of the cloud storage services. We observe that the root cause of the problem is the unawareness of varying and imbalanced workload inefficiency of resource management in the multi-tenant cloud storage setting. In this paper, we present FinNVMe, a new throughput-oriented NVMe storage virtualization management mechanism, that (1) passes-through I/O performance-critical resources and emulates privileged resources to provide high throughput in a workload-aware manner among multi-tenant VMs, (2) enables fine-grained scheduling for I/O resources to achieve promising flexibility and scalability with respective to virtualization, and (3) adopts the queue binding and the queue shuffling to reduce the virtualization and management overhead, and involves active polling for further I/O acceleration. This article subsequently evaluates FinNVMe with micro benchmarks on two typical scenarios (both balanced and imbalanced workload) and the real-world storage workloads to show its high throughput performance, along with the flexibility and scalability of virtualization and resource management. For example, FinNVMe achieves up to 20 percent throughput improvement with more stable latency in the varying and imbalanced workload. Bo Peng 0043, Ming Yang 0021, Jianguo Yao 0002, Haibing Guan |
IEEE Trans. Computers | 3 |
| 2021 | Enabling Cloud Applications to Negotiate Multiple Resources in a Cost-Efficient MannerabstractCloud applications can achieve similar performance with diverse multi-resource configurations, allowing cloud service providers to benefit from optimal resource allocation for reducing their operation cost. This paper aims to solve the problem of multi-resource negotiation with considerations of both the service-level agreement (SLA) and the cost efficiency so that the performance requirement for cloud services is satisfied and the cost of resource usage is also minimized. The performance and resource demand are usually application-dependent, making the optimization problem complicated, especially when the dimension of multi-resource configurations is large. To this end, we use reinforcement learning to solve the optimal problem of multi-resource configuration with simultaneous optimization of the learning efficiency and performance guarantee. The developed prototype named SmartYARN is an extended Apache YARN equipped with our learning algorithm which can enable cloud applications to negotiate multiple resources cost-effectively. The extensive evaluations with four typical benchmarks show that SmartYARN performs well in reducing the cost of resource usage while maintaining compliance with the SLA constraints of cloud service simultaneously. Jianguo Yao 0002, Hans-Arno Jacobsen, Haibing Guan |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | gQoS: A QoS-Oriented GPU Virtualization with Adaptive Capacity SharingabstractCurrently, the virtualization technologies for cloud computing infrastructures supporting extra devices, such as GPU, require additional development and refinement. This requirement is particularly evident in the area of resource sharing and allocation under some performance constraints, like the quality of service (QoS) guarantee, in light of the closed GPU platform. This deficiency significantly limits the applicability range of the cloud platform, which aims to support the efficient and fluent execution of business and academic workloads. This paper introduces gQoS, an adaptive virtualized GPU resource capacity sharing system under the QoS target, which can share and allocate the virtualized GPU resource among workloads adaptively, guaranteeing the QoS level with stability and accuracy. We evaluate the workloads and compare our gQoS strategy with other allocation strategies. The experiments show that our strategy guarantees much better accuracy and stability in QoS control and that the total GPU resource utilization under gQoS can be rewarded with at most a 25.85 percent reduction compared with other strategies. Qiumin Lu, Jianguo Yao 0002, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | gMig: Efficient vGPU Live Migration with Overlapped Software-Based Dirty Page VerificationabstractThis paper introduces gMig, an open-source and practical vGPU live migration solution for full virtualization. Taking the advantage of the dirty pattern of GPU workloads, gMig presents the One-Shot Pre-Copy mechanism combined with the hashing based Software Dirty Page technique to achieve efficient vGPU live migration. Particularly, we propose three core techniques for gMig: 1) Dynamic Graphics Address Remapping, which parses and manipulates GPU commands to adjust the address mapping and adapt to a different environment after migration, 2) Software Dirty Page, which utilizes a hashing based approach with sampling pre-filtering to detect page modification, overcomes the commodity GPU's hardware limitation, and speeds up the migration by only sending the dirtied pages, 3) Overlapped Migration Process, which significantly compresses the hanging overhead by overlapping the dirty page verification and transmission concurrently. Our evaluation shows that gMig achieves GPU live migration with an average downtime of 302 ms on Windows and 119 ms on Linux. With the help of Software Dirty Page, the number of GPU pages transferred during the downtime is effectively reduced by up to 80.0 percent . The design of sampling filter and overlapped processing can bring about further 30.0 and 10.0 percent improvements in page processing. Qiumin Lu, Jiacheng Ma 0001, Yaozu Dong, Zhengwei Qi, Jianguo Yao 0002, Bingsheng He, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2019 | Proactive coordination for low-congestion multi-path datacenter networks
Bo Peng 0043, Jianguo Yao 0002, Haibing Guan |
J. Syst. Archit. | 2 |
| 2019 | Fairness-Efficiency Allocation of CPU-GPU Heterogeneous ResourcesabstractConsidering the performance improvement the cloud technology provides by processing workloads in parallel, applications and services are now migrating to online clouds. In a cloud platform, workloads can be executed in a virtualized environment to have a great improvement of the resource utilization. However, there is a new challenge in the allocation problem, which is quantifying and optimizing the fairness and efficiency of heterogeneous resources (CPUs and GPUs) required by applications such as cloud gaming. The solving approach needs scalarization methods of the requirement vector, relevant functions for fairness metrics, and an acceptable algorithm to solve that, where the difficulties mainly locate. We design an iterative, dynamic-adaptive heuristic solving algorithm Fairness-Efficiency Allocation (FEA) and optimize the implementation on a virtualized platform, which collects runtime data, allocates resources and reports differences. Data are recorded and analyzed to discover the effect of the allocation in different situations, including the promotion of fairness and the effect on the frame rate of the workloads. The result indicates that there is a considerable fairness improvement after the resource allocation, especially in situations that many virtual machines are executing simultaneously. Compared with the VGASA strategy, the fairness metric value improved 45 percent in three virtual machines' situation. Qiumin Lu, Jianguo Yao 0002, Zhengwei Qi, Bingsheng He, Haibing Guan |
IEEE Trans. Serv. Comput. | 2 |
| 2019 | HyperCo: Optimizing Network Performance in ARM-Based Mobile VirtualizationabstractIn the ARM-based mobile virtualized environment, optimizing both the network throughput and the latency for network-intensive applications is of great importance. Our experimental studies have shown that the context switch between the host and guest triggered by a hypercall results in CPU overload and thus degrades the I/O performance under an intensive workload. To address this problem, we propose the Adaptive Hypercall Coalescing (HyperCo) algorithm, a software-only approach to optimizing the network I/O performance by reducing the number of hypercalls and achieving a trade-off between throughput and latency. We implement HyperCo by modifying the front-end driver of Virtio-net on the KVM/ARM platform, and we carry out extensive performance evaluations, which indicate that the number of hypercalls and the load on the CPU can be significantly reduced. We show that HyperCo can significantly improve the network throughput under an intensive workload while keeping a low penalty by adapting a coalescing interval dynamically at a low frequency of network I/O requests. Jianguo Yao 0002, Ting Deng, Xue (Steve) Liu, Hans-Arno Jacobsen, Haibing Guan |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | A Two-Layer Algorithmic Framework for Service Provider Configuration and Planning with Optimal Spatial MatchingabstractIndustrial telecommunication applications prefer to run at the optimal capacity configuration to achieve the required Quality of Service (QoS) at the minimum cost. The optimal capacity configuration is usually achieved through the selection of cell towers capacities and locations. Given a set of service providers (e.g., cell towers) and a set of customers (e.g., major residential areas), where each customer has an amount of demand and each provider has multiple candidate capacities and corresponding costs, the optimal capacity selection is configured through spatial matching to satisfy the demand of each customer at the minimum cost. However, existing solutions developed for spatial matching, in which each provider's capacity is fixed, cannot be directly applied to the capacity configuration problem with multiple capacities and location selections. In this paper, we are the first to study Service Provider Configuration and Planning with Optimal Spatial Matching (SPC-POSM) problem, in which the objectives are 1) to select the proper capacity for each provider at the minimum total cost and 2) to assign providers' service to satisfy the demand of each customer on a condition that the matching distance is no more than service quality requirement. We prove that SPC-POSM problem is NP-hard and design an efficient two-layer meta-heuristic framework to solve the problem. Unsupervised learning technique is designed to accelerate the calculation and a novel local search mechanism is embedded to further improve solution quality. Extensive experimental results on both real and synthetic datasets verify the effectiveness and efficiency of the proposed framework. Xijun Li, Jianguo Yao 0002, Mingxuan Yuan |
CIKM | 2 |
| 2018 | Qualitative Instead of Quantitative: Towards Practical Data Analysis Under Differential Privacy
Xuanyu Bai, Jianguo Yao 0002, Mingxuan Yuan, Haibing Guan |
DASFAA (2) | 2 |
| 2018 | Efficient Sharing and Fine-Grained Scheduling of Virtualized GPU ResourcesabstractGraphics Processing Unit (GPU) provides acceleration services to many applications, such as AI, games, media transcoding, etc. Virtualization on GPU is an enabling technology which facilitates the hardware resource sharing among multiple virtual machines (VMs). Sharing a GPU not only brings pros such as high utilization but also introduces cons such as resource contention and performance degradation. Although the existing GPU scheduling policies have been to some extent optimized, there are still some deficiencies, such as inefficient GPU sharing among multiple VMs, and high overhead within VM switching. As a result, the performance of GPU virtualization is limited by the current design, which lacks fine-grained scheduling supports. In this paper, we propose the Fine-grained schEduLing of vIrtualized gPu rEsources (FELIPE) to fully utilize and efficiently share a physical GPU among multiple VMs. To this end, we achieve the FELIPE optimization by introducing fine-grained scheduling mechanisms for virtualized GPU resources in three aspects: 1) We design a mixed time/event-based scheduling policy to reduce the idle time within VM switching. 2) We create a seamless VM assignment process, which enables VMs to switch seamlessly by stages. 3) We develop a hybrid per-ring/VM scheduling strategy, which schedules workloads to different GPU engines to run simultaneously. Then we implement the FELIPE with Intel Graphics Virtualization Technology for shared vGPU technology (GVT-g). Finally, the experimental evaluations show that the performance of the first two scheduling policies can respectively achieve up to 21.5% and 19.7% improvement, and the last one can improve the performance from 57.9% to 98.5% compared with the native design for two virtual machines. Jianguo Yao 0002, Haibing Guan |
ICDCS | 2 |
| 2018 | Maximizing Profit of Cloud Service Brokerage with Economic Demand ResponseabstractCloud service brokerage (CSB), which procures cloud services from multiple cloud service providers (CSPs) and resells them to cloud customers, has been put forward to facilitate the delivery of cloud services. However, it is challenging to address the economic issues of CSB incurred by insufficient provisioning problem in response to dynamic conditions, for example, dynamic customer demands, dynamic cloud service prices and different availabilities of CSPs. In this paper, we propose a novel mechanism called CSB Demand Response (DR-CSB), which aims to maximize the profit of CSB under dynamic customer demands with respect to the capacity and availability constraints, to mitigate the insufficient provisioning problem. To this end, we formulate an optimization problem of profit maximization for the CSB, and employ economic demand response mechanism to allow cloud customers to adjust their consumptions with dynamic cloud service prices. Our evaluations driven by Google cluster-usage traces have verified that the DR-CSB not only can help the CSB to achieve the profit maximization, but also can handle the impact of the dynamic conditions in CSB. As the result shows, the profit of CSB with implementing DR-CSB can increase by up to 20%, and customers also achieve a 37% aggregated cost saving, compared with the scenario without DR-CSB. Ting Deng, Jianguo Yao 0002, Haibing Guan |
INFOCOM | 2 |
| 2018 | HybridPass: Hybrid Scheduling for Mixed Flows in Datacenter NetworksabstractModern cloud applications generate millions of mixed flows transmitted between distributed nodes, and the typical latency-sensitive and throughput-intensive flows coexist in datacenter networks. Scheduling those mixed flows presents new challenges when meeting both low latency and high throughput requirements. This paper introduces HybridPass, a novel hybrid network architecture for datacenter networks, which is the first attempt to support both time-triggered and event-triggered scheduling in respect to latency-sensitive and throughput-intensive flows. To this end, we develop an arbiter, which uses a loosely synchronized time-triggered manner to allocate the network bandwidth for latency-sensitive and throughput-intensive flows from a global perspective. Specifically, the time-triggered scheduling aims to minimize the latency through establishing flow-level and task-level models. Then the event-triggered scheduling is developed to utilize the leftover bandwidth for throughput-intensive flows without any impact on latency-sensitive flows. Our experiments show that HybridPass can achieve up to 40.71% latency reduction for latency-sensitive flows compared with the baseline DCTCP while maintaining the high throughput with inappreciable 0.77% throughput sacrifice for throughput-intensive flows. Bo Peng 0043, Jianguo Yao 0002, Zhengwei Qi, Haibing Guan |
IPDPS | 2 |
| 2018 | A Data-Driven Three-Layer Algorithm for Split Delivery Vehicle Routing Problem with 3D Container Loading ConstraintabstractSplit Delivery Vehicle Routing Problem with 3D Loading Constraints (3L-SDVRP) can be seen as the most important problem in large-scale manufacturing logistics. The goal is to devise a strategy consisting of three NP-hard planning components: vehicle routing, cargo splitting and container loading, which shall be jointly optimized for cost savings. The problem is an enhanced variant of the classical logistics problem 3L-CVRP, and its complexity leaps beyond current studies of solvability. Our solution employs a novel data-driven three-layer search algorithm (DTSA), which we designed to improve both the efficiency and effectiveness of traditional meta-heuristic approaches, through learning from data and from simulation. Xijun Li, Mingxuan Yuan, Jianguo Yao 0002 |
KDD | 4 |
| 2018 | MDev-NVMe: A NVMe Storage Virtualization Solution with Mediated Pass-Through
Bo Peng 0043, Haozhong Zhang, Jianguo Yao 0002, Yaozu Dong, Haibing Guan |
USENIX ATC | 3 |
| 2018 | Demon: An Efficient Solution for on-Device MMU Virtualization in Mediated Pass-ThroughabstractMemory Management Units (MMUs) for on-device address translation are widely used in modern devices. However, conventional solutions for on-device MMU virtualization, such as shadow page table implemented in mediated pass-through, still suffer from high complexity and low performance. Jianguo Yao 0002, Yaozu Dong, Haibing Guan |
VEE | 2 |
| 2018 | D2FL: Design and Implementation of Distributed Dynamic Fault LocalizationabstractCompromised or misconfigured routers have been a major concern in large-scale networks. Such routers sabotage packet delivery, and thus hurt network performance. Data-plane fault localization (FL) promises to solve this problem. Regrettably, the path-based FL fails to support dynamic routing, and the neighbor-based FL requires a centralized trusted administrative controller (AC) or global clock synchronization in each router and introduces storage overhead for caching packets. To address these problems, we introduce a dynamic distributed and low-cost model, D2FL. Using random two-hop neighborhood authentication, D2FL supports volatile path without the AC or global clock synchronization. Besides, D2FL requires only constant tens of KB for caching which is independent of the packet transmission rate. This is much less than the cache size of DynaFL or DFL which consumes several MB. The simulations show that D2FL achieves low false positive and false negative rate with no more than 3 percent bandwidth overhead. We also implement an open source prototype and evaluate its effect. The result shows that the performance burden in user space is less than 10 percent with the dynamic sampling algorithm. Fanfu Zhou, Zhengwei Qi, Jianguo Yao 0002, Ruhui Ma, Bin Wang 0062, Athanasios V. Vasilakos, Haibing Guan |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2018 | Power Peak Shaving With Data Transmission Delays for Thermal Management in Smart BuildingsabstractThis paper presents a scheme aimed at mitigating the influence of random data transmission delays in networked thermal appliance control systems in smart buildings. The impact of this type of delays is first analyzed, and it is proposed to utilize loose timing synchronization and add blank gaps between the consecutive appliance operations to avoid the possible violation of the given power budget. A cooperative control of thermal appliance operation is developed using a networked model predictive control-based controller to deal with delays. It is also shown that the schedulability of such a control scheme can be assessed online. The performance of the proposed control scheme is assessed by a simulation study based on the thermal dynamics of an eight-room office building. The obtained results show that the proposed solution can achieve an efficient power peaks shaving in the presence of random network delays. Jianguo Yao 0002, Guchuan Zhu |
IEEE Trans. Ind. Informatics | 3 |
| 2017 | MobiXen: Porting Xen on Android devices for mobile virtualizationabstractThe mobile virtualization technology provides a feasible way to improve the manageability and security for embedded systems. This paper presents an architecture named MobiXen to address these challenges. In the MobiXen, both Xen's physical memory space and virtual address space are shrunk as much as possible and thus Android owns more memory resource; optimizations are developed to reduce the virtualization overhead when Android is accessing system resources; new policies are implemented to achieve low suspend/resume latency. With these work adopted, MobiXen is customized as a high efficient mobile hypervisor. Detailed implementations shows that, most of the performance degradation brought by MobiXen is less than 3%, which is imperceptible by end users. Yaozu Dong, Jianguo Yao 0002, Haibing Guan, R. Ananth Krishna, Yunhong Jiang |
DATE | 2 |
| 2017 | Traffic Prediction Based Power Saving in Cellular Networks: A Machine Learning MethodabstractIn smart cities, green cellular networks play a crucial role to support wireless access for numerous devices anywhere and anytime with efficiency and sustainability. Because base stations (BSes) consume more than 70% of overall cellular network infrastructure energy, saving the power consumption of BSes is the key task to build a green cellular network. Except for low power design of the BS hardware and software, the traffic-driven BS sleeping operation is an economical way to improve existing cellular networks, which can reduce the BS power consumption at low traffic load. However, prior BS sleeping strategies establish on the static temporal characteristics of traffic load, which ignore the fact that network traffic is influenced by many factors such as time, human mobility, holiday, weather, etc. Hence, prior traffic estimation is coarse, and the BS sleeping strategies cannot apply to the changing network traffic. In this paper, we exploit a machine learning method to estimate the BS traffic and propose a BS sleeping strategy based on predicted traffic for power saving in the cellular network. We analyze network traffic in multi-views: temporal influence, spatial influence, and event influence. Then, we propose a multi-view ensemble learning model to predict network traffic load, which learns the traffic in multi-views and combine the results with ensemble. Furthermore, we formulate a BS sleeping strategy based on the predicted traffic load. Finally, we evaluate our traffic prediction algorithm on real cellular network data. The evaluation shows that our traffic prediction algorithm improves about 40% than state-of-the-art machine learning methods. Also, we evaluate the proposed BS sleeping strategy, which yields about 10% more energy savings and less device damage than the competitors in the simulated environment. Shenglin Zhao, Mingxuan Yuan, Jianguo Yao 0002, Michael R. Lyu, Irwin King |
SIGSPATIAL/GIS | 5 |
| 2017 | A First Look at Information Entropy-Based Data PricingabstractDistribution of intangible information goods is experiencing tremendous growth in recent years, which has facilitated a blossoming of information goods economics. As big data develops, there are more and more information goods markets for data trading. In the current of data pricing policies in data trading, there are many metrics to measure the value of data goods, such as the data generation date, data volume, and data integrity, etc. However, it is very challenging to identify the amount of data information and its distribution, and the corresponding data pricing has rarely been discussed. In this paper, we propose a new data pricing metric, i.e., the data information entropy, which helps to make a reasonable price in the data trading. We first demonstrate a data information measurement method based on information entropy, and then propose a pricing function based on the result of data information measurement. To comprehensively understand the new data pricing metric and facilitate its application in data trading, we verify the rationality of the data information measurement method and give three concrete pricing functions. It is the first time to look at the information entropy-based data pricing, which can inspire the research concerning the pricing mechanism of data goods, further promoting the development of data products business. Xijun Li, Jianguo Yao 0002, Xue (Steve) Liu, Haibing Guan |
ICDCS | 2 |
| 2017 | Cost-efficient negotiation over multiple resources with reinforcement learningabstractCloud applications can achieve similar performance with diverse multi-resource configurations, allowing cloud service providers to benefit from optimal resource allocation for reducing their operation cost. This paper aims to solve the problem of multi-resource negotiation with considerations of both the service-level agreement (SLA) and the cost efficiency. The performance and resource demand are usually application-dependent, making the optimization problem complicated, especially when the dimension of multiresource configuration is large. To this end, we use reinforcement learning to solve the optimization problem of multi-resource configuration with simultaneous optimization of the learning efficiency and performance guarantee. The developed prototype named SmartYARN is extended Apache YARN equipped with our learning algorithm which can enable cloud applications to negotiate multiple resources cost-effectively. The extensive evaluations show that SmartYARN performs well in reducing the cost of resource usage while maintaining compliance with the SLA constraints of cloud service simultaneously. Jianguo Yao 0002, Hans-Arno Jacobsen, Haibing Guan |
IWQoS | 2 |
| 2017 | Robust Multi-Resource Allocation with Demand Uncertainties in Cloud SchedulerabstractCloud scheduler manages multi-resources (e.g., CPU, GPU, memory, storage etc.) in cloud platform to improve resource utilization and achieve cost-efficiency for cloud providers. The optimal allocation for multi-resources has become a key technique in cloud computing and attracted more and more researchers' attentions. The existing multi-resource allocation methods are developed based on a condition that the job has constant demands for multi-resources. However, these methods may not apply in a real cloud scheduler due to the dynamic resource demands in jobs' execution. In this paper, we study a robust multi-resource allocation problem with uncertainties brought by varying resource demands. To this end, the cost function is chosen as either of two multi-resource efficiency-fairness metrics called Fairness on Dominant Shares and Generalized Fairness on Jobs, and we model the resource demand uncertainties through three typical models, i.e., scenario demand uncertainty, box demand uncertainty and ellipsoidal demand uncertainty. By solving an optimization problem we get the solution for robust multi-resource allocation with uncertainties for cloud scheduler. The extensive simulations show that the proposed approach can handle the resource demand uncertainties and the cloud scheduler runs in an optimized and robust manner. Jianguo Yao 0002, Qiumin Lu, Hans-Arno Jacobsen, Haibing Guan |
SRDS | 1 |
| 2017 | Automated Resource Sharing for Virtualized GPU with Self-ConfigurationabstractIn this paper, we propose Auto-vGPU, a framework of automated resource sharing for virtualized GPU with self-configuration, to reduce manual intervention in system management while ensuring Service Level Agreement (SLA) targets. Auto-vGPU automatically collects the measurements of system metrics and learns a linear model for each application with dimension reduction. In order to fulfill the automated configuration of controller parameters, we propose a self-control-configuration method featuring the theory of automatic tuning of proportional-integral (PI) regulators. The experimental results of cloud gaming implementation demonstrate that Auto-vGPU is able to automatically build the low-dimension model and configure the control parameters without any manual interventions and the derived controller can adaptively allocate virtualized GPU resource to ensure the high performance of cloud applications. Jianguo Yao 0002, Qiumin Lu, Zhengwei Qi |
SRDS | 1 |
| 2017 | Embedding differential privacy in decision tree algorithm with different depths
Xuanyu Bai, Jianguo Yao 0002, Mingxuan Yuan, Xike Xie, Haibing Guan |
Sci. China Inf. Sci. | 2 |
| 2016 | A user mode CPU-GPU scheduling framework for hybrid workloads
Bin Wang 0062, Ruhui Ma, Zhengwei Qi, Jianguo Yao 0002, Haibing Guan |
Future Gener. Comput. Syst. | 4 |
| 2016 | MixCPS: Mixed Time/Event-Triggered Architecture of Cyber-Physical SystemsabstractCyber–physical system (CPS) integrates variety of applications where some work in a time-triggered manner while others work in an event-triggered manner. Supporting unified designs of both time-triggered and event-triggered tasks facilitates the integration of CPS applications. This paper presents a mixed time/event-triggered architecture called MixCPS to make CPS applications work in the integrated and optimized manner. To optimize MixCPS, we propose a new timing performance metric for CPS applications, i.e., the application-level delay, which is induced from sensor to actuator and captures the coupling between cyber and physical worlds. Using the system-level time-triggered scheduling based on the synchronization between computation and communication nodes, we optimize the assignment of computation tasks and the packet transmissions in order to minimize the application-level delays. Furthermore, we discuss the scheduling policy for event-trigged tasks after the resource reservation for time-trigged tasks. Finally, the proposed MixCPS is evaluated by several simulations which show that the application-level delays can be significantly reduced through the time-triggered cooperation mechanism of computation and communication. Jianguo Yao 0002, Xue (Steve) Liu |
Proc. IEEE | 1 |
| 2015 | Comprehensive understanding of operation cost reduction using energy storage for IDCsabstractTo reduce the operation cost incurred by the rapidly growing energy consumption in internet data centers (IDCs), more and more internet service providers have spontaneously started using energy storage in various forms. The approach of energy storage is used to store cheap electricity energy when the electricity price from smart gird is low or the renewable energy is used. There are two typical forms of energy storage equipped in many IDCs, i.e., the battery storage and thermal energy storage. Recent work shows the energy storage can significantly reduce the operation cost for IDCs. However, the cost of the energy storage devices are still at a high level, and it may increase the operation cost for IDCs. In this paper, we investigate the comprehensive understanding on the operation cost reduction for IDCs using the energy storage. To this end, we conduct a quantitative analysis on the normalized electricity price in the two energy storage forms. The experiments demonstrate that the cost of the storage devices and renewable energy supply are largely affected by the storage capacity and the location of data centers, and we conclude that it does not always reduce operation cost using energy storage for IDCs. Haihang Zhou, Jianguo Yao 0002, Haibing Guan, Xue (Steve) Liu |
INFOCOM | 2 |
| 2015 | End-to-end delay analysis for networked systemsabstractEnd-to-end delay measurement has been an essential element in the deployment of real-time services in networked systems. Traditional methods of delay measurement based on time domain analysis, however, are not efficient as the network scale and the complexity increase. We propose a novel theoretical framework to analyze the end-to-end delay distributions of networked systems from the frequency domain. We use a signal flow graph to model the delay distribution of a networked system and prove that the end-to-end delay distribution is indeed the inverse Laplace transform of the transfer function of the signal flow graph. Two efficient methods, Cramer’s rule-based method and the Mason gain rule-based method, are adopted to obtain the transfer function. By analyzing the time responses of the transfer function, we obtain the end-to-end delay distribution. Based on our framework, we propose an efficient method using the dominant poles of the transfer function to work out the bottleneck links of the network. Moreover, we use the framework to study the network protocol performance. Theoretical analysis and extensive evaluations show the effectiveness of the proposed approach. Jie Shen 0011, Wenbo He 0003, Xue (Steve) Liu, Zhibo Wang 0001, Zhi Wang 0003, Jianguo Yao 0002 |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2015 | Differential Privacy in Telco Big Data PlatformabstractDifferential privacy (DP) has been widely explored in academia recently but less so in industry possibly due to its strong privacy guarantee. This paper makes the first attempt to implement three basic DP architectures in the deployed telecommunication (telco) big data platform for data mining applications. We find that all DP architectures have less than 5% loss of prediction accuracy when the weak privacy guarantee is adopted (e.g., privacy budget parameter ε ≥ 3). However, when the strong privacy guarantee is assumed (e.g., privacy budget parameter ε ≤ 0:1), all DP architectures lead to 15% ~ 30% accuracy loss, which implies that real-word industrial data mining systems cannot work well under such a strong privacy guarantee recommended by previous research works. Among the three basic DP architectures, the Hybridized DM (Data Mining) and DB (Database) architecture performs the best because of its complicated privacy protection design for the specific data mining algorithm. Through extensive experiments on big data, we also observe that the accuracy loss increases by increasing the variety of features, but decreases by increasing the volume of training data. Therefore, to make DP practically usable in large-scale industrial systems, our observations suggest that we may explore three possible research directions in future: (1) Relaxing the privacy guarantee (e.g., increasing privacy budget ε) and studying its effectiveness on specific industrial applications; (2) Designing specific privacy scheme for specific data mining algorithms; and (3) Using large volume of data but with low variety for training the classification models. Xueyang Hu, Mingxuan Yuan, Jianguo Yao 0002, Lei Chen 0002, Qiang Yang 0001, Haibing Guan |
Proc. VLDB Endow. | 3 |
| 2015 | Energy-Efficient SLA Guarantees for Virtualized GPU in Cloud GamingabstractBoth power consumption and SLA guarantees are important concerns for cloud gaming. Recently, various approaches have been developed to effectively reduce GPU power consumption by making GPU run at low frequencies. However, virtual machines (VMs) running on the same physical GPU with virtualization technology are correlated, because the change of GPU frequencies will affect the SLA performance of all the VMs. In fact, both reducing power consumption and guaranteeing SLA should work together under the considerations of their correlations. This paper proposes a novel two-layer control architecture called energy-efficient SLA guarantees for virtualized GPU (EvGPU) based on well-established feedback control techniques. The first control loop adopts a proportional-integral (PI) controller to ensure SLA guarantees, which in particular are measured at a predefined level of the frames per second (FPS) for each online game. The secondary power control loop then adjusts GPU frequency through dynamic voltage/frequency scaling (DVFS) to reduce power consumption based on the current FPS achieved by the first loop. Empirical results demonstrate that the proposed solution can effectively reduce GPU power consumption, while achieving the required SLA performance in virtualized GPU for cloud gaming. Haibing Guan, Jianguo Yao 0002, Zhengwei Qi |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Adaptive Power Management through Thermal Aware Workload Balancing in Internet Data CentersabstractThe past decade witnessed the tremendous growth of online services and applications. Together with the increase of cloud computing, more and more computation are hosted by Internet data centers (IDCs). Today's IDCs are achieving significant advances in communication and computation capabilities. However, along with the increasing demand from IDC clients, power consumption for powering up and cooling these IDCs has been skyrocketing. Most existing works optimize the power consumption of either servers or Computer Room Air Conditioners (CRACs), and overlook the correlation between the power consumption of these two types of equipment. In this paper, we propose an adaptive power control method which leverages the correlation between the power consumption of servers and CRACs. To capture the workload uncertainties and thermal dynamics, we exploit Recursive-Least Square based Model Predictive Control (MPC) to solve the power control problem. Performance evaluations shows the effective power peak reduction using our approach. Jianguo Yao 0002, Haibing Guan, Jianying Luo, Lei Rao, Xue (Steve) Liu |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | COMIC: Cost Optimization for Internet Content MultihomingabstractContent service is a type of Internet cloud service that provides end-users plentiful contents. To ensure high performance for content delivering, content service utilizes a technology known as content multihoming: contents are generated from multiple geographically distributed data centers and delivered by multiple distributed content distribution networks (CDNs). The electricity costs for data centers and the usage costs for CDNs are major contributors to the contents service cost. As electricity prices vary across data centers and usage costs vary across CDNs, scheduling data centers and CDNs has a tremendous consequence for optimizing content service cost. In this paper, we propose a novel framework named Cost Optimization for Internet Content Multihoming (COMIC). COMIC dynamically balances end-users' loads among data centers and CDNs so as to minimize the content service cost. Using real-life electricity prices and CDN traces, the experiments demonstrate that COMIC effectively reduces the content service cost by more than 20 percent. Jianguo Yao 0002, Haihang Zhou, Jianying Luo, Xue (Steve) Liu, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Control of Large-Scale Systems through Dimension ReductionabstractAutomated physical resource management of large-scale Internet Technology (IT) systems requires dynamic configuration of both application-level and system-level parameters. The existence of large number of tunable parameters makes it difficult to design a feedback controller that adjusts these parameters effectively in order to achieve application-level performance targets. In this paper, we introduce a new approach for simplified control architecture of large-scale IT systems based on dimension reduction techniques. It combines online selection of critical control knobs through LASSO-a powerful$L_1$-constrained fitting method/Compressive Sensing (CS)-a$L_1$-optimization method, and adaptive control of the identified knobs. The latter relies on the online estimation of the input-output model with the selected control knobs using the recursive least square (RLS) method and a self-tuning linear quadratic (LQ) optimal controller for output regulation. The results of both a numerical simulation in Matlab and a realistic case are presented to demonstrate the effectiveness of our approach. Jianguo Yao 0002, Xue (Steve) Liu, Xiaoyun Zhu, Haibing Guan |
IEEE Trans. Serv. Comput. | 1 |
| 2014 | VGRIS: Virtualized GPU Resource Isolation and Scheduling in Cloud GamingabstractTo achieve efficient resource management on a graphics processing unit (GPU), there is a demand to develop a framework for scheduling virtualized resources in cloud gaming. In this article, we propose VGRIS, a resource management framework for virtualized GPU resource isolation and scheduling in cloud gaming. A set of application programming interfaces (APIs) is provided so that a variety of scheduling algorithms can be implemented within the framework without modifying the framework itself. Three scheduling algorithms are implemented by the APIs within VGRIS. Experimental results show that VGRIS can effectively schedule GPU resources among various workloads. Zhengwei Qi, Jianguo Yao 0002, Chao Zhang 0115, Zhizhou Yang, Haibing Guan |
ACM Trans. Archit. Code Optim. | 2 |
| 2014 | vGASA: Adaptive Scheduling Algorithm of Virtualized GPU Resource in Cloud GamingabstractAs the virtualization technology for GPUs matures, cloud gaming has become an emerging application among cloud services. In addition to the poor default mechanisms of GPU resource sharing, the performance of cloud games is inevitably undermined by various runtime uncertainties such as rendering complex game scenarios. The question of how to handle the runtime uncertainties for GPU resource sharing remains unanswered. To address this challenge, we propose vGASA, a virtualized GPU resource adaptive scheduling algorithm in cloud gaming. vGASA interposes scheduling algorithms in the graphics API of the operating system, and hence the host graphic driver or the guest operating system remains unmodified. To fulfill the service level agreement as well as maximize GPU usage, we propose three adaptive scheduling algorithms featuring feedback control that mitigates the impact of the runtime uncertainties on the system performance. The experimental results demonstrate that vGASA is able to maintain frames per second of various workloads at the desired level with the performance overhead limited to 5-12 percent. Chao Zhang 0115, Jianguo Yao 0002, Zhengwei Qi, Haibing Guan |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2013 | VGRIS: virtualized GPU resource isolation and scheduling in cloud gaming
Chao Zhang 0115, Zhengwei Qi, Jianguo Yao 0002, Yin Wang 0001, Haibing Guan |
HPDC | 4 |
| 2013 | A Better Understanding of Event-Triggered Control from a CPS PerspectiveabstractIn cybernetics filed, the benefit of event-triggered control reaches a consensus, but in practice the event-triggered control is hampered by the lack of a system theory. This paper provides a Cyber-Physical System (CPS) viewpoint of the relation between the controller and the physical plant to explain the event detection logic of the event-triggered control. The relation between the controller and the physical plant is considered as communication, and based on this communication process the controller and the physical plant exchange the information entropy and thermodynamic entropy. The physical plant generates the thermodynamic entropy and the controller provides the negative thermodynamic entropy based on the predictability of the mathematical model for the physical plant, where there is an entropy balance process between the thermodynamic entropy and information entropy. The adjustment of the controller cycle period is the process to adjust the entropy balance process, which implies the event detection logic. Jie An 0003, Jianguo Yao 0002, Haihang Zhou |
ICPADS | 2 |
| 2013 | NetSimplex: Controller Fault Tolerance Architecture in Networked Control SystemsabstractThe assurance of reliability becomes increasingly challenging as the complexity of networked control systems (NCS) rapidly increases. Simplex architecture was designed to tolerate control software design and implementation. This architecture consists of a high assurance controller (HAC) and a high performance controller (HPC). The HAC uses the linear state feedback control to create a large maximum stability region (MSR). The HPC aims at achieving a better control performance and may use any design. However, the plant's states under HPC must stay within the MSR, or the control is switched to HAC. Jianguo Yao 0002, Xue (Steve) Liu, Guchuan Zhu, Lui Sha |
IEEE Trans. Ind. Informatics | 1 |
| 2013 | System-level calibration for data fusion in wireless sensor networksabstractWireless sensor networks are typically composed of low-cost sensors that are deeply integrated in physical environments. As a result, the sensing performance of a wireless sensor network is inevitably undermined by biases in imperfect sensor hardware and the noises in data measurements. Although a variety of calibration methods have been proposed to address these issues, they often adopt the device-level approach that becomes intractable for moderate-to large-scale networks. In this article, we propose a two-tier system-level calibration approach for a class of sensor networks that employ data fusion to improve the sensing performance. In the first tier of our calibration approach, each sensor learns its local sensing model from noisy measurements using an online algorithm and only transmits a few model parameters. In the second tier, sensors' local sensing models are then calibrated to a common system sensing model. Our approach fairly distributes computation overhead among sensors and significantly reduces the communication overhead of calibration compared with the device-level approach. Based on this approach, we develop an optimal model calibration scheme that maximizes the target detection probability of a sensor network under bounded false alarm rate. Our approach is evaluated by both experiments on a testbed of TelosB motes and extensive simulations based on synthetic datasets as well as data traces collected in a real vehicle detection experiment. The results demonstrate that our system-level calibration approach can significantly boost the detection performance of sensor networks in scenarios with low signal-to-noise ratios. Rui Tan 0001, Guoliang Xing, Zhaohui Yuan, Xue (Steve) Liu, Jianguo Yao 0002 |
ACM Trans. Sens. Networks | 5 |
| 2012 | Dynamic Control of Electricity Cost with Power Demand Smoothing and Peak Shaving for Distributed Internet Data CentersabstractInternet based service providers, such as Amazon, Google, Yahoo etc, build their data centers (IDC) across multiple regions to provide reliable and low latency of services to clients. Ever-increasing service demand, complexity of services and growing client population cause enormous power consumptions by these IDCs incurring a major part of their running costs. Modern electric power grid provides a feasible way to dynamically and efficiently manage the electricity cost of distributed IDCs based on the Locational Marginal Pricing (LMP) policy. While recent works exploit LMP by electricity-price based geographic load distribution, the dynamic workload and high volatility of electricity prices induce highly volatile power demand and critical power peak problem. The benefit of cost minimization via geographic load distribution is counterbalanced with the high cost incurred by violating the peak power. In this paper, we study the dynamic control of electricity cost to provide low volatility in power demand and shaving of power peaks. To this end, a Model Predictive Control (MPC) electricity cost minimization problem is formulated based on a time-continuous differential model. The proposed solution minimizes electricity costs, provides low variation in power demand by penalizing the change in workload and alleviates the power peaks by tracking the available power budget. By providing extensive simulation results based on real-life electricity price traces we show the effectiveness of our approach. Jianguo Yao 0002, Xue (Steve) Liu, Wenbo He 0003, Ashikur Rahman |
ICDCS | 1 |
| 2012 | Adaptive calibration for fusion-based cyber-physical systemsabstractMany Cyber-Physical Systems (CPS) are composed of low-cost devices that are deeply integrated with physical environments. As a result, the performance of a CPS system is inevitably undermined by various physical uncertainties , which include stochastic noises, hardware biases, unpredictable environment changes, and dynamics of the physical process of interest. Traditional solutions to these issues (e.g., device calibration and collaborative signal processing) work in an open-loop fashion and hence often fail to adapt to the uncertainties after system deployment. In this article, we propose an adaptive system-level calibration approach for a class of CPS systems whose primary objective is to detect events or targets of interest. Through collaborative data fusion, our calibration approach features a feedback control loop that exploits system heterogeneity to mitigate the impact of aforementioned uncertainties on the system performance. In contrast to existing heuristic-based solutions, our control-theoretical calibration algorithm can ensure provable system stability and convergence. We also develop a routing algorithm for fusion-based multihop CPS systems that is robust to communication unreliability and delay. Our approach is evaluated by both experiments on a testbed of Tmotes as well as extensive simulations based on data traces gathered from a real vehicle detection experiment. The results demonstrate that our calibration algorithm enables a CPS system to maintain the optimal sensing performance in the presence of various system and environmental dynamics. Rui Tan 0001, Guoliang Xing, Xue (Steve) Liu, Jianguo Yao 0002, Zhaohui Yuan |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2010 | Adaptive Calibration for Fusion-based Wireless Sensor NetworksabstractWireless sensor networks (WSNs) are typically composed of low-cost sensors that are deeply integrated with physical environments. As a result, the sensing performance of a WSN is inevitably undermined by various physical uncertainties, which include stochastic sensor noises, unpredictable environment changes and dynamics of the monitored phenomenon. Traditional solutions (e.g., sensor calibration and collaborative signal processing) work in an open-loop fashion and hence fail to adapt to these uncertainties after system deployment. In this paper, we propose an adaptive system-level calibration approach for a class of sensor networks that employ data fusion to improve system sensing performance. Our approach features a feedback control loop that exploits sensor heterogeneity to deal with the aforementioned uncertainties in calibrating system performance. In contrast to existing heuristic based solutions, our control-theoretical calibration algorithm can ensure provable system stability and convergence. We also systematically analyze the impacts of communication reliability and delay, and propose an optimal routing algorithm that minimizes the impact of packet loss on system stability. Our approach is evaluated by both experiments on a testbed of Tmotes as well as extensive simulations based on data traces gathered from a real vehicle detection experiment. The results demonstrate that our calibration algorithm enables a network to maintain the optimal detection performance in the presence of various system and environmental dynamics. Rui Tan 0001, Guoliang Xing, Xue (Steve) Liu, Jianguo Yao 0002, Zhaohui Yuan |
INFOCOM | 4 |
| 2010 | System-Level Calibration for Fusion-Based Wireless Sensor NetworksabstractWireless sensor networks are typically composed of low-cost sensors that are deeply integrated in physical environments. As a result, the sensing performance of a wireless sensor network is inevitably undermined by biases in imperfect sensor hardware and the noises in data measurements. Although a variety of calibration methods have been proposed to address these issues, they often adopt the device-level approach that becomes intractable for moderate- to large-scale networks. In this paper, we propose a two-tier system-level calibration approach for a class of sensor networks that employ data fusion to improve the sensing performance. In the first tier of our calibration approach, each sensor learns its local sensing model from noisy measurements using an online algorithm and only transmits a few model parameters. In the second tier, sensors' local sensing models are then calibrated to a common system sensing model. Our approach fairly distributes computation overhead among sensors and significantly reduces the communication overhead of calibration. Based on this approach, we develop an optimal model calibration scheme that maximizes the target detection probability of a sensor network under bounded false alarm rate. Our approach is evaluated by both experiments on a testbed of TelosB motes and extensive simulations based on data traces collected in a real vehicle detection experiment. The results demonstrate that our system-level calibration approach can significantly boost the detection performance of sensor networks in the scenarios with low signal-to-noise ratios. Rui Tan 0001, Guoliang Xing, Zhaohui Yuan, Xue (Steve) Liu, Jianguo Yao 0002 |
RTSS | 5 |
| 2010 | Online adaptive utilization control for real-time embedded multiprocessor systems
Jianguo Yao 0002, Xue (Steve) Liu, Zonghua Gu 0001, Jian Li 0021 |
J. Syst. Archit. | 1 |