Liang Zhou 0006

dblp:81/4761-6 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
3since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 6 first-author · 3 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2022 Cottage: Coordinated Time Budget Assignment for Latency, Quality and Power Optimization in Web Search
abstract
Most CPU power management techniques for web search assume that the time budget for a query is given a priori. However, determining the time budget on a per query granularity is challenging, because a difficult trade-off between the search latency, quality and power consumption has to be made. In this paper, we present Cottage, a coordinated time budget assignment framework between the aggregator and Index Serving Nodes (ISNs), which employs two distinct distributed search latency and quality predictors. The prediction results are integrated at a centralized optimizer for selecting the proper search time budget, while cutting off slow and low quality ISNs. Cottage also accelerates slow ISNs that have a high quality contribution, thus improving search quality. The implementation results on the Solr search engine show that Cottage outperforms state-of-the-art approaches with a 54% latency reduction and 41.3% less consumed power. In addition, the P@10 search quality with Cottage can still be as good as 0.947.
Liang Zhou 0006, Laxmi N. Bhuyan, K. K. Ramakrishnan
HPCA1
2021 Balancing Latency and Quality in Web Search
abstract
Selecting the right time budget for a search query is challenging because a proper balance between the search latency, quality and efficiency has to be maintained. State-of-the-art approaches leverage a centralized sample index at the aggregator to select the Index Serving Nodes (ISNs) to maintain quality and responsiveness. In this paper, we propose Cottage, a coordinated framework between the aggregator and ISNs for latency and quality optimization in web search. Cottage has two separate neural network models at each ISN to predict the quality contribution and latency, respectively. Then, these prediction results are sent back to the aggregator for latency and quality optimizations. The key task is integration of the predictions at the aggregator in determining an optimal dynamic time budget for identifying slow and low quality ISNs to improve latency and search efficiency. Our experiments on the Solr search engine prove that Cottage can reduce the average query latency by 54% and achieve a good P@10 search quality of 0.947.
Liang Zhou 0006, K. K. Ramakrishnan
NAS1
2021 PAVER: Locality Graph-Based Thread Block Scheduling for GPUs
abstract
The massive parallelism present in GPUs comes at the cost of reduced L1 and L2 cache sizes per thread, leading to serious cache contention problems such as thrashing. Hence, the data access locality of an application should be considered during thread scheduling to improve execution time and energy consumption. Recent works have tried to use the locality behavior of regular and structured applications in thread scheduling, but the difficult case of irregular and unstructured parallel applications remains to be explored. We present PAVER , a P riority- A ware V ertex schedul ER , which takes a graph-theoretic approach toward thread scheduling. We analyze the cache locality behavior among thread blocks ( TBs ) through a just-in-time compilation, and represent the problem using a graph representing the TBs and the locality among them. This graph is then partitioned to TB groups that display maximum data sharing, which are then assigned to the same streaming multiprocessor by the locality-aware TB scheduler. Through exhaustive simulation in Fermi, Pascal, and Volta architectures using a number of scheduling techniques, we show that PAVER reduces L2 accesses by 43.3%, 48.5%, and 40.21% and increases the average performance benefit by 29%, 49.1%, and 41.2% for the benchmarks with high inter-TB locality.
Devashree Tripathy, AmirAli Abdolrashidi, Laxmi N. Bhuyan, Liang Zhou 0006, Daniel Wong 0001
ACM Trans. Archit. Code Optim.4
2020 Swan: a two-step power management for distributed search engines
abstract
The service quality of web search depends considerably on the request tail latency from Index Serving Nodes (ISNs), prompting data centers to operate them at low utilization and wasting server power. ISNs can be made more energy efficient utilizing Dynamic Voltage and Frequency Scaling (DVFS) or sleep states techniques to take advantage of slack in latency of search queries. However, state-of-the-art frameworks use a single distribution to predict a request's service time and select a high percentile tail latency to derive the CPU's frequency or sleep states. Unfortunately, this misses plenty of energy saving opportunities. In this paper, we develop a simple linear regression predictor to estimate each individual search request's service time, based on the length of the request's posting list. To use this prediction for power management, the major challenge lies in reducing miss rates for deadlines due to prediction errors, while improving energy efficiency. We present Swan, a two-Step poWer mAnagement for distributed search eNgines. For each request, Swan selects an initial, lower frequency to optimize power, and then appropriately boosts the CPU frequency just at the right time to meet the deadline. Additionally, we re-configure the time instant for boosting frequency, when a critical request arrives and avoid deadline violations. Swan is implemented on the widely-used Solr search engine and evaluated with two representative, large query traces. Evaluations show Swan outperforms state-of-the-art approaches, saving at least 39% CPU power on average.
Liang Zhou 0006, Laxmi N. Bhuyan, K. K. Ramakrishnan
ISLPED1
2020 Gemini: Learning to Manage CPU Power for Latency-Critical Search Engines
abstract
Saving energy for latency-critical applications like web search can be challenging because of their strict tail latency constraints. State-of-the-art power management frameworks use Dynamic Voltage and Frequency Scaling (DVFS) and Sleep states techniques to slow down the request processing and finish the search just-in-time. However, accurately predicting the compute demand of a request can be difficult. In this paper, we present Gemini, a novel power management framework for latency-critical search engines. Gemini has two unique features to capture the per query service time variation. First, at light loads without request queuing, a two-step DVFS is used to manage the CPU power. Our two-step DVFS selects the initial CPU frequency based on the query specific service time prediction and then judiciously boosts the initial frequency at the right time to catch-up to the deadline. The determination of boosting time further relies on estimating the error in the prediction of individual query's service time. At high loads, where there is request queuing, only the current request being executed and the critical request in the queue adopt a two-step DVFS. All the other requests in-between use the same frequency to reduce the frequency transition overhead. Second, we develop two separate neural network models, one for predicting the service time and the other for the error in the prediction. The combination of these two predictors significantly improves the power saving and tail latency results of our two-step DVFS. Gemini is implemented on the Solr search engine. Evaluations on three representative query traces show that Gemini saves 41% of the CPU power, and is better than other state-of-the-art techniques.
Liang Zhou 0006, Laxmi N. Bhuyan, K. K. Ramakrishnan
MICRO1
2020 PacketUsher: Exploiting DPDK to accelerate compute-intensive packet processing
Qingqing Ren, Liang Zhou 0006, Zhijun Xu, Yujun Zhang 0001, Lei Zhang 0202
Comput. Commun.2
2019 Goldilocks: Adaptive Resource Provisioning in Containerized Data Centers
abstract
Power management in data centers is challenging because of fluctuating workloads and strict task completion time requirements. Recent resource provisioning systems, such as Borg and RC-Informed, pack tasks on servers to save power. However, current power optimization frameworks based on packing leave very little headroom for spikes, and the task completion times are compromised. In this paper, we design Goldilocks, a novel resource provisioning system for optimizing both power and task completion time by allocating tasks to servers in groups. Tasks hosted in containers are grouped together by running a graph partitioning algorithm. Containers communicating frequently are placed together, which improves the task completion times. We also leverage new findings on power consumption of modern-day servers to ensure that their utilizations are in a range where they are power-proportional. Both testbed implementation measurements and large-scale trace-driven simulations prove that Goldilocks outperforms all the previous works on data center power saving. Goldilocks saves power by 11.7%-26.2% depending on the workload, whereas the best of the implemented alternatives, Borg, saves 8.9%-22.8%. The energy per request for the Twitter content caching workload in Goldilocks is only 33% of RC-Informed. Finally, the best alternative in terms of task completion time, E-PVM, has 1.17-3.29 times higher task completion times than Goldilocks across different workloads.
Liang Zhou 0006, Laxmi N. Bhuyan, K. K. Ramakrishnan
ICDCS1
2018 Joint Server and Network Energy Saving in Data Centers for Latency-Sensitive Applications
abstract
Achieving energy proportionality in data centers supporting latency-sensitive applications is challenging because of the strict Service Level Agreements. Previous works individually focus on making the server energy proportional or reducing the data center network's power consumption for latency-tolerant applications. In this paper, we propose EPRONS to minimize the overall data center's power consumption with latency-sensitive applications by trading-off network slack in favor of providing additional slack for computations. We utilize the linear programming model to consolidate latency-sensitive search queries and latency-tolerant background flows to a minimal subnet of the topology by turning off unused switches and links without violating the application deadlines. Servers take advantage of the additional 'network-provided' slack to allow slowing down request processing. For servers, we design a novel power saving technique using Dynamic Voltage and Frequency Scaling (DVFS) based on the average tail latency of a request. If needed, we turn on a minimal number of additional network links and switches to reduce network latency while still maximizing entire data center's power saving. Experimental results show that our scheme saves up to 31.25% of a data center's total power budget.
Liang Zhou 0006, Chih-Hsun Chou, Laxmi N. Bhuyan, K. K. Ramakrishnan, Daniel Wong 0001
IPDPS1