Cong Li 0010

dblp:74/3487-10 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
4since 2021 · last 2024
0000-0001-9732-9598ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2024 Foreseer: Knowledge-Driven Acceleration of Memory-Bound Matrix Multiplications for Large Language Model Inference
abstract
The majority of the latency in large language model inference lies in a few memory-bound matrix multiplication kernels. Traditionally, those kernels are optimized by extensive offline experiments of different tiling, striding, and slicing parameter choices on the target hardware. Scaling to different models, inference settings, and GPU hardware then becomes difficult. In this paper, we propose Foreseer, a new policy to generate the high-performing parameters for different kernels on the fly without offline experiments on the hardware. Foreseer leverages the specific characteristics of memory-bound matrix multiplication kernels, and synthesizes various prior knowledge such as keeping the adequate memory access pressure, avoiding the potential tail effect from load imbalance in dispatching, etc., into a simple penalty function. For a given parameter choice, Foreseer provides an analytical estimation of the impacting factors such as the memory access pressure, the potential tail effect, etc., to score the choice. A high-performing parameter choice is then discriminated from the penalty function. Micro-benchmarks of matrix multiplication kernels from popular language models on two different high-end Intel Data Center GPUs demonstrate that Foreseer outperforms different baselines including the vendor libraries, achieving >90% of the best-known performance obtained from the exhaustive offline parameter search experiments. End-to-end inference experiments on those language models with different settings also show that Foreseer accelerates the baseline by 20% on average, being competitive with and quite often exceeding the performance of the comprehensively-optimized vendor inference engine.
Cong Li 0010, Yutao Xu
SYSTOR1
2022 From Correctable Memory Errors to Uncorrectable Memory Errors: What Error Bits Tell
abstract
Uncorrectable memory errors are one of the major failure causes in datacenters. In this paper, we present an empirical study correlating correctable errors (CEs) and uncorrectable errors (UEs) using the large-scale field data across 3 major dual in-line memory module (DIMM) manufacturers from a contemporary server farm of ByteDance. Different from the previous studies, our study is the first to comprehend the error-bit information of CEs and the DIMM part numbers. Unlike the traditional chipkill error correction code (ECC), in contemporary Intel server platforms the ECC gets weakened, not able to tolerate some error-bit patterns from a single chip. Using obtainable coarse-grained ECC knowledge, we derive a new indicator from the error-bit information: risky CE occurrence in terms of ECC guaranteed coverage. From the data, we show that the new indicator has a consistently high sensitivity and specificity in the test of future UE occurrences across DIMMs from different manufacturers. This leads us to conjecture that the weakened ECC substantially contributes to many UEs today. The new risky CE indicator is then applied in predicting the future UE occurrence based on the CE history. We empirically demonstrate how practically useful predictors are constructed in conjunction with other useful attributes such as certain micro-level fault indicators and DIMM part numbers, achieving the state-of-the-art performance.
Cong Li 0010, Yu Zhang 0209, Tai Huang, Shen Zhou 0002, Shijian Ge
SC1
2021 Fault-Aware Prediction-Guided Page Offlining for Uncorrectable Memory Error Prevention
abstract
Uncorrectable memory errors are the major causes of hardware failures in datacenters leading to server crashes. Page offlining is an error-prevention mechanism implemented in modern operating systems. Traditional offlining policies are based on correctable error (CE) rate of a page in a past period. However, CEs are just the observations while the underlying causes are memory circuit faults. A certain fault such as a row fault can impact quite a few pages. Meanwhile, not all faults are equally prone to uncorrectable errors (UEs). In this paper, we propose a fault-aware prediction-guide policy for page offlining. In the proposed policy, we first identify row faults based on CE observations as the preliminary candidates for offlining. Leveraging the knowledge of the error correction code, we design a predictor based on error-bit patterns to predict whether a row fault is prone to UEs or not. Pages impacted by the UE-prone rows are then offlined. Empirical evaluation using the error log from a modern large-scale cluster in ByteDance demonstrates that the proposed policy avoids several times more UEs than the traditional policy does at a comparable cost of memory capacity loss due to page offlining.
Cong Li 0010, Shen Zhou 0002, Shijian Ge
ICCD2
2021 SHARC: improving adaptive replacement cache with shadow recency cache management
abstract
Adaptive Replacement Cache (ARC) is a state-of-the-art cache replacement policy with a constant-time complexity per request. It uses a recency list and a frequency list to balance between access recency and access frequency. In this paper, we re-examine the ARC policy and demonstrate its weaknesses: 1) some entries in the recency list are not recent; and 2) the constraint of the recency list length limits the capability in identifying weak locality. We then propose a new policy, Shadow ARC (SHARC), to overcome those weaknesses with shadow recency cache management. In SHARC, we track the virtual time of the accesses. We allow the shadow recency cache to grow on demand, but proactively identify unpromising entries for eviction based on a comprehensive eviction criterion. While the criterion is calculated from the virtual time, we provide the theoretical justification that in scenarios of strong locality, it tightly bounds the recency distance of the entries. In scenarios of relatively weak locality, the criterion dynamically determines the size of the shadow recency cache based on the activeness of the frequency cache items and the promotion activities of the recency items. Experimental results indicate that SHARC outperforms the state-of-the-art policies of ARC, Low Inter-Reference Recency Set (LIRS), and Dynamic LIRS.
Cong Li 0010
Middleware2
2020 DPCLS: Improving Partial Cache Line Sparing with Dynamics for Memory Error Prevention
abstract
On modern systems, memory failures constitute around half of the total hardware failures and negatively impact system reliability, availability, and serviceability. Partial cache line sparing (PCLS) is an error-prevention mechanism in memory controllers. PCLS statically encodes the locations of the faulty nibbles of bits into a sparing directory along with the corresponding data content for replacement during memory accesses. Due to the limited number of spare entries in memory controllers, the error-prevention capability of PCLS is weak. In this paper, we propose a new approach, dynamic PCLS (DPCLS), to overcome the weakness of PCLS. Different from the static error location encoding in the sparing directory in PCLS, DPCLS exploits the temporal localities in memory errors and uses a simple policy to dynamically admit and evict the faulty nibbles spared in the directory. Empirical evaluation demonstrates that DPCLS outperforms PCLS by avoiding more errors at the same cost of snare resources.
Cong Li 0010
ICCD2
2019 CLOCK-pro+: improving CLOCK-pro cache replacement with utility-driven adaptation
abstract
CLOCK-Pro is the low-overhead approximation of the state-of-the-art cache replacement policy, Low Inter-Reference Recency Set (LIRS). It also improves the static cache space allocation in LIRS with simple heuristics to adapt to LRU-friendly workloads. However, the heuristics do not perform well in certain cases. Inspired by the idea of utility-driven adaptation from another state-of-the-art policy, CLOCK for Adaptive Replacement (CAR), we propose a new CLOCK-Pro+ policy. The new policy directly evaluates the utility of growing the number of cold pages against that of growing hot pages. It then dynamically adjusts the cache space allocation driven by the utility comparison. Experiments are performed on traces from the UMass Trace Repository as well as a synthetic trace drawn from a stack-depth distribution. While sometimes CLOCK-Pro substantially outperforms CAR and sometimes vice versa, the new CLOCK-Pro+ policy consistently performs close to the winner between the two in all the cases.
Cong Li 0010
SYSTOR1
2018 Robust Distributed Anomaly Detection Using Optimal Weighted One-Class Random Forests
abstract
Wireless sensor networks (WSNs) have been widely deployed in various applications, e.g., agricultural monitoring and industrial monitoring, for their ease-of-deployment. The low-cost nature makes WSNs particularly vulnerable to changes of extrinsic factors, i.e., the environment, or changes of intrinsic factors, i.e., hardware or software failures. The problem can, often times, be uncovered via detecting unexpected behaviors (anomalies) of devices. However, anomaly detection in WSNs is subject to the following challenges: (1) the limited computation and connectivity, (2) the dynamicity of the environment and network topology, and (3) the need of taking real-time actions in response to anomalies. In this paper, we propose a novel framework using optimal weighted one-class random forests for unsupervised anomaly detection to address the aforementioned challenges in WSNs. The ample experiments showed that our framework not only is feasible but also outperforms the state-of-the-art unsupervised methods in terms of both detection accuracy and resource utilization.
Yu-Lin Tsou, Hong-Min Chu, Cong Li 0010, Shao-Wen Yang
ICDM3
2018 DLIRS: Improving Low Inter-Reference Recency Set Cache Replacement Policy with Dynamics
abstract
As one of the state-of-the-art policies for buffer cache replacement, Low Inter-Reference Recency Set (LIRS) uses Inter-Reference Recency (IRR) to predict future access behaviors of blocks. With a static allocation of most cache space to low IRR blocks, it does not perform well in some LRU-friendly workloads. Inspired by the idea of dynamic cache space partitioning from another state-of-the-art policy, Adaptive Replacement Cache (ARC), we propose a new Dynamic LIRS (DLIRS) policy. The new policy uses a simple mechanism to perform an approximated online estimation on how well IRR predicts future access behaviors, and then dynamically adapts the space allocation of low IRR blocks against high IRR blocks. Experiments are performed on traces from the UMass Trace Repository as well as a synthetic trace drawn from a stack depth distribution. While sometimes LIRS outperforms ARC with a significant margin and sometimes vice versa, the new DLIRS policy consistently performs close to the winner between ARC and LIRS in all the cases.
Cong Li 0010
SYSTOR1
2017 Optimizing low memory killers for mobile devices using reinforcement learning
abstract
Different from memory management mechanisms in legacy systems, those in mobile devices leverage the use of app caching to accelerate app launch performance. The responsiveness of switching back to a recently used app becomes significantly worse if the cached app process has been killed in the past due to a low memory constraint. We propose a new approach to optimize the low memory killer with reinforcement learning. The new low memory killer acts as an autonomic decision maker in an uncertain environment, continuously observing various indicators and metrics for memory management, making the process-killing decisions, and taking app launch latencies as the penalties from the decision-making environment. Through a trial-and-error exploration, the killer interacts with the dynamic environment and automatically learns a holistic policy through reinforcement learning optimizing the expected app launch latency over a long run. Preliminary experimental results show that the new approach consistently and significantly improves the app launch performance, outperforming the baselines.
Cong Li 0010, Jia Bao
IWCMC1
2004 Word Translation Disambiguation Using Bilingual Bootstrapping
abstract
This article proposes a new method for word translation disambiguation, one that uses a machine-learning technique called bilingual bootstrapping. In learning to disambiguate words to be translated, bilingual bootstrapping makes use of a small amount of classified data and a large amount of unclassified data in both the source and the target languages. It repeatedly constructs classifiers in the two languages in parallel and boosts the performance of the classifiers by classifying unclassified data in the two languages and by exchanging information regarding classified data between the two languages. Experimental results indicate that word translation disambiguation based on bilingual bootstrapping consistently and significantly outperforms existing methods that are based on monolingual bootstrapping.
Cong Li 0010
Comput. Linguistics2
2003 Text Classification Using Stochastic Keyword Generation
Cong Li 0010, Ji-Rong Wen, Hang Li 0001
ICML1
2002 Word Translation Disambiguation Using Bilingual Bootstrapping
abstract
This paper proposes a new method for word translation disambiguation using a machine learning technique called 'Bilingual Bootstrapping'. Bilingual Bootstrapping makes use of in learning, a small number of classified data and a large number of unclassified data in the source and the target languages in translation. It constructs classifiers in the two languages in parallel and repeatedly boosts the performances of the classifiers by further classifying data in each of the two languages and by exchanging between the two languages information regarding the classified data. Experimental results indicate that word translation disambiguation based on Bilingual Bootstrapping consistently and significantly outperforms the existing methods based on 'Monolingual Bootstrapping'.
Cong Li 0010
ACL1