VLDB 2026 Research / reviewers in the wild / expert
Fan Fred Lin
dblp:264/3611
· DBLP profile ↗
5ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0001-3749-4193ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Understanding Recommendation System Robustness Against Silent Data Corruption: An Empirical StudyabstractModern deep learning-based recommendation system (DRS) is a dominant workload in industrial data centers. However, with continuous transistor scaling and increasing hardware complexity, silent data corruption (SDC) has become a notable threat to the reliability of data center workloads. Multiple industry hyperscalars have reported the difficulty in addressing SDC due to their “stealthy” nature and elusive manifestation. Given the critical role that DRS plays in maintaining the quality of online services, understanding and enhancing its robustness against SDCs are imperative. To the best our knowledge, this paper presents the first empirical study on understanding DRS robustness against SDCs. Specifically, we develop PyTEI, a PyTorch-based, user-friendly, and highly-efficient error injection framework, based on which we perform large-scale error injection experiments to the parameters of five representative DRS models under three datasets. Experimental results reveal that, the sparsity level of input data and feature affect the robustness of DRS, and in particular, MLP modules inside DRS are especially vulnerable to SDCs. Further, we evaluate the effectiveness of three representative error mitigation methods – algorithm based fault tolerance (ABFT), activation clipping, and selective bit protection (SBP), in enhancing DRS robustness. Experimental results reveal that activation clipping obtains the best result by recovering up to 30% of the degraded DRS performance under SDCs. This study provides valuable insights for industry practitioners in developing robust fault-tolerant strategies for DRS workloads. We open source PyTEI at https://github.com/facebookresearch/PyTEI. Dongning Ma, Fan Fred Lin, Sriram Sankar |
ISSRE | 3 |
| 2025 | CP-Bench: A PyTorch Test Suite to Detect AI Hardware Failure, Performance Degradation, and Silent Data CorruptionabstractThe growing complexity in manufacturing and operating the hardware in AI clusters leads to significant challenges in reliability. Hyperscalars have reported various AI hardware failures during high-stake jobs such as GenAI model training, where one GPU failure could bring down the entire training job. To tackle this issue, we present CP-Bench, an open-source, Configurable and Parameterizable, PyTorch-level test suite designed to test AI hardware failure, performance degradation, and silent data corruption (SDC). Built upon open-source projects, CP-Bench contains 30+ AI workloads (e.g., Llama), and implements various checks (e.g., SDC check) within these workloads. We have deployed CP-Bench throughout Meta’s AI hardware lifecycle, spanning manufacturing, in-production diagnostics, and device RMA; CP-Bench identified various hardware issues, some of which were not caught by vendor’s tooling. Notably, vendor has acknowledged to establish CP-Bench as a valid RMA criteria and plan to integrate CP-Bench into its tooling. CP-Bench is open-sourced at https://github.com/facebookincubator/CP-Bench. Sunny Yang, Suman Gumudavelli, Shreya Varshini, Abhinav Pandey, Abhinav Jauhri, Francesco Caggioni, Gautham Vunnam, Harish Dattatraya Dixit, Jason Liang, Philip Henzler, Sameeksha Gupta, Tyler Graf, Venkat Ramesh, Fan Fred Lin |
ITC | 15 |
| 2024 | Dr. DNA: Combating Silent Data Corruptions in Deep Learning using Distribution of Neuron ActivationsabstractDeep neural networks (DNNs) have been widely-adopted in various safety-critical applications such as computer vision and autonomous driving. However, as technology scales and applications diversify, coupled with the increasing heterogeneity of underlying hardware architectures, silent data corruption (SDC) has been emerging as a pronouncing threat to the reliability of DNNs. Recent reports from industry hyperscalars underscore the difficulty in addressing SDC due to their "stealthy" nature and elusive manifestation. In this paper, we propose Dr. DNA, a novel approach to enhance the reliability of DNN systems by detecting and mitigating SDCs. Specifically, we formulate and extract a set of unique SDC signatures from the Distribution of Neuron Activations (DNA), based on which we propose early-stage detection and mitigation of SDCs during DNN inference. We perform an extensive evaluation across 3 vision tasks, 5 different datasets, and 10 different models, under 4 different error models. Results show that Dr. DNA achieves 100% SDC detection rate for most cases, 95% detection rate on average and >90% detection rate across all cases, representing 20% - 70% improvement over baselines. Dr. DNA can also mitigate the impact of SDCs by effectively recovering DNN model performance with <1% memory overhead and <2.5% latency overhead. Dongning Ma, Fan Fred Lin, Alban Desmaison, Joel Coburn, Sriram Sankar, Xun Jiao 0001 |
ASPLOS (3) | 2 |
| 2023 | Brief Industry Paper: Evaluating Robustness of Deep Learning-Based Recommendation Systems Against Hardware Errors: A Case StudyabstractDeep learning-based recommendation systems (DL-RMs) are industry-scale recommendation models developed by Meta, designed to make use of both categorical and numerical inputs to make personalized recommendations. To serve billions of users in real-time, DLRMs rely on high-performance hardware and accelerators within our data centers, optimizing for execution latency and recommendation quality. However, continuous technology scaling, expanding workload, and increasing hardware heterogeneity could lead to increased risk of hardware errors. Addressing this risk often involves introducing extra design redundancy, which can pose a non-negligible overhead in performance and latency. In this paper, we present a case study of evaluating DLRM robustness against hardware errors by performing an extensive error injection campaign to DLRM. Our findings unveil that DLRM is notably robust to hardware errors and we further find that embedding tables in DLRM show an especially strong robustness. Additionally, we explore a software-level error mitigation techniques, activation clipping, for mitigating the hardware errors, which improves the DLRM robustness further. This industrial case study of understanding and improving DLRM robustness can enable the system to continue to deliver timely recommendations even in the presence of hardware challenges, or reduce the timing latency overhead posed by design redundancy, enhancing overall recommendation system performance. Fan Fred Lin, Matt Xiao, Alban Desmaison, Sriram Sankar |
RTSS | 2 |
| 2020 | Optimizing Interrupt Handling Performance for Memory Failures in Large Scale Data CentersabstractIntermittent hardware failures are generally non-catastrophic and typical large-scale service infrastructures are designed to tolerate them while still serving user traffic. However, intermittent errors cause performance aberrations if they are not handled appropriately. System error reporting mechanisms send hardware interrupts to the Central Processing Unit (CPU) for handling the hardware errors. This disrupts the CPU's normal operation, which impacts the performance of the server. In this paper, we describe common intermittent hardware errors observed on server systems in a large-scale data center environment. We discuss two methodologies of handling interrupts in server systems - System Management Interrupt (SMI) and Corrected Machine Check Interrupt (CMCI). We characterize the performance of these methods in live environments as compared to prior studies that used error injection to simulate error behavior. Our experience shows that error injection methods are not reflective of production behavior. We also present a hybrid approach for handling error interrupts that achieves better performance, while preserving monitoring granularity, in large scale data center environments. Harish Dattatraya Dixit, Fan Fred Lin, Bill Holland, Matt Beadon, Zhengyu Yang 0003, Sriram Sankar |
ICPE | 2 |