VLDB 2026 Research / reviewers in the wild / expert
Shreya Varshini
dblp:291/2871
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0007-2235-5821ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEVI: Silent Data Corruption of Vector Instructions in Hyper-Scale DatacentersabstractSilent Data Corruption (SDC) poses a reliability threat in modern datacenters. These insidious errors evade detections and propagate incorrect results throughout the system. Companies including Google, Meta, and Alibaba have reported SDC incidents affecting their production. In this paper, we present the first comprehensive instruction- and application-level analysis of vector instruction SDCs in hyper-scale datacenters using a two-stage approach. We perform over 78 trillion test rounds in more than 14 billion CPU seconds. Our observations reveal undocumented SDC patterns that provide insights into possible underlying causes and inspire new mitigation strategies. Based on these findings, we propose a low-overhead SDC detection mechanism leveraging in-application algorithm-based fault tolerance. Our method achieves 88% to 100% SDC machine detection rate with a time overhead of only 1.35% even for modestly sized inputs. Yixuan Mei, Shreya Varshini, Harish Dattatraya Dixit, Sriram Sankar, K. V. Rashmi |
ASPLOS (2) | 2 |
| 2025 | Meta's Second Generation AI Chip: Model-Chip Co-Design and Productionization ExperiencesabstractThe rapid growth of AI workloads at Meta has motivated our inhouse development of AI chips, aiming to significantly reduce the total cost of ownership and mitigate risks posed by unpredictable GPU supplies.At ISCA'23, we presented Meta's first-generation AI chip, MTIA 1.This paper describes its successor, MTIA 2i, now deployed at scale and serving billions of users.MTIA 2i significantly improves upon MTIA 1, reducing total cost of ownership by 44% compared to GPUs while delivering competitive performance per watt.A key differentiator is its memory hierarchy: instead of costly HBM, it uses large SRAM alongside LPDDR.Although there has been a proliferation of publications on AI chips, they often focus on architectural design and overlook three critical aspects:(1) co-designing and optimizing ML models to work effectively with the AI chip; (2) demonstrating sufficient flexibility to support a wide range of models; and (3) during the productionization process, addressing challenges unanticipated or decisions deferred at design time, such as dealing with memory errors, safe overclocking, reducing provisioned power, and implementing real-time firmware updates to mitigate silicon design defects.A key contribution of this paper is sharing our experience with these aspects, based on our journey of productionizing MTIA 2i at scale. Joel Coburn, Chunqiang Tang, Sameer Abu Asal, Neeraj Agrawal, Raviteja Chinta, Harish Dattatraya Dixit, Brian Dodds, Saritha Dwarakapuram, Amin Firoozshahian, Cao Gao, Kaustubh Gondkar, Tyler Graf, Junhan Hu, Sterling Hughes, Adam Hutchin, Bhasker Jakka, Guoqiang Jerry Chen, Indu Kalyanaraman, Ashwin Kamath, Pankaj Kansal, Erum Kazi, Roman Levenstein, Mahesh Maddury, Alex Mastro, Siji Medaiyese, Pritesh Modi, Jack Montgomery, Nadathur Satish, Amit Nagpal, Ashwin Narasimha, Maxim Naumov, Eleanor Ozer, Jongsoo Park, Poorvaja Ramani, Harikrishna Reddy, David Reiss, Deboleena Roy, Sathish Sekar, Pavan Shetty, Aravind Sukumaran-Rajam, Eran Tal, Mike Tsai, Shreya Varshini, Richard Wareing, Olívia Wu, Xiaolong Xie, Hangchen Yu, Tanmay Zargar, Zitong Zeng, Feixiong Zhang, Ajit Mathews, Jiyuan Zhang 0008, Emmanuel Menage, Truls Edvard Stokke, Mohammed Sourouri |
ISCA | 45 |
| 2025 | CP-Bench: A PyTorch Test Suite to Detect AI Hardware Failure, Performance Degradation, and Silent Data CorruptionabstractThe growing complexity in manufacturing and operating the hardware in AI clusters leads to significant challenges in reliability. Hyperscalars have reported various AI hardware failures during high-stake jobs such as GenAI model training, where one GPU failure could bring down the entire training job. To tackle this issue, we present CP-Bench, an open-source, Configurable and Parameterizable, PyTorch-level test suite designed to test AI hardware failure, performance degradation, and silent data corruption (SDC). Built upon open-source projects, CP-Bench contains 30+ AI workloads (e.g., Llama), and implements various checks (e.g., SDC check) within these workloads. We have deployed CP-Bench throughout Meta’s AI hardware lifecycle, spanning manufacturing, in-production diagnostics, and device RMA; CP-Bench identified various hardware issues, some of which were not caught by vendor’s tooling. Notably, vendor has acknowledged to establish CP-Bench as a valid RMA criteria and plan to integrate CP-Bench into its tooling. CP-Bench is open-sourced at https://github.com/facebookincubator/CP-Bench. Sunny Yang, Suman Gumudavelli, Shreya Varshini, Abhinav Pandey, Abhinav Jauhri, Francesco Caggioni, Gautham Vunnam, Harish Dattatraya Dixit, Jason Liang, Philip Henzler, Sameeksha Gupta, Tyler Graf, Venkat Ramesh, Fan Fred Lin |
ITC | 4 |
| 2021 | CompOFA - Compound Once-For-All Networks for Faster Multi-Platform Deployment
Manas Sahni, Shreya Varshini, Alind Khare, Alexey Tumanov |
ICLR | 2 |