VLDB 2026 Research / reviewers in the wild / expert
Shobhit O. Kanaujia
dblp:70/4587
· DBLP profile ↗
8ranked-venue papers
0as first author
4since 2021 · last 2026
0009-0003-0542-3002ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads
Albert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia, Sultan Mahmud Sajal, Gefei Zuo, Krishna T. Malladi, Devon Akers, Kalyan Subramanian, Shobhit O. Kanaujia, Alexandros Daglis |
ISCA | 10 |
| 2026 | Vistara: Making CXL Real-Full Path From ASIC Design and OS Support to Hyperscale Deployment
Neha Gholkar, Jovan Stojkovic, Hasan Al Maruf, Gregory Price, Prakash Chauhan, Hiral Patel, Cedric Van Goethem Kiran Vemuri, Kiran Malwankar, Kishore Sriadibhatla, Kalyan Subramanian, Shobhit O. Kanaujia, Chunqiang Tang, Abhishek Dhanotia |
ISCA | 11 |
| 2025 | DCPerf: An Open-Source, Battle-Tested Performance Benchmark Suite for Datacenter WorkloadsabstractWe present DCPerf, the first open-source performance benchmark suite actively used to inform procurement decisions for millions of CPU in hyperscale datacenters.Although numerous benchmarks exist, our evaluation reveals that they inaccurately project server performance for datacenter workloads or fail to scale to resemble production workloads on modern many-core servers.DCPerf distinguishes itself in two aspects: (1) it faithfully models essential software architectures and features of datacenter applications, such as microservice architecture and highly optimized multi-process or multi-thread concurrency; and (2) it strives to align its performance characteristics with those of production workloads, at both the system level and microarchitecture level.Both are made possible by our direct access to the source code and hyperscale production deployments of datacenter workloads.Additionally, we share real-world examples of using DCPerf in critical decision-making, such as selecting future CPU SKUs and guiding CPU vendors in optimizing their designs.Our evaluation demonstrates that DCPerf accurately projects the performance of representative production workloads within a 3.3% error margin across four generations of production servers introduced over a span of six years, with core counts varying widely from 36 to 176. Wei Su 0005, Abhishek Dhanotia, Jayneel Gandhi, Neha Gholkar, Shobhit O. Kanaujia, Maxim Naumov, Kalyan Subramanian, Valentin Andrei, Chunqiang Tang |
ISCA | 6 |
| 2023 | TPP: Transparent Page Placement for CXL-Enabled Tiered-MemoryabstractThe increasing demand for memory in hyperscale applications has led to memory becoming a large portion of the overall datacenter spend. The emergence of coherent interfaces like CXL enables main memory expansion and offers an efficient solution to this problem. In such systems, the main memory can constitute different memory technologies with varied characteristics. In this paper, we characterize memory usage patterns of a wide range of datacenter applications across the server fleet of Meta. We, therefore, demonstrate the opportunities to offload colder pages to slower memory tiers for these applications. Without efficient memory management, however, such systems can significantly degrade performance. Hasan Al Maruf, Hao Wang 0011, Abhishek Dhanotia, Johannes Weiner, Niket Agarwal, Pallab Bhattacharya, Chris Petersen 0002, Mosharaf Chowdhury, Shobhit O. Kanaujia, Prakash Chauhan |
ASPLOS (3) | 9 |
| 2019 | Who's Afraid of Uncorrectable Bit Errors? Online Recovery of Flash Errors with Distributed Redundancy
Amy Tai, Andrew Kryczka, Shobhit O. Kanaujia, Kyle Jamieson, Michael J. Freedman, Asaf Cidon |
USENIX ATC | 3 |
| 2007 | Compression in cache designabstractIncreasing cache capacity via compression enables designers to improve performance of existing designs for small incremental cost, further leveraging the large die area invested in last level caches. This paper explores the compressed cache design space with focus on implementation feasibility. Ali-Reza Adl-Tabatabai, Anwar M. Ghuloum, Shobhit O. Kanaujia |
ICS | 3 |
| 2005 | Supporting Demanding Hard-Real-Time Systems with STIabstractSoftware thread integration (STI) is a compilation technique which enables the efficient use of an application's fine-grain idle time on generic processors without special hardware support. With STI, a primary function is automatically interleaved with a secondary function to create a single implicitly multithreaded function which minimizes context switching and, hence, both improves performance and also offers very fine-grain concurrency. In this work, we extend STI techniques to address two challenges. First, we reduce response time for interrupts or other high-priority threads by introducing polling servers into integrated threads. Second, we enable integration with long host threads, expanding the domain of STI. We derive methods to evaluate the response time for threads in systems with and without these new integration methods. We demonstrate these concepts with the integration of various threads in a sample hard-real-time system on a highly-constrained microcontroller. We use an inexpensive 20 MHz AVR 8-bit microcontroller to generate monochrome NTSC video while servicing a high-speed (115,2 kbaud) serial communication link. We have built and tested this system, achieving graphics rendering speed-ups of 3.99/spl times/ to 13.5/spl times/. Benjamin J. Welch, Shobhit O. Kanaujia, Adarsh Seetharam, Deepaksrivats Thirumalai, Alexander G. Dean |
IEEE Trans. Computers | 2 |
| 2003 | Extending STI for demanding hard-real-time systemsabstractSoftware thread integration (STI) is a compilation technique which enables the efficient use of an application's fine-grain idle time on generic processors without special hardware support. With STI, a primary function (with real-time requirements on specific instructions) is automatically interleaved with a secondary function to create a single implicitly multithreaded function which minimizes context switching and hence both improves performance and also offers very fine-grain concurrency.In this paper we extend STI techniques to address two challenges. First, we reduce response time for interrupts or other high-priority threads by introducing polling servers into integrated threads. Currently integrated threads disable interrupts, delaying all other work until their completion. Second, we enable integration with long host threads, expanding the domain of STI. With current techniques, if there are frequent interrupts, only host threads which can finish execution before the next interrupt can be integrated. We derive methods to evaluate the response time for threads in systems with and without these new integration methods.We demonstrate these concepts with the integration of various threads in a sample hard-real-time system on a highly-constrained microcontroller. We use an inexpensive 20 MHz AVR 8-bit microcontroller to generate monochrome NTSC video while servicing a high-speed (115.2 kbaud) serial communication link. We have built and tested this system and demonstrate graphics rendering speed-ups of 3.99x to 13.5x. Benjamin J. Welch, Shobhit O. Kanaujia, Adarsh Seetharam, Deepaksrivats Thirumalai, Alexander G. Dean |
CASES | 2 |