EDBT 2026 Demo / reviewers in the wild / expert
Mohit Saxena
dblp:45/5501
· DBLP profile ↗
14ranked-venue papers
10as first author
1since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 4 first-authorComputer networks · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Storage systems · 62% Memory systems · 17% Cloud and datacenter computing · 15% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% | |
| Computer networks
2 papers |
Internet of things and sensor networks · 38% Network measurement and analytics · 32% Software-defined and programmable networks · 25% |
Topics — the 21 heaviest of 24, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data integration and cleaning
extract-transform-load |
0.7 | 1 | 2023 | The Story of AWS Glue · Proc. VLDB Endow. 2023 |
Storage systems
flash and SSD |
0.4 | 3 | 2014 | Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014 FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012 FlashVM: Virtual Memory Management on Flash · USENIX ATC 2010 |
Storage systems
distributed storage |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems › storage reliability
erasure coding |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems › file systems › distributed file system
HDFS |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Storage systems
storage reliability |
0.2 | 1 | 2015 | A Tale of Two Erasure Codes in HDFS · FAST 2015 |
Memory systems
cache |
0.2 | 2 | 2014 | FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012 Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014 |
Data integration and cleaning
metadata management |
0.2 | 1 | 2023 | The Story of AWS Glue · Proc. VLDB Endow. 2023 |
Cloud and datacenter computing › serverless computing
serverless analytics |
0.2 | 1 | 2023 | The Story of AWS Glue · Proc. VLDB Endow. 2023 |
Cloud and datacenter computing
serverless computing |
0.2 | 1 | 2023 | The Story of AWS Glue · Proc. VLDB Endow. 2023 |
Storage systems › flash and SSD › flash memory management
flash translation layer |
0.2 | 1 | 2014 | Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014 |
Performance modeling and evaluation
simulation |
0.2 | 1 | 2013 | Getting real: lessons in transitioning research simulations into hardware systems · FAST 2013 |
Storage systems › flash and SSD
SSD cache |
0.1 | 1 | 2012 | FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012 |
Memory systems › cache management › storage caching
storage cache |
0.1 | 1 | 2012 | FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012 |
Memory systems
virtual memory management |
0.1 | 1 | 2010 | FlashVM: Virtual Memory Management on Flash · USENIX ATC 2010 |
Software-defined and programmable networks › SDN measurement
flow statistics collection |
0.1 | 1 | 2009 | A Framework for Efficient Class-Based Scheduling · INFOCOM 2009 |
Network measurement and analytics
traffic classification |
0.1 | 1 | 2009 | A Framework for Efficient Class-Based Scheduling · INFOCOM 2009 |
Internet of things and sensor networks › wireless sensor network
testbed |
0.1 | 1 | 2007 | A sensor-cyber network testbed for plume detection, identification, and tracking · IPSN 2007 |
Internet of things and sensor networks
wireless sensor network |
0.1 | 1 | 2007 | A sensor-cyber network testbed for plume detection, identification, and tracking · IPSN 2007 |
Storage systems
storage hierarchy |
0.1 | 1 | 2014 | Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014 |
Network measurement and analytics › sampling
traffic sampling |
0.0 | 1 | 2009 | A Framework for Efficient Class-Based Scheduling · INFOCOM 2009 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.2prototyping · 0.2cache design · 0.1packet sampling · 0.1composite bloom filter · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | The Story of AWS GlueabstractAWS Glue is Amazon's serverless data integration cloud service that makes it simple and cost effective to extract, clean, enrich, load, and organize data. Originally launched in August 2017, AWS Glue began as an extract-transform-load (ETL) service designed to relieve developers and data engineers of the undifferentiated heavy lifting needed to load databases, data warehouses, and build data lakes on Amazon S3. Since then, it has evolved to serve a larger audience including ETL specialists and data scientists, and includes a broader suite of data integration capabilities. Today, hundreds of thousands of customers use AWS Glue every month. In this paper, we describe the use cases and challenges cloud customers face in preparing data for analytics and the tenets we chose to drive Glue's design. We chose early on to focus on ease-of-use, scale, and extensibility. At its core, Glue offers serverless Apache Spark and Python engines backed by a purpose-built resource manager for fast startup and auto-scaling. In Spark, it offers a new data structure --- DynamicFrames --- for manipulating messy schema-free semi-structured data such as event logs, a variety of transformations and tooling to simplify data preparation, and a new shuffle plugin to offload to cloud storage. It also includes a Hivemetastore compatible Data Catalog with Glue crawlers to build and manage metadata, e.g. for data lakes on Amazon S3. Finally, Glue Studio is its visual interface for authoring Spark and Python-based ETL jobs. We describe the innovations that differentiate AWS Glue and drive its popularity and how it has evolved over the years. Mohit Saxena, Ben Sowell, Daiyan Alamgir, Nitin Bahadur, Bijay Bisht, Santosh Chandrachood, Chitti Keswani, G2 Krishnamoorthy, Austin Lee, Bohou Li, Zach Mitchell, Vaibhav Porwal, Maheedhar Reddy Chappidi, Brian Ross, Noritaka Sekiyama, Omer Zaki, Linchi Zhang, Mehul A. Shah |
Proc. VLDB Endow. | 1 |
| 2016 | Neutrino: Revisiting Memory Caching for Iterative Data Analytics
Erci Xu, Mohit Saxena, Lawrence Chiu |
HotStorage | 2 |
| 2015 | A Tale of Two Erasure Codes in HDFS
Mingyuan Xia 0001, Mohit Saxena, Mario Blaum, David Pease |
FAST | 2 |
| 2015 | Regions-of-interest based automated diagnosis of Parkinson's disease using T1-weighted MRI
Bharti Rana 0001, Akanksha Juneja, Mohit Saxena, Sunita Gudwani, S. Senthil Kumaran, R. K. Agrawal 0001, Madhuri Behari |
Expert Syst. Appl. | 3 |
| 2014 | Design and Prototype of a Solid-State CacheabstractThe availability of high-speed solid-state storage has introduced a new tier into the storage hierarchy. Low-latency and high-IOPS solid-state drives (SSDs) cache data in front of high-capacity disks. However, most existing SSDs are designed to be a drop-in disk replacement, and hence are mismatched for use as a cache. This article describes FlashTier , a system architecture built upon a solid-state cache (SSC), which is a flash device with an interface designed for caching. Management software at the operating system block layer directs caching. The FlashTier design addresses three limitations of using traditional SSDs for caching. First, FlashTier provides a unified logical address space to reduce the cost of cache block management within both the OS and the SSD. Second, FlashTier provides a new SSC block interface to enable a warm cache with consistent data after a crash. Finally, FlashTier leverages cache behavior to silently evict data blocks during garbage collection to improve performance of the SSC. We first implement an SSC simulator and a cache manager in Linux to perform an in-depth evaluation and analysis of FlashTier's design techniques. Next, we develop a prototype of SSC on the OpenSSD Jasmine hardware platform to investigate the benefits and practicality of FlashTier design. Our prototyping experiences provide insights applicable to managing modern flash hardware, implementing other SSD prototypes and new OS storage stack interface extensions. Overall, we find that FlashTier improves cache performance by up to 168% over consumer-grade SSDs and up to 52% over high-end SSDs. It also improves flash lifetime for write-intensive workloads by up to 60% compared to SSD caches with a traditional flash interface. Mohit Saxena, Michael M. Swift |
ACM Trans. Storage | 1 |
| 2013 | Getting real: lessons in transitioning research simulations into hardware systems
Mohit Saxena, Yiying Zhang 0005, Michael M. Swift, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau |
FAST | 1 |
| 2012 | Hathi: durable transactions for memory using flashabstractRecent architectural trends---cheap, fast solid-state storage, inexpensive DRAM, and multi-core CPUs---provide an opportunity to rethink the interface between applications and persistent storage. To leverage these advances, we propose a new system architecture called Hathi that provides an in-memory transactional heap made persistent using high-speed flash drives. With Hathi, programmers can make consistent concurrent updates to in-memory data structures that survive system failures. Mohit Saxena, Mehul A. Shah, Stavros Harizopoulos, Michael M. Swift, Arif Merchant |
DaMoN | 1 |
| 2012 | FlashTier: a lightweight, consistent and durable storage cacheabstractThe availability of high-speed solid-state storage has introduced a new tier into the storage hierarchy. Low-latency and high-IOPS solid-state drives (SSDs) cache data in front of high-capacity disks. However, most existing SSDs are designed to be a drop-in disk replacement, and hence are mismatched for use as a cache. Mohit Saxena, Michael M. Swift, Yiying Zhang 0005 |
EuroSys | 1 |
| 2010 | FlashVM: Virtual Memory Management on Flash
Mohit Saxena, Michael M. Swift |
USENIX ATC | 1 |
| 2010 | CLAMP: Efficient class-based sampling for flexible flow monitoring
Mohit Saxena, Ramana Rao Kompella |
Comput. Networks | 1 |
| 2009 | FlashVM: Revisiting the Virtual Memory Hierarchy
Mohit Saxena, Michael M. Swift |
HotOS | 1 |
| 2009 | A Framework for Efficient Class-Based SchedulingabstractWith an increasing requirement for network monitoring tools to classify traffic and track security threats, newer and efficient ways are needed for collecting traffic statistics and monitoring of network flows. However, traditional solutions based on random packet sampling treat all flows as equal and therefore, do not provide the flexibility required for these applications. In this paper, we propose a novel architecture called CLAMP that provides an efficient framework to implement size-based sampling. At the heart of CLAMP is a novel data structure called composite bloom filter (CBF) that consists of a set of bloom filters that work together to encapsulate various class definitions. In comparison to previous approaches that implement simple size-based sampling, our architecture requires substantially lower memory (upto 80x) and results in higher flow coverage (upto 8x more flows) under specific configurations. Mohit Saxena, Ramana Rao Kompella |
INFOCOM | 1 |
| 2008 | Analyzing video services in Web 2.0: a global perspectiveabstractServing multimedia content over the Internet with negligible delay remains a challenge. With the advent of Web 2.0, numerous video sharing sites using different storage and content delivery models have become popular. Yet, little is known about these models from a global perspective. Such an understanding is important for designing systems which can efficiently serve video content to users all over the world. In this paper, we analyze and compare the underlying distribution frameworks of three video sharing services - YouTube, Dailymotion and Metacafe - based on traces collected from measurements over a period of 23 days. We investigate the variation in service delay with the user's geographical location and with video characteristics such as age and popularity. We leverage multiple vantage points distributed around the globe to validate our observations. Our results represent some of the first measurements directed towards analyzing these recently popular services. Mohit Saxena, Umang Sharan, Sonia Fahmy |
NOSSDAV | 1 |
| 2007 | A sensor-cyber network testbed for plume detection, identification, and trackingabstractNo abstract available. Jren-Chit Chin, I-Hong Hou, Jennifer C. Hou, Chris Y. T. Ma, Nageswara S. V. Rao, Mohit Saxena, Mallikarjun Shankar, Yong Yang 0009, David K. Y. Yau |
IPSN | 6 |