Mohit Saxena

dblp:45/5501 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
1since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-authorComputer networks · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 62% Memory systems · 17% Cloud and datacenter computing · 15%
Databases, data mining, and information retrieval
1 paper
Data integration and cleaning · 100%
Computer networks
2 papers
Internet of things and sensor networks · 38% Network measurement and analytics · 32% Software-defined and programmable networks · 25%

Topics — the 21 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data integration and cleaning
extract-transform-load
0.712023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Storage systems
flash and SSD
0.432014
Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014
FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012
FlashVM: Virtual Memory Management on Flash · USENIX ATC 2010
Storage systems
distributed storage
0.212015
A Tale of Two Erasure Codes in HDFS · FAST 2015
Storage systems › storage reliability
erasure coding
0.212015
A Tale of Two Erasure Codes in HDFS · FAST 2015
Storage systems › file systems › distributed file system
HDFS
0.212015
A Tale of Two Erasure Codes in HDFS · FAST 2015
Storage systems
storage reliability
0.212015
A Tale of Two Erasure Codes in HDFS · FAST 2015
Memory systems
cache
0.222014
FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012
Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014
Data integration and cleaning
metadata management
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Cloud and datacenter computing › serverless computing
serverless analytics
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Cloud and datacenter computing
serverless computing
0.212023
The Story of AWS Glue · Proc. VLDB Endow. 2023
Storage systems › flash and SSD › flash memory management
flash translation layer
0.212014
Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014
Performance modeling and evaluation
simulation
0.212013
Getting real: lessons in transitioning research simulations into hardware systems · FAST 2013
Storage systems › flash and SSD
SSD cache
0.112012
FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012
Memory systems › cache management › storage caching
storage cache
0.112012
FlashTier: a lightweight, consistent and durable storage cache · EuroSys 2012
Memory systems
virtual memory management
0.112010
FlashVM: Virtual Memory Management on Flash · USENIX ATC 2010
Software-defined and programmable networks › SDN measurement
flow statistics collection
0.112009
A Framework for Efficient Class-Based Scheduling · INFOCOM 2009
Network measurement and analytics
traffic classification
0.112009
A Framework for Efficient Class-Based Scheduling · INFOCOM 2009
Internet of things and sensor networks › wireless sensor network
testbed
0.112007
A sensor-cyber network testbed for plume detection, identification, and tracking · IPSN 2007
Internet of things and sensor networks
wireless sensor network
0.112007
A sensor-cyber network testbed for plume detection, identification, and tracking · IPSN 2007
Storage systems
storage hierarchy
0.112014
Design and Prototype of a Solid-State Cache · ACM Trans. Storage 2014
Network measurement and analytics › sampling
traffic sampling
0.012009
A Framework for Efficient Class-Based Scheduling · INFOCOM 2009

Methods — techniques the papers use, named apart from their topics

simulation · 0.2prototyping · 0.2cache design · 0.1packet sampling · 0.1composite bloom filter · 0.1
YearPublicationVenuePosition
2023 The Story of AWS Glue
abstract
AWS Glue is Amazon's serverless data integration cloud service that makes it simple and cost effective to extract, clean, enrich, load, and organize data. Originally launched in August 2017, AWS Glue began as an extract-transform-load (ETL) service designed to relieve developers and data engineers of the undifferentiated heavy lifting needed to load databases, data warehouses, and build data lakes on Amazon S3. Since then, it has evolved to serve a larger audience including ETL specialists and data scientists, and includes a broader suite of data integration capabilities. Today, hundreds of thousands of customers use AWS Glue every month. In this paper, we describe the use cases and challenges cloud customers face in preparing data for analytics and the tenets we chose to drive Glue's design. We chose early on to focus on ease-of-use, scale, and extensibility. At its core, Glue offers serverless Apache Spark and Python engines backed by a purpose-built resource manager for fast startup and auto-scaling. In Spark, it offers a new data structure --- DynamicFrames --- for manipulating messy schema-free semi-structured data such as event logs, a variety of transformations and tooling to simplify data preparation, and a new shuffle plugin to offload to cloud storage. It also includes a Hivemetastore compatible Data Catalog with Glue crawlers to build and manage metadata, e.g. for data lakes on Amazon S3. Finally, Glue Studio is its visual interface for authoring Spark and Python-based ETL jobs. We describe the innovations that differentiate AWS Glue and drive its popularity and how it has evolved over the years.
Mohit Saxena, Ben Sowell, Daiyan Alamgir, Nitin Bahadur, Bijay Bisht, Santosh Chandrachood, Chitti Keswani, G2 Krishnamoorthy, Austin Lee, Bohou Li, Zach Mitchell, Vaibhav Porwal, Maheedhar Reddy Chappidi, Brian Ross, Noritaka Sekiyama, Omer Zaki, Linchi Zhang, Mehul A. Shah
Proc. VLDB Endow.1
2016 Neutrino: Revisiting Memory Caching for Iterative Data Analytics
Erci Xu, Mohit Saxena, Lawrence Chiu
HotStorage2
2015 A Tale of Two Erasure Codes in HDFS
Mingyuan Xia 0001, Mohit Saxena, Mario Blaum, David Pease
FAST2
2015 Regions-of-interest based automated diagnosis of Parkinson's disease using T1-weighted MRI
Bharti Rana 0001, Akanksha Juneja, Mohit Saxena, Sunita Gudwani, S. Senthil Kumaran, R. K. Agrawal 0001, Madhuri Behari
Expert Syst. Appl.3
2014 Design and Prototype of a Solid-State Cache
abstract
The availability of high-speed solid-state storage has introduced a new tier into the storage hierarchy. Low-latency and high-IOPS solid-state drives (SSDs) cache data in front of high-capacity disks. However, most existing SSDs are designed to be a drop-in disk replacement, and hence are mismatched for use as a cache. This article describes FlashTier , a system architecture built upon a solid-state cache (SSC), which is a flash device with an interface designed for caching. Management software at the operating system block layer directs caching. The FlashTier design addresses three limitations of using traditional SSDs for caching. First, FlashTier provides a unified logical address space to reduce the cost of cache block management within both the OS and the SSD. Second, FlashTier provides a new SSC block interface to enable a warm cache with consistent data after a crash. Finally, FlashTier leverages cache behavior to silently evict data blocks during garbage collection to improve performance of the SSC. We first implement an SSC simulator and a cache manager in Linux to perform an in-depth evaluation and analysis of FlashTier's design techniques. Next, we develop a prototype of SSC on the OpenSSD Jasmine hardware platform to investigate the benefits and practicality of FlashTier design. Our prototyping experiences provide insights applicable to managing modern flash hardware, implementing other SSD prototypes and new OS storage stack interface extensions. Overall, we find that FlashTier improves cache performance by up to 168% over consumer-grade SSDs and up to 52% over high-end SSDs. It also improves flash lifetime for write-intensive workloads by up to 60% compared to SSD caches with a traditional flash interface.
Mohit Saxena, Michael M. Swift
ACM Trans. Storage1
2013 Getting real: lessons in transitioning research simulations into hardware systems
Mohit Saxena, Yiying Zhang 0005, Michael M. Swift, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST1
2012 Hathi: durable transactions for memory using flash
abstract
Recent architectural trends---cheap, fast solid-state storage, inexpensive DRAM, and multi-core CPUs---provide an opportunity to rethink the interface between applications and persistent storage. To leverage these advances, we propose a new system architecture called Hathi that provides an in-memory transactional heap made persistent using high-speed flash drives. With Hathi, programmers can make consistent concurrent updates to in-memory data structures that survive system failures.
Mohit Saxena, Mehul A. Shah, Stavros Harizopoulos, Michael M. Swift, Arif Merchant
DaMoN1
2012 FlashTier: a lightweight, consistent and durable storage cache
abstract
The availability of high-speed solid-state storage has introduced a new tier into the storage hierarchy. Low-latency and high-IOPS solid-state drives (SSDs) cache data in front of high-capacity disks. However, most existing SSDs are designed to be a drop-in disk replacement, and hence are mismatched for use as a cache.
Mohit Saxena, Michael M. Swift, Yiying Zhang 0005
EuroSys1
2010 FlashVM: Virtual Memory Management on Flash
Mohit Saxena, Michael M. Swift
USENIX ATC1
2010 CLAMP: Efficient class-based sampling for flexible flow monitoring
Mohit Saxena, Ramana Rao Kompella
Comput. Networks1
2009 FlashVM: Revisiting the Virtual Memory Hierarchy
Mohit Saxena, Michael M. Swift
HotOS1
2009 A Framework for Efficient Class-Based Scheduling
abstract
With an increasing requirement for network monitoring tools to classify traffic and track security threats, newer and efficient ways are needed for collecting traffic statistics and monitoring of network flows. However, traditional solutions based on random packet sampling treat all flows as equal and therefore, do not provide the flexibility required for these applications. In this paper, we propose a novel architecture called CLAMP that provides an efficient framework to implement size-based sampling. At the heart of CLAMP is a novel data structure called composite bloom filter (CBF) that consists of a set of bloom filters that work together to encapsulate various class definitions. In comparison to previous approaches that implement simple size-based sampling, our architecture requires substantially lower memory (upto 80x) and results in higher flow coverage (upto 8x more flows) under specific configurations.
Mohit Saxena, Ramana Rao Kompella
INFOCOM1
2008 Analyzing video services in Web 2.0: a global perspective
abstract
Serving multimedia content over the Internet with negligible delay remains a challenge. With the advent of Web 2.0, numerous video sharing sites using different storage and content delivery models have become popular. Yet, little is known about these models from a global perspective. Such an understanding is important for designing systems which can efficiently serve video content to users all over the world. In this paper, we analyze and compare the underlying distribution frameworks of three video sharing services - YouTube, Dailymotion and Metacafe - based on traces collected from measurements over a period of 23 days. We investigate the variation in service delay with the user's geographical location and with video characteristics such as age and popularity. We leverage multiple vantage points distributed around the globe to validate our observations. Our results represent some of the first measurements directed towards analyzing these recently popular services.
Mohit Saxena, Umang Sharan, Sonia Fahmy
NOSSDAV1
2007 A sensor-cyber network testbed for plume detection, identification, and tracking
abstract
No abstract available.
Jren-Chit Chin, I-Hong Hou, Jennifer C. Hou, Chris Y. T. Ma, Nageswara S. V. Rao, Mohit Saxena, Mallikarjun Shankar, Yong Yang 0009, David K. Y. Yau
IPSN6