VLDB 2026 Research / reviewers in the wild / expert
Raghul Gunasekaran
dblp:86/1044
· DBLP profile ↗
10ranked-venue papers
1as first author
0since 2021 · last 2017
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1Computer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Storage systems · 34% High-performance computing · 19% Performance modeling and evaluation · 14% | |
| Databases, data mining, and information retrieval
1 paper |
Data integration and cleaning · 100% |
Topics — the 7 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems › distributed file system
parallel file system |
0.5 | 2 | 2017 | Scientific user behavior and data-sharing trends in a petascale file system · SC 2017 Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 |
Cloud and datacenter computing
log analysis |
0.3 | 1 | 2017 | GUIDE: a scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facility · SC 2017 |
Performance modeling and evaluation
workload characterization |
0.3 | 2 | 2017 | Automatic identification of application I/O signatures from noisy server-side traces · FAST 2014 Scientific user behavior and data-sharing trends in a petascale file system · SC 2017 |
Parallel and multicore computing › parallel scheduling › resource-aware scheduling
i/o-aware scheduling |
0.2 | 1 | 2016 | Server-side log data analytics for I/O workload characterization and coordination on large shared storage systems · SC 2016 |
Storage systems
i/o workload characterization |
0.2 | 1 | 2016 | Server-side log data analytics for I/O workload characterization and coordination on large shared storage systems · SC 2016 |
High-performance computing
supercomputing |
0.2 | 1 | 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 |
Storage systems
storage reliability |
0.1 | 1 | 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File Systems · SC 2014 |
Methods — techniques the papers use, named apart from their topics
log federation · 0.6data warehousing · 0.6quantitative file system metrics · 0.3metadata snapshot analysis · 0.3pattern mining · 0.2log analytics · 0.2technology evaluation · 0.2benchmarking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Applying Graph Analytics to Understand Compute Core Usage and Publication Trends in a Petascale Supercomputing FacilityabstractThe Oak Ridge Leadership Computing Facility (OLCF) runs Titan, the No. 4 supercomputer in the world, to deliver over four billion compute core hours every year to several scientific domains, in their pursuit of leadership science. In this paper, we analyze four years worth of heterogeneous log data sources from the OLCF resource fabric, capturing metadata on entities such as users (2,546), scientific project allocations (674), jobs (1,352,402) and publications (1,146), to derive insights into the trends in core hour usage and publications, across 35 science domains. We have constructed a scalable graph to represent the OLCF entities and apply rich graph analytics for our analysis. Based on this, we have analyzed the metadata across five dimensions, namely (1) quantitative analysis of Titan system usage, (2) quantitative analysis of OLCF publications, (3) correlation analysis between system usage and publications, (4) text analysis to derive OLCF research trends, and (5) utilization of graph mining for association analysis. To the best of our knowledge, our work is the first of its kind to apply graph- based big data techniques to provide comprehensive insights on an HPC center's core hour usage and users' publication trends. Our results provide valuable details into an HPC center's core allocation program, measuring the productivity of scientific domains, the interplay between core usage and research output, accelerating collaboration, and in predicting new connections between resource entities. Sangkeun Matt Lee, Sudharshan S. Vazhkudai, Raghul Gunasekaran |
HiPC | 3 |
| 2017 | Scientific user behavior and data-sharing trends in a petascale file systemabstractThe Oak Ridge Leadership Computing Facility (OLCF) runs the No. 4 supercomputer in the world, supported by a petascale file system, to facilitate scientific discovery. In this paper, using the daily file system metadata snapshots collected over 500 days, we have studied the behavioral trends of 1, 362 active users and 380 projects across 35 science domains. In particular, we have analyzed both individual and collective behavior of users and projects, highlighting needs from individual communities and the overall requirements to operate the file system. We have analyzed the metadata across three dimensions, namely (i) the projects' file generation and usage trends, using quantitative file system-centric metrics, (ii) scientific user behavior on the file system, and (iii) the data sharing trends of users and projects. To the best of our knowledge, our work is the first of its kind to provide comprehensive insights on user behavior from multiple science domains through metadata analysis of a large-scale shared file system. We envision that this OLCF case study will provide valuable insights for the design, operation, and management of storage systems at scale, and also encourage other HPC centers to undertake similar such efforts. Seung-Hwan Lim, Hyogi Sim, Raghul Gunasekaran, Sudharshan S. Vazhkudai |
SC | 3 |
| 2017 | GUIDE: a scalable information directory service to collect, federate, and analyze logs for operational insights into a leadership HPC facilityabstractIn this paper, we describe the GUIDE framework used to collect, federate, and analyze log data from the Oak Ridge Leadership Computing Facility (OLCF), and how we use that data to derive insights into facility operations. We collect system logs and extract monitoring data at every level of the various OLCF subsystems, and have developed a suite of pre-processing tools to make the raw data consumable. The cleansed logs are then ingested and federated into a central, scalable data warehouse, Splunk, that offers storage, indexing, querying, and visualization capabilities. We have further developed and deployed a set of tools to analyze these multiple disparate log streams in concert and derive operational insights. We describe our experience from developing and deploying the GUIDE infrastructure, and deriving valuable insights on the various subsystems, based on two years of operations in the production OLCF environment. Sudharshan S. Vazhkudai, Ross G. Miller, Devesh Tiwari, Christopher Zimmer 0001, Feiyi Wang, Sarp Oral, Raghul Gunasekaran, Deryl Steinert |
SC | 7 |
| 2016 | Constellation: A science graph network for scalable data and knowledge discovery in extreme-scale scientific collaborationsabstractConstellation's overarching goal is the federation of information from resources within an extreme-scale scientific collaboration to enable the scalable discovery of data and new knowledge pathways. The resource fabric is comprised of petascale supercomputers and storage systems, users, jobs, datasets and lifecycle artifacts. For an extreme-scale supercomputing center, normal operations can generate hundreds of millions of data products and metadata entries describing the resource fabric. Constellation federates the information extracted from the resources using a custom, transformative science graph network; constructs rich metadata indexes and higher-order derived metadata from the extracted information; and conducts scalable graph analytics to unravel hidden data pathways. Our implementation and deployment for a production, supercomputing facility shows that the graph can scale to more than 750 million vertices, its domain agnostic indexing can answer interesting science queries, and its analytics can aid in structural, topological and temporal analysis to identify usage hotspots. Sudharshan S. Vazhkudai, John Harney, Raghul Gunasekaran, Dale Stansberry, Seung-Hwan Lim, Tom Barron, Andrew Nash, Arvind Ramanathan |
IEEE BigData | 3 |
| 2016 | Server-side log data analytics for I/O workload characterization and coordination on large shared storage systemsabstractInter-application I/O contention and performance interference have been recognized as severe problems. In this work, we demonstrate, through measurement from Titan (world's No. 3 supercomputer), that high I/O variance co-exists with the fact that individual storage units remain under-utilized for the majority of the time. This motivates us to propose AID, a system that performs automatic application I/O characterization and I/O-aware job scheduling. AID analyzes existing I/O traffic and batch job history logs, without any prior knowledge on applications or user/developer involvement. It identifies the small set of I/O-intensive candidates among all applications running on a supercomputer and subsequently mines their I/O patterns, using more detailed per-I/O-node traffic logs. Based on such auto-extracted information, AID provides online I/O-aware scheduling recommendations to steer I/O-intensive applications away from heavy ongoing I/O activities. We evaluate AID on Titan, using both real applications (with extracted I/O patterns validated by contacting users) and our own pseudo-applications. Our results confirm that AID is able to (1) identify I/O-intensive applications and their detailed I/O characteristics, and (2) significantly reduce these applications' I/O performance degradation/variance by jointly evaluating outstanding applications' I/O pattern and real-time system l/O load. Yang Liu 0129, Raghul Gunasekaran, Xiaosong Ma, Sudharshan S. Vazhkudai |
SC | 2 |
| 2015 | Understanding I/O workload characteristics of a Peta-scale storage system
Youngjae Kim 0001, Raghul Gunasekaran |
J. Supercomput. | 2 |
| 2014 | Automatic identification of application I/O signatures from noisy server-side traces
Yang Liu 0129, Raghul Gunasekaran, Xiaosong Ma, Sudharshan S. Vazhkudai |
FAST | 2 |
| 2014 | Best Practices and Lessons Learned from Deploying and Operating Large-Scale Data-Centric Parallel File SystemsabstractThe Oak Ridge Leadership Computing Facility (OLCF) has deployed multiple large-scale parallel file systems (PFS) to support its operations. During this process, OLCF acquired significant expertise in large-scale storage system design, file system software development, technology evaluation, benchmarking, procurement, deployment, and operational practices. Based on the lessons learned from each new PFS deployment, OLCF improved its operating procedures, and strategies. This paper provides an account of our experience and lessons learned in acquiring, deploying, and operating large-scale parallel file systems. We believe that these lessons will be useful to the wider HPC community. Sarp Oral, James Simmons, Jason Hill, Dustin Leverman, Feiyi Wang, Matthew Ezell, Ross G. Miller, Douglas Fuller, Raghul Gunasekaran, Youngjae Kim 0001, Saurabh Gupta 0002, Devesh Tiwari, Sudharshan S. Vazhkudai, James H. Rogers, David Dillow, Galen M. Shipman, Arthur S. Bland |
SC | 9 |
| 2008 | Control-Based Real-Time Metadata Matching for Information DisseminationabstractReal-time information dissemination is of increasing importance to our society. Existing work mainly focuses on delivering information from sources to sinks in a timely manner based on established subscriptions, with the assumption that those subscriptions are persistent. However, the bottleneck of many real-time information dissemination systems is actually the matching process to continuously reevaluate such subscriptions between numerous sources and numerous sinks, in response to dynamically varying information attributes at runtime. In this paper, we propose a feedback controller to adaptively meet the response time constraints on metadata matching in an example information dissemination system. Our controller features a rigorous design based on well-established feedback control theory for guaranteed control accuracy and system stability. Empirical results on a physical test-bed demonstrate that our controller outperforms both an open-loop solution and a typical heuristic solution, by having more accurate control and better system quality of service. Ming Chen 0002, Raghul Gunasekaran, Hairong Qi 0001, Mallikarjun Shankar |
RTCSA | 3 |
| 2008 | XLRP: Cross Layer Routing Protocol for Wireless Sensor NetworksabstractThe developments in the field of wireless sensor networks (WSNs) have been accompanied by a paradigm shift from the layered protocol design to a cross layer design, which has shown its promise in effectively preserving energy, the most constraint resource in sensor networks. The proposed cross layer protocol, XLRP, explores an efficient routing strategy based on the application layer information along with the capabilities of the physical layer. Protocol design in wireless sensor networks have always resorted to low transmitter power levels as an energy efficient strategy with minimum interference. The proposed routing algorithm substantiates on switching transmission power levels based on the volume of the data being transmitted as an energy efficient mechanism. The cross layer design proposes a back-off mechanism by switching OFF unintended receivers based on the power of the received radio signal. In addition, the protocol reduces control message exchange by piggy-backing information and extracting information from packets received by unintended receivers. The effectiveness of the proposed protocol is demonstrated through simulation in ns2. The concept of operating the transmitter at various power levels is shown to be an energy efficient, scalable approach with a rational throughput. Raghul Gunasekaran, Hairong Qi 0001 |
WCNC | 1 |