Mallikarjun Shankar

dblp:66/1246 · also Mallikarjun Arjun Shankar · DBLP profile ↗
← Back
24ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-5289-7460ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 4 since 2021Computer networks · 8Databases, data management, data science and information retrieval · 5 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Mixed-precision numerics in scientific applications: survey and perspectives
abstract
Abstract The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8 $$\times$$ × compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.
Aditya Kashi, Hao Lu 0001, Wesley Brewer, David Rogers, Michael A. Matheson, Mallikarjun Shankar, Feiyi Wang
J. Supercomput.6
2025 RingX: Scalable Parallel Attention for Long-Context Learning on HPC
abstract
The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.
Junqi Yin, Mijanur Palash, Mallikarjun Shankar, Feiyi Wang
SC3
2023 DeepThermo: Deep Learning Accelerated Parallel Monte Carlo Sampling for Thermodynamics Evaluation of High Entropy Alloys
abstract
Since the introduction of Metropolis Monte Carlo (MC) sampling, it and its variants have become standard tools used for thermodynamics evaluations of physical systems. However, a long-standing problem that hinders the effectiveness and efficiency of MC sampling is the lack of a generic method (a.k.a. MC proposal) to update the system configurations. Consequently, current practices are not scalable. Here we propose a parallel MC sampling framework for thermodynamics evaluation—DeepThermo. By using deep learning–based MC proposals that can globally update the system configurations, we show that DeepThermo can effectively evaluate the phase transition behaviors of high entropy alloys, which have an astronomical configuration space. For the first time, we directly evaluate a density of states expanding over a range of ~e10,000for a real material. We also demonstrate DeepThermo’s performance and scalability up to 3,000 GPUs on both NVIDIA V100 and AMD MI250X-based supercomputers.
Junqi Yin, Feiyi Wang, Mallikarjun Shankar
IPDPS3
2023 FORGE: Pre-Training Open Foundation Models for Science
abstract
Large language models (LLMs) are poised to revolutionize the way we conduct scientific research. However, both model complexity and pre-training cost are impeding effective adoption for the wider science community. Identifying suitable scientific use cases, finding the optimal balance between model and data sizes, and scaling up model training are among the most pressing issues that need to be addressed. In this study, we provide practical solutions for building and using LLM-based foundation models targeting scientific research use cases. We present an end-to-end examination of the effectiveness of LLMs in scientific research, including their scaling behavior and computational requirements on Frontier, the first Exascale supercomputer. We have also developed for release to the scientific community a suite of open foundation models called FORGE with up to 26B parameters using 257B tokens from over 200M scientific articles, with performance either on par or superior to other state-of-the-art comparable models. We have demonstrated the use and effectiveness of FORGE on scientific downstream tasks. Our research establishes best practices that can be applied across various fields to take advantage of LLMs for scientific discovery.
Junqi Yin, Sajal Dash, Feiyi Wang, Mallikarjun Shankar
SC4
2021 Enabling discovery data science through cross-facility workflows
abstract
Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.
Katie Antypas, Deborah Bard, Johannes P. Blaschke, Shane Canon, Bjoern Enders, Mallikarjun Shankar, Suhas Somnath, Dale Stansberry, Thomas D. Uram, Sean R. Wilkinson
IEEE BigData6
2020 GPU lifetimes on titan supercomputer: survival analysis and reliability
abstract
The Cray XK7 Titan was the top supercomputer system in the world for a long time and remained critically important throughout its nearly seven year life. It was an interesting machine from a reliability viewpoint as most of its power came from 18,688 GPUs whose operation was forced to execute three rework cycles, two on the GPU mechanical assembly and one on the GPU circuitboards. We write about the last rework cycle and a reliability analysis of over 100,000 years of GPU lifetimes during Titan’s 6-year-long productive period. Using time between failures analysis and statistical survival analysis techniques, we find that GPU reliability is dependent on heat dissipation to an extent that strongly correlates with detailed nuances of the cooling architecture and job scheduling. We describe the history, data collection, cleaning, and analysis and give recommendations for future supercomputing systems. We make the data and our analysis codes publicly available.
George Ostrouchov, Don E. Maxwell, Rizwan A. Ashraf, Christian Engelmann, Mallikarjun Shankar, James H. Rogers
SC5
2018 The design, deployment, and evaluation of the CORAL pre-exascale systems
Sudharshan S. Vazhkudai, Bronis R. de Supinski, Arthur S. Bland, Al Geist, James C. Sexton, James A. Kahle, Christopher Zimmer 0001, Scott Atchley, Sarp Oral, Don E. Maxwell, Verónica G. Vergara Larrea, Adam Bertsch, Robin Goldstone, Wayne Joubert, Christopher M. Chambreau, David Appelhans, Robert Blackmore, Ben Casses, George Chochia, Gene Davison, Matthew Ezell, Thomas Gooding, Elsa Gonsiorowski, Leopold Grinberg, Bill Hanson, Bill Hartner, Ian Karlin, Matthew L. Leininger, Dustin Leverman, Chris Marroquin, Adam Moody, Martin Ohmacht, Ramesh Pankajakshan, Fernando Pizzano, James H. Rogers, Bryan S. Rosenburg, Drew Schmidt, Mallikarjun Shankar, Feiyi Wang, Py Watson, Bob Walkup, Lance D. Weems, Junqi Yin
SC38
2018 Best effort broadcast under cascading failures in interdependent critical infrastructure networks
Sisi Duan, Sankeun Lee 0001, Supriya Chinthavali, Mallikarjun Shankar
Pervasive Mob. Comput.4
2016 URBAN-NET: A network-based infrastructure monitoring and analysis system for emergency management and public safety
abstract
Critical Infrastructures (CIs) such as energy, water, and transportation are complex networks that are crucial for sustaining day-to-day commodity flows vital to national security, economic stability, and public safety. The nature of these CIs is such that failures caused by an extreme weather event or a man-made incident can trigger widespread cascading failures, sending ripple effects at regional or even national scales. To minimize such effects, it is critical for emergency responders to identify existing or potential vulnerabilities within CIs during such stressor events in a systematic and quantifiable manner and take appropriate mitigating actions. We present here a novel critical infrastructure monitoring and analysis system named URBAN-NET. The system includes a software stack and tools for monitoring CIs, pre-processing data, interconnecting multiple CI datasets as a heterogeneous network, identifying vulnerabilities through graph-based topological analysis, and predicting consequences based on “what-if” simulations along with visualization. As a proof-of-concept, we present several case studies to show the capabilities of our system. We also discuss remaining challenges and future work.
Sangkeun Matt Lee, Liangzhe Chen, Sisi Duan, Supriya Chinthavali, Mallikarjun Shankar, B. Aditya Prakash
IEEE BigData5
2015 Enabling graph appliance for genome assembly
abstract
In recent years, there has been a huge growth in the amount of genomic data available as reads generated from various genome sequencers. The number of reads generated can be huge, ranging from hundreds to billions of nucleotide, each varying in size. Assembling such large amounts of data is one of the challenging computational problems for both biomedical and data scientists. Most of the genome assemblers that have developed use de Bruijn graph techniques. A de Bruijn graph represents a collection of read sequences by billions of vertices and edges, which require large amounts of memory and computational power to store and process. This is the major drawback to de Bruijn graph assembly. Massively parallel, multithreaded, shared memory systems can be leveraged to overcome some of these issues. The objective of our research is to investigate the feasibility and scalability issues of de Bruijn graph assembly on Cray's Urika-GD system; Urika-GD is a high performance graph appliance with a large shared memory and massively multithreaded custom processor designed for executing SPARQL queries over large-scale RDF data sets. However, to the best of our knowledge, there is no research on representing a de Bruijn graph as an RDF graph or finding Eulerian paths in RDF graphs using SPARQL for potential genome discovery. In this paper, we address the issues involved in representing de Bruin graphs as RDF graphs and propose an iterative querying approach for searching cycles to find Eulerian paths in large RDF graphs. We evaluate the performance of our implementation on real world ebola genome datasets and illustrate how genome assembly can be accomplished with Urika-GD using iterative SPARQL queries.
Rina Singh, Jeffrey A. Graves, Sankeun Lee 0001, Sreenivas R. Sukumar 0001, Mallikarjun Shankar
IEEE BigData5
2012 Adaptive response time control for metadata matching in information dissemination systems
Ming Chen 0002, Hairong Qi 0001, Mallikarjun Shankar
J. Syst. Archit.4
2010 Identification of low-level point radioactive sources using a sensor network
abstract
Identification of a low-level point radioactive source amidst background radiation is achieved by a network of radiation sensors using a two-step approach. Based on measurements from three or more sensors, a geometric difference triangulation method or an N -sensor localization method is used to estimate the location and strength of the source. Then a sequential probability ratio test based on current measurements and estimated parameters is employed to finally decide: (1) the presence of a source with the estimated parameters, or (2) the absence of the source, or (3) the insufficiency of measurements to make a decision. This method achieves specified levels of false alarm and missed detection probabilities, while ensuring a close-to-minimal number of measurements for reaching a decision. This method minimizes the ghost-source problem of current estimation methods, and achieves a lower false alarm rate compared with current detection methods. This method is tested and demonstrated using: (1) simulations, and (2) a test-bed that utilizes the scaling properties of point radioactive sources to emulate high intensity ones that cannot be easily and safely handled in laboratory experiments.
Jren-Chit Chin, Nageswara S. V. Rao, David K. Y. Yau, Mallikarjun Shankar, Yong Yang 0009, Jennifer C. Hou, Srinivasagopalan Srivathsan, S. Sitharama Iyengar
ACM Trans. Sens. Networks4
2010 Quality of monitoring of stochastic events by periodic and proportional-share scheduling of sensor coverage
abstract
We analyze the quality of monitoring (QoM) of stochastic events by a periodic sensor which monitors a point of interest (PoI) for q time every p time. We show how the amount of information captured at a PoI is affected by the proportion q/p , the time interval p over which the proportion is achieved, the event type in terms of its stochastic arrival dynamics and staying times and the utility function. The periodic PoI sensor schedule happens in two broad contexts. In the case of static sensors, a sensor monitoring a PoI may be periodically turned off to conserve energy, thereby extending the lifetime of the monitoring until the sensor can be recharged or replaced. In the case of mobile sensors, a sensor may move between the PoIs in a repeating visit schedule. In this case, the PoIs may vary in importance, and the scheduling objective is to distribute the sensor's coverage time in proportion to the importance levels of the PoIs. Based on our QoM analysis, we optimize a class of periodic mobile coverage schedules that can achieve such proportional sharing while maximizing the QoM of the total system.
David K. Y. Yau, Nung Kwan Yip, Chris Y. T. Ma, Nageswara S. V. Rao, Mallikarjun Shankar
ACM Trans. Sens. Networks5
2009 Improved SPRT detection using localization with application to radiation sources
Nageswara S. V. Rao, Charles W. Glover, Mallikarjun Shankar, Jren-Chit Chin, David K. Y. Yau, Chris Y. T. Ma, Yong Yang 0009, Sartaj Sahni
FUSION3
2009 Sensor Placement for Detecting Propagative Sources in Populated Environments
abstract
We consider the placement of sensors to detect propagative sources where the sensing area of each sensor is anisotropic and arbitrarily-shaped due to the terrain and meteorological conditions. The propagation and detection times are non-negligible due to the propagation of source effects through space at a slow speed. We formulate the problem as placing the minimum number of sensors to ensure a detection time T and the coverage utility C. Both the sensing areas of sensors and the utility function U(ldr) are chosen to capture the environmental factors and the population distribution. We show this problem to be NP-hard, and present heuristic algorithms for 1-coverage and fc-coverage by adopting exiting methods. We evaluate the proposed algorithms in the realistic setting of Port of Memphis where the objective is to protect the population against chemical leaks or attacks. We utilize the SCIPUFF dispersion model to determine the sensing areas by accounting for the terrain and meteorological conditions, and use the real-life population distribution as the utility function. Based on empirical study, we make several important observations.
Yong Yang 0009, I-Hong Hou, Jennifer C. Hou, Mallikarjun Shankar, Nageswara S. V. Rao
INFOCOM4
2009 Matching and Fairness in Threat-Based Mobile Sensor Coverage
abstract
Mobile sensors can be used to effect complete coverage of a surveillance area for a given threat over time, thereby reducing the number of sensors necessary. The surveillance area may have a given threat profile as determined by the kind of threat, and accompanying meteorological, environmental, and human factors. In planning the movement of sensors, areas that are deemed higher threat should receive proportionately higher coverage. We propose a coverage algorithm for mobile sensors to achieve a coverage that will match - over the long term and as quantified by an RMSE metric - a given threat profile. Moreover, the algorithm has the following desirable properties: 1) stochastic, so that it is robust to contingencies and makes it hard for an adversary to anticipate the sensor's movement, 2) efficient, and 3) practical, by avoiding movement over inaccessible areas. Further to matching, we argue that a fairness measure of performance over the shorter time scale is also important. We show that the RMSE and fairness are, in general, antagonistic, and argue for the need of a combined measure of performance, which we call efficacy. We show how a pause time parameter of the coverage algorithm can be used to control the trade-off between the RMSE and fairness, and present an efficient offline algorithm to determine the optimal pause time maximizing the efficacy. Finally, we discuss the effects of multiple sensors, under both independent and coordinated operation. Extensive simulation results - under realistic coverage scenarios - are presented for performance evaluation.
Chris Y. T. Ma, David K. Y. Yau, Jren-Chit Chin, Nageswara S. V. Rao, Mallikarjun Shankar
IEEE Trans. Mob. Comput.5
2008 Quality of monitoring of stochastic events by periodic & proportional-share scheduling of sensor coverage
abstract
We analyze the quality of monitoring (QoM) of stochastic events by a periodic sensor which monitors a point of interest (PoI) for q time every p time. We show how the amount of information captured at a PoI is affected by the proportion q/p, the time interval p over which the proportion is achieved, the event type, and the stochastic event arrival dynamics and staying times. The periodic PoI sensor schedule happens in two broad contexts. In the case of static sensors, a sensor monitoring a PoI may be periodically turned off to conserve energy, thereby extending the lifetime of the monitoring until the sensor can be recharged or replaced. In the case of mobile sensors, a sensor may move between the PoIs in a repeating visit schedule. In this case, the PoIs may vary in importance, and the scheduling objective is to distribute the sensor's coverage time in proportion to the importance levels of the PoIs. Based on our QoM analysis, we optimize a class of periodic mobile coverage schedules that can achieve such proportional sharing while maximizing the QoM of the total system.
David K. Y. Yau, Nung Kwan Yip, Chris Y. T. Ma, Nageswara S. V. Rao, Mallikarjun Shankar
CoNEXT5
2008 Localization under random measurements with application to radiation sources
Nageswara S. V. Rao, Mallikarjun Shankar, Jren-Chit Chin, David K. Y. Yau, Chris Y. T. Ma, Yong Yang 0009, Jennifer C. Hou, Xiaochun Xu, Sartaj Sahni
FUSION2
2008 Identification of Low-Level Point Radiation Sources Using a Sensor Network
abstract
Identification of a low-level point radiation source amidst background radiation is achieved by a network of radiation sensors using a two-step approach. Based on measurements from three sensors, the geometric difference triangulation method is used to estimate the location and strength of the source. Then a sequential probability ratio test based on current measurements and estimated parameters is employed to finally decide: (1) the presence of a source with the estimated parameters, or (2) the absence of the source, or (3) the insufficiency of measurements to make a decision. This method achieves specified levels of false alarm and missed detection probabilities, while ensuring a close-to-minimal number of measurements for reaching a decision. This method minimizes the ghost-source problem of current estimation methods, and achieves a lower false alarm rate compared with current detection methods. This method is tested and demonstrated using: (1) simulations, and (2) a test-bed that utilizes the scaling properties of point radiation sources to emulate high intensity ones that cannot be easily and safely handled in laboratory experiments.
Nageswara S. V. Rao, Mallikarjun Shankar, Jren-Chit Chin, David K. Y. Yau, Srinivasagopalan Srivathsan, S. Sitharama Iyengar, Yong Yang 0009, Jennifer C. Hou
IPSN2
2008 Control-Based Real-Time Metadata Matching for Information Dissemination
abstract
Real-time information dissemination is of increasing importance to our society. Existing work mainly focuses on delivering information from sources to sinks in a timely manner based on established subscriptions, with the assumption that those subscriptions are persistent. However, the bottleneck of many real-time information dissemination systems is actually the matching process to continuously reevaluate such subscriptions between numerous sources and numerous sinks, in response to dynamically varying information attributes at runtime. In this paper, we propose a feedback controller to adaptively meet the response time constraints on metadata matching in an example information dissemination system. Our controller features a rigorous design based on well-established feedback control theory for guaranteed control accuracy and system stability. Empirical results on a physical test-bed demonstrate that our controller outperforms both an open-loop solution and a typical heuristic solution, by having more accurate control and better system quality of service.
Ming Chen 0002, Raghul Gunasekaran, Hairong Qi 0001, Mallikarjun Shankar
RTCSA5
2008 Accurate localization of low-level radioactive source under noise and measurement errors
abstract
The localization of a radioactive source can be solved in closed-form using 4 ideal sensors and the Apollonius circle in a noise- and error-free environment. When measurement errors and noise such as background radiation are considered, a larger number of sensors is needed to produce accurate results, particularly for extremely low source intensities. In this paper, we present an efficient fusion algorithm that can exploit measurements from n sensors to improve the localization accuracy, and show how the accuracy scales with n. We report testbed results for a 0.911 μCi source to illustrate the effectiveness of our algorithm, in particular performance comparisons with state-of-the-art fusion algorithms based on Mean of Estimates (MoE) and Maximum Likelihood Estimation (MLE). We show that ITP is more accurate than MoE, whereas the choice between ITP and MLE is generally a tradeoff between accuracy and run time efficiency. Higher-intensity radioactive sources are not safe for actual experiments. In this case, we present simulation results based on a validated simulation model. We show that a low-intensity 400 μCi source, similar to the radioactivity of a concealed dirty bomb, can be localized to within 32.5 m using a sensor density of about 1 per 1100 m 2 in a surveillance area.
Jren-Chit Chin, David K. Y. Yau, Nageswara S. V. Rao, Yong Yang 0009, Chris Y. T. Ma, Mallikarjun Shankar
SenSys6
2007 A sensor-cyber network testbed for plume detection, identification, and tracking
abstract
No abstract available.
Jren-Chit Chin, I-Hong Hou, Jennifer C. Hou, Chris Y. T. Ma, Nageswara S. V. Rao, Mohit Saxena, Mallikarjun Shankar, Yong Yang 0009, David K. Y. Yau
IPSN7
2005 SensorNet Operational Prototypes: Building Wide-Area Interoperable Sensor Networks - Extended Abstract
Mallikarjun Shankar, Bryan L. Gorman, Cyrus M. Smith
DCOSS1
1996 Algorithms and Optimality of Scheduling Soft Aperiodic Requests in Fixed-Priority Preemptive Systems
Too-Seng Tia, Jane W.-S. Liu, Mallikarjun Shankar
Real Time Syst.3