Samarth Swarup

dblp:21/4106 · DBLP profile ↗
← Back
8ranked-venue papers in the field
0as first author
4since 2021 · last 2025
0000-0003-3615-1663ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 5Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 IrrMap: A Large-Scale Comprehensive Dataset for Irrigation Method Mapping
abstract
We introduce IrrMap, the first large-scale dataset (1.1 million patches) for irrigation method mapping across regions. IrrMap consists of multi-resolution satellite imagery from LandSat and Sentinel, along with key auxiliary data such as crop type, land use, and vegetation indices. The dataset spans 1,668,899 farms and 11,443,492 acres across multiple western U.S. states from 2013 to 2023, providing a rich and diverse foundation for irrigation analysis and ensuring geospatial alignment and quality control. The dataset is ML-ready, with standardized 224×224 GeoTIFF patches, the multiple input modalities, carefully chosen train-test-split data, and accompanying dataloaders for seamless deep learning model training and benchmarking in irrigation mapping. The dataset is also accompanied by a complete pipeline for dataset generation, enabling researchers to extend IrrMap to new regions for irrigation data collection or adapt it with minimal effort for other similar applications in agricultural and geospatial analysis. We also analyze the irrigation method distribution across crop groups, spatial irrigation patterns (using Shannon diversity indices), and irrigated area variations for both LandSat and Sentinel, providing insights into regional and resolution-based differences. To promote further exploration, we openly release IrrMap, along with the derived datasets, benchmark models, and pipeline code, through a GitHub repository: https://github.com/Nibir088/IrrMap and Data repository: https://huggingface.co/Nibir/IrrMap, providing comprehensive documentation and implementation details.
Nibir Chandra Mandal, Oishee Bintey Hoque, Abhijin Adiga, Samarth Swarup, Mandy L. Wilson, Lu Feng 0001, Yangfeng Ji, Miaomiao Zhang 0002, Geoffrey C. Fox, Madhav V. Marathe
KDD (2)4
2024 A Scalable Game-theoretic Approach to Urban Evacuation Routing and Scheduling
abstract
Evacuation planning is an essential part of disaster management where the goal is to relocate people under imminent danger to safety. However, finding jointly optimal evacuation routes and a schedule that minimizes the average evacuation time or evacuation completion time, is a computationally hard problem. As a result, large-scale evacuation routing and scheduling continues to be a challenge. In this paper, we present a game-theoretic approach to tackle this problem. We start by formulating a strategic routing and scheduling game, named the Evacuation Game: Routing and Scheduling (EGRES), where players choose their route and time of departure. We show that: (i) every instance of EGRES has at least one pure strategy Nash equilibrium, and (ii) an optimal outcome in an instance will always be an equilibrium in that instance. We then provide bounds on how bad an equilibrium can be compared to an optimal outcome. Additionally, we present a polynomial-time algorithm, the Sequential Action Algorithm (SAA), for finding equilibria in a given instance under a special condition. We use Virginia Beach City in Virginia, and Harris County in Houston, Texas as study areas and construct two EGRES instances. Our results show that, by utilizing SAA, we can efficiently find equilibria in these instances that have social objective close to the optimal value.
Kazi Ashik Islam, Da Qi Chen, Madhav V. Marathe, Henning S. Mortveit, Samarth Swarup, Anil Vullikanti
IEEE Big Data5
2022 Fidelity and diversity metrics for validating hierarchical synthetic data: Application to residential energy demand
abstract
Synthetic data is gaining rapid importance in many application domains due to privacy issues, bias, lack, or simply unavailability of real data. It is important that the synthetic data be a good representation of real data for successfully completing the task at hand. Thus, devising characteristic validation metrics is crucial and remains an open problem in many domains (e.g., image generation). Good validation metrics must be able to disentangle the differences between the quality and the variability coverage of the synthetic data. We propose to use a 3-dimensional metric (precision α, recall β, coverage γ) to describe the fidelity and diversity of the synthetic data. In this paper, we improve on existing definitions of precision, recall, and coverage to extend to large scale time series data. Traditional nearest neighbor manifolds from the literature are replaced by unsupervised learning techniques such as clustering to deal with large scale fine resolution time series while computing the validation metrics. The proposed metrics are employed to validate synthetic data in the domain of residential energy demand. In addition, we extend these definitions to datasets that have a natural hierarchical structure. We propose a hierarchical data-tree model in which precision, recall, and coverage can be computed at multiple inherent (and/or custom) levels of groupings of the data.
Swapna Thorve, Anil Vullikanti, Henning S. Mortveit, Samarth Swarup, Madhav V. Marathe
IEEE Big Data4
2022 Incorporating Fairness in Large-scale Evacuation Planning
abstract
Evacuation planning is an essential part of disaster management where the goal is to relocate people in a safe and orderly manner. Existing research has shown that such problems are hard to approximate and current methods are difficult to scale to real-life applications. We introduce a notion of fairness and two related objectives while studying evacuation planning, namely: minimizing maximum inconvenience and minimizing average inconvenience. We show that both problems are not just NP-hard to solve exactly, but in fact are NP-hard to approximate. On the positive side, we present a heuristic optimization method MIP-LNS, based on the well-known Large Neighborhood Search framework, that can find good approximate solutions in reasonable amount of time. We also consider a multi-objective problem where the goal is to minimize both objectives and solve it using MIP-LNS. We use real-world road network and population data from Harris County in Houston, Texas (a region that needed large-scale evacuations in the past), and apply MIP-LNS to calculate evacuation plans for the area. We compare the quality of the plans in terms of evacuation efficiency and fairness. We find that the solutions to the multi-objective problem are superior in both of these aspects. We also perform statistical tests to show that the solutions are significantly different.
Kazi Ashik Islam, Da Qi Chen, Madhav V. Marathe, Henning S. Mortveit, Samarth Swarup, Anil Vullikanti
CIKM5
2020 A Simulation-based Approach for Large-scale Evacuation Planning
abstract
Evacuation planning methods aim to design routes and schedules to relocate people to safety in the event of natural or man-made disasters. The primary goal is to minimize casualties which often requires the evacuation process to be completed as soon as possible. In this paper, we present QueST, an agent-based discrete event queuing network simulation system, and STEERS, an iterative routing algorithm that uses QueST for designing and evaluating large scale evacuation plans in terms of total egress time and congestion/bottlenecks occurring during evacuation. We use the Houston Metropolitan Area, which consists of nine US counties and spans an area of 9,444 square miles as a case study, and compare the performance of STEERS with two existing route planning methods. We find that STEERS is either better or comparable to these methods in terms of total evacuation time and congestion faced by the evacuees. We also analyze the large volume of data generated by the simulation process to gain insights about the scenarios arising from following the evacuation routes prescribed by these methods.
Kazi Ashik Islam, Madhav V. Marathe, Henning S. Mortveit, Samarth Swarup, Anil Vullikanti
IEEE BigData4
2020 Creating Realistic Power Distribution Networks using Interdependent Road Infrastructure
abstract
It is well known that physical interdependencies exist between networked civil infrastructures such as transportation and power system networks. In order to analyze complex nonlinear correlations between such networks, datasets pertaining to such real infrastructures are required. However, such data are not readily available due to their proprietary nature. This work proposes a methodology to generate realistic synthetic power distribution networks for a given geographical region. A network generated in this manner is not the actual distribution system, but its functionality is very similar to the real distribution network. The synthetic network connects high voltage substations to individual residential consumers through primary and secondary distribution networks. Here, the distribution network is generated by solving an optimization problem which minimizes the overall length of the network subject to structural and power flow constraints. This work also incorporates identification of long high voltage feeders originating from substations and connecting remotely situated customers in rural geographic locations while maintaining voltage regulation within acceptable limits. The proposed methodology is applied to the state of Virginia and creates synthetic distribution networks which are validated by comparing them to actual power distribution networks at the same location.
Rounak Meyur, Madhav V. Marathe, Anil Vullikanti, Henning S. Mortveit, Samarth Swarup, Virgilio Centeno, Arun G. Phadke
IEEE BigData5
2018 An Empirical Assessment of the Complexity and Realism of Synthetic Social Contact Networks*
abstract
We use multiple measures of graph complexity to evaluate the realism of synthetically-generated networks of human activity, in comparison with several stylized network models as well as a collection of empirical networks from the literature. The synthetic networks are generated by integrating data about human populations from several sources, including the Census, transportation surveys, and geographical data. The resulting networks represent an approximation of daily or weekly human interaction. Our results indicate that the synthetically generated graphs according to our methodology are closer to the real world graphs, as measured across multiple structural measures, than a range of stylized graphs generated using common network models from the literature.
Kiran Karra, Samarth Swarup, Justus Graham
IEEE BigData2
2013 Blocking Simple and Complex Contagion by Edge Removal
abstract
Eliminating interactions among individuals is an important means of blocking contagion spread, e.g., closing schools during an epidemic or shutting down electronic communication channels during social unrest. We study contagion blocking in networked populations by identifying edges to remove from a network, thus blocking contagion transmission pathways. We formulate various problems to minimize contagion spread and show that some are efficiently solvable while others are formally hard. We also compare our hardness results to those from node blocking problems and show interesting differences between the two. Our main problem is not only hard, but also has no approximation guarantee, unless P=NP. Therefore, we devise a heuristic for the problem and compare its performance to state-of-the-art heuristics from the literature. We show, through results of 12 (network, heuristic) combinations on three real social networks, that our method offers considerable improvement in the ability to block contagions in weighted and unweighted networks. We also conduct a parametric study to understand the limitations of our approach.
Chris J. Kuhlman, Gaurav Tuli, Samarth Swarup, Madhav V. Marathe, S. S. Ravi
ICDM3