Samir Gupta

dblp:24/6461 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 AVATAR: Autonomy Aware Routing for On-demand Transit Applications
abstract
Autonomous vehicles (AVs) are becoming integral to on-demand micro transit, offering the potential for safer, efficient, and sustainable transportation. However, AV deployment faces several challenges, including the lack of suitable roadways, varying travel conditions. Traditional routers prioritize speed and not reliability, leading to unpredictable operations and complications in planning. To address these, we introduce AVATAR, an autonomy-aware routing framework that prioritizes dependable, low-variance routes. Our approach encodes multiple objectives including road speed, speed variability, zoning areas, pedestrian encounters, and operator preferred roadways into edge-level routing engines. Objective optimized routes are generated, then scored using a multi-criteria decision-making process. User-configurable preference profiles, allow operators to define a balance between reliability and speed. AVATAR is a data-driven framework that supports both real-time AV operations and offline analysis, enabling transit operators to assess and refine routing strategies. Our experiments using real-world data from Silicon Valley, California, and Yokohama, Japan show that our approach significantly improves AV reliability and performance and advances the sustainable and scalable integration of AVs into future transportation networks.
David Rogers, Samir Gupta, Jose Paolo Talusan, Ammar Bin Zulqarnain, Mirza Baig, Arti Ramesh, Natsu Takahashi, Naoki Kojo, Abhishek Dubey
SMARTCOMP2
2024 PhD Forum: Towards Efficient Urban Mobility: Leveraging GNN and MTL for Demand Forecasting
abstract
Efficient public transit is essential for urban sustainability, helping to reduce carbon emissions and alleviate congestion. However, challenges such as overcrowding and delays degrade service quality and discourage ridership. Existing studies in predicting public transit ridership consider only a static depiction of bus networks or typically analyze occupancy and delays independently, missing the correlation between these factors. Our experiments with real-world data from Chattanooga and Nashville, Tennessee, demonstrate that our approach, which uses Graph Neural Networks (GNN) and Multi-Task Learning for stop-level day-ahead and same-day demand prediction, respectively, outperforms state-of-the-art models in terms of accuracy and robustness.
Samir Gupta
SMARTCOMP1
2024 A Graph Neural Network Framework for Imbalanced Bus Ridership Forecasting
abstract
Public transit systems are paramount in lowering carbon emissions and reducing urban congestion for environmental sustainability. However, overcrowding has adverse effects on the quality of service, passenger experience, and overall efficiency of public transit causing a decline in the usage of public transit systems. Therefore, it is crucial to identify and forecast potential windows of overcrowding to improve passenger experience and encourage higher ridership. Predicting ridership is a complex task, due to the inherent noise of collected data and the sparsity of overcrowding events. Existing studies in predicting public transit ridership consider only a static depiction of bus networks. We address these issues by first applying a data processing pipeline that cleans noisy data and engineers several features for training. Then, we address sparsity by converting the network to a dynamic graph and using a graph convolutional network, incorporating temporal, spatial, and auto-regressive features, to learn generalizable patterns for each route. Finally, since conventional loss functions like categorical cross-entropy have limitations in addressing class imbalance inherent in ridership data, our proposed approach uses focal loss to refine the prediction focus on less frequent yet task-critical overcrowding instances. Our experiments, using real-world data from our partner agency, show that the proposed approach outperforms existing state-of-the-art baselines in terms of accuracy and robustness.
Samir Gupta, Agrima Khanna, Jose Paolo Talusan, Anwar Said, Daniel Freudberg, Ayan Mukhopadhyay, Abhishek Dubey
SMARTCOMP1
2023 Addressing APC Data Sparsity in Predicting Occupancy and Delay of Transit Buses: A Multitask Learning Approach
abstract
Public transit is a vital mode of transportation in urban areas, and its efficiency is crucial for the daily commute of millions of people. To improve the reliability and predictability of transit systems, researchers have developed separate single-task learning models to predict the occupancy and delay of buses at the stop or route level. However, these models provide a narrow view of delay and occupancy at each stop and do not account for the correlation between the two. We propose a novel approach that leverages broader generalizable patterns governing delay and occupancy for improved prediction. We introduce a multitask learning toolchain that takes into account General Transit Feed Specification feeds, Automatic Passenger Counter data, and contextual temporal and spatial information. The toolchain predicts transit delay and occupancy at the stop level, improving the accuracy of the predictions of these two features of a trip given sparse and noisy data. We also show that our toolchain can adapt to fewer samples of new transit data once it has been trained on previous routes/trips as compared to state-of-the-art methods. Finally, we use actual data from Chattanooga, Tennessee, to validate our approach. We compare our approach against the state-of-the-art methods and we show that treating occupancy and delay as related problems improves the accuracy of the predictions. We show that our approach improves delay prediction significantly by as much as 4% in F1 scores while producing equivalent or better results for occupancy.
Ammar Bin Zulqarnain, Samir Gupta, Jose Paolo Talusan, Daniel Freudberg, Philip Pugliese, Ayan Mukhopadhyay, Abhishek Dubey
SMARTCOMP2
2021 A strategy for validation of variables derived from large-scale electronic health record data
abstract
PURPOSE: Standardized approaches for rigorous validation of phenotyping from large-scale electronic health record (EHR) data have not been widely reported. We proposed a methodologically rigorous and efficient approach to guide such validation, including strategies for sampling cases and controls, determining sample sizes, estimating algorithm performance, and terminating the validation process, hereafter referred to as the San Diego Approach to Variable Validation (SDAVV). METHODS: We propose sample size formulae which should be used prior to chart review, based on pre-specified critical lower bounds for positive predictive value (PPV) and negative predictive value (NPV). We also propose a stepwise strategy for iterative algorithm development/validation cycles, updating sample sizes for data abstraction until both PPV and NPV achieve target performance. RESULTS: We applied the SDAVV to a Department of Veterans Affairs study in which we created two phenotyping algorithms, one for distinguishing normal colonoscopy cases from abnormal colonoscopy controls and one for identifying aspirin exposure. Estimated PPV and NPV both reached 0.970 with a 95% confidence lower bound of 0.915, estimated sensitivity was 0.963 and specificity was 0.975 for identifying normal colonoscopy cases. The phenotyping algorithm for identifying aspirin exposure reached a PPV of 0.990 (a 95% lower bound of 0.950), an NPV of 0.980 (a 95% lower bound of 0.930), and sensitivity and specificity were 0.960 and 1.000. CONCLUSIONS: A structured approach for prospectively developing and validating phenotyping algorithms from large-scale EHR data can be successfully implemented, and should be considered to improve the quality of "big data" research.
Ranier Bustamante, Ashley Earles, Joshua Demb, Karen Messer, Samir Gupta
J. Biomed. Informatics6
2020 Automated Identification of Patients with Immune-related Adverse Events from Clinical Notes using Machine Learning
Samir Gupta, Anas Belouali, Neil J. Shah, Michael Atkins, Subha Madhavan
AMIA1
2020 Linking Polyps to Jars: information loss across colonoscopy and pathology reports
Olga V. Patterson, Samir Gupta, Andrew Gawron, Tonya Kaltenbach, Ranier Bustamante, Daniel W. Denhalter, Ashley Earles, Scott L. DuVall
AMIA2
2020 A system uptake analysis and GUIDES checklist evaluation of the Electronic Asthma Management System: A point-of-care computerized clinical decision support system
abstract
OBJECTIVE: Computerized clinical decision support systems (CCDSSs) promise improvements in care quality; however, uptake is often suboptimal. We sought to characterize system use, its predictors, and user feedback for the Electronic Asthma Management System (eAMS)-an electronic medical record system-integrated, point-of-care CCDSS for asthma-and applied the GUIDES checklist as a framework to identify areas for improvement. MATERIALS AND METHODS: The eAMS was tested in a 1-year prospective cohort study across 3 Ontario primary care sites. We recorded system usage by clinicians and patient characteristics through system logs and chart reviews. We created multivariable models to identify predictors of (1) CCDSS opening and (2) creation of a self-management asthma action plan (AAP) (final CCDSS step). Electronic questionnaires captured user feedback. RESULTS: Over 1 year, 490 asthma patients saw 121 clinicians. The CCDSS was opened in 205 of 1033 (19.8%) visits and an AAP created in 121 of 1033 (11.7%) visits. Multivariable predictors of opening the CCDSS and producing an AAP included clinic site, having physician-diagnosed asthma, and presenting with an asthma- or respiratory-related complaint. The system usability scale score was 66.3 ± 16.5 (maximum 100). Reported usage barriers included time and system accessibility. DISCUSSION: The eAMS was used in a minority of asthma patient visits. Varying workflows and cultures across clinics, physician beliefs regarding asthma diagnosis, and relevance of the clinical complaint influenced uptake. CONCLUSIONS: Considering our findings in the context of the GUIDES checklist helped to identify improvements to drive uptake and provides lessons relevant to CCDSS design across diseases.
Jeffrey Lam Shin Cheung, Natalie Paolucci, Courtney Price, Jenna Sykes, Samir Gupta
J. Am. Medical Informatics Assoc.5
2019 OncoMX: an Integrated Cancer Mutation and Expression Knowledgebase for Biomarker Evaluation and Discovery
Amanda Bell, Raja Mazumder, Daniel J. Crichton, Vijay K. Shanker, Frederic B. Bastian, Hayley Dingerdissen, Samir Gupta, Evan Holmes, Robel Y. Kahsay, Heather Kincaid, A. S. M. Ashique Mahmood, Marc Robinson-Rechavi, Stephanie S. Singleton
AMIA7
2016 Development of the Parkland-UT Southwestern Colonoscopy Reporting System (CoRS) for evidence-based colon cancer surveillance recommendations
abstract
OBJECTIVE: Through colonoscopy, polyps can be identified and removed to reduce colorectal cancer incidence and mortality. Appropriate use of surveillance colonoscopy, post polypectomy, is a focus of healthcare reform. MATERIALS AND METHODS: The authors developed and implemented the first electronic medical record-based colonoscopy reporting system (CoRS) that matches endoscopic findings with guideline-consistent surveillance recommendations and generates tailored results and recommendation letters for patients and providers. RESULTS: In its first year, CoRS was used in 98.6% of indicated cases. Via a survey, colonoscopists agreed/strongly agreed it is easy to use (83%), provides guideline-based recommendations (89%), improves quality of Spanish letters (94%), they would recommend it for other institutions (78%), and it made their work easier (61%), and led to improved practice (56%). DISCUSSION: CoRS' widespread adoption and acceptance likely resulted from stakeholder engagement throughout the development and implementation process. CONCLUSION: CoRS is well-accepted by clinicians and provides guideline-based recommendations and results communications to patients and providers.
Celette Sugg Skinner, Samir Gupta, Ethan A. Halm, Shaun Wright, Katharine McCallister, Wendy Bishop, Noel Santini, Christian Mayorga, Brett A. Moran, Joanne M. Sanders, Amit G. Singal
J. Am. Medical Informatics Assoc.2
2013 Part-of-speech tagging of program identifiers for improved text-based software engineering tools
abstract
To aid program comprehension, programmers choose identifiers for methods, classes, fields and other program elements primarily by following naming conventions in software. These software “naming conventions” follow systematic patterns which can convey deep natural language clues that can be leveraged by software engineering tools. For example, they can be used to increase the accuracy of software search tools, improve the ability of program navigation tools to recommend related methods, and raise the accuracy of other program analyses. After splitting multi-word names into their component words, the next step to extracting accurate natural language information is tagging each word with its part of speech (POS) and then chunking the name into natural language phrases. State-of-theart approaches, most of which rely on “traditional POS taggers” trained on natural language documents, do not capture the syntactic structure of program elements. In this paper, we present a POS tagger and syntactic chunker for source code names that takes into account programmers' naming conventions to understand the regular, systematic ways a program element is named. We studied the naming conventions used in Object Oriented Programming and identified different grammatical constructions that characterize a large number of program identifiers. This study then informed the design of our POS tagger and chunker. Our evaluation results show a significant improvement in accuracy(11%-20%) of POS tagging of identifiers, over the current approaches. With this improved accuracy, both automated software engineering tools and developers will be able to better capture and understand the information available in code.
Samir Gupta, Sana Malik, Lori L. Pollock, K. Vijay-Shanker
ICPC1
2013 Automatically mining software-based, semantically-similar words from comment-code mappings
abstract
Many software development and maintenance tools involve matching between natural language words in different software artifacts (e.g., traceability) or between queries submitted by a user and software artifacts (e.g., code search). Because different people likely created the queries and various artifacts, the effectiveness of these tools is often improved by expanding queries and adding related words to textual artifact representations. Synonyms are particularly useful to overcome the mismatch in vocabularies, as well as other word relations that indicate semantic similarity. However, experience shows that many words are semantically similar in computer science situations, but not in typical natural language documents. In this paper, we present an automatic technique to mine semantically similar words, particularly in the software context. We leverage the role of leading comments for methods and programmer conventions in writing them. Our evaluation of our mined related comment-code word mappings that do not already occur in WordNet are indeed viewed as computer science, semantically-similar word pairs in high proportions.
Samir Gupta, Lori L. Pollock, K. Vijay-Shanker
MSR2