Karandeep Singh

dblp:11/5722 · also Karandeep Singh Brar · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 Self-Supervised Vision for Climate Downscaling
Karandeep Singh, Chaeyoon Jeong, Naufal Shidqi, Sungwon Park 0001, Arjun Nellikkattil, Elke Zeller, Meeyoung Cha
IJCAI1
2023 GraphFC: Customs Fraud Detection with Label Scarcity
abstract
Customs officials across the world encounter huge volumes of transactions. Associated with customs transactions is customs fraud-the intentional manipulation of goods declarations to avoid taxes and duties. Due to limited manpower, the customs offices can only manually inspect a small number of declarations, necessitating the automation of customs fraud detection by machine learning techniques. The limited availability of manually inspected ground truth data makes it essential for the ML approach to generalize well on unseen data. However, current customs fraud detection models are not well suited or designed for this setting. In this work, we propose GraphFC (Graph Neural networks for Customs Fraud), a model-agnostic, domain-specific, graph neural network based customs fraud detection model that is designed to work in a real-world setting with limited ground truth data. Extensive experimentation using real customs data from two countries demonstrates that GraphFC generalizes well over unseen data and outperforms various baselines and other models by a large margin.
Karandeep Singh, Yu-Che Tsai, Cheng-Te Li, Meeyoung Cha, Shou-De Lin
CIKM1
2023 Assessing the net benefit of machine learning models in the presence of resource constraints
abstract
OBJECTIVE: The objective of this study is to provide a method to calculate model performance measures in the presence of resource constraints, with a focus on net benefit (NB). MATERIALS AND METHODS: To quantify a model's clinical utility, the Equator Network's TRIPOD guidelines recommend the calculation of the NB, which reflects whether the benefits conferred by intervening on true positives outweigh the harms conferred by intervening on false positives. We refer to the NB achievable in the presence of resource constraints as the realized net benefit (RNB), and provide formulae for calculating the RNB. RESULTS: Using 4 case studies, we demonstrate the degree to which an absolute constraint (eg, only 3 available intensive care unit [ICU] beds) diminishes the RNB of a hypothetical ICU admission model. We show how the introduction of a relative constraint (eg, surgical beds that can be converted to ICU beds for very high-risk patients) allows us to recoup some of the RNB but with a higher penalty for false positives. DISCUSSION: RNB can be calculated in silico before the model's output is used to guide care. Accounting for the constraint changes the optimal strategy for ICU bed allocation. CONCLUSIONS: This study provides a method to account for resource constraints when planning model-based interventions, either to avoid implementations where constraints are expected to play a larger role or to design more creative solutions (eg, converted ICU beds) to overcome absolute constraints when possible.
Karandeep Singh, Nigam H. Shah, Andrew J. Vickers
J. Am. Medical Informatics Assoc.1
2023 Multi-Stage Machine Learning Model for Hierarchical Tie Valence Prediction
abstract
Individuals interacting in organizational settings involving varying levels of formal hierarchy naturally form a complex network of social ties having different tie valences (e.g., positive and negative connections). Social ties critically affect employees’ satisfaction, behaviors, cognition, and outcomes—yet identifying them solely through survey data is challenging because of the large size of some organizations or the often hidden nature of these ties and their valences. We present a novel deep learning model encompassing NLP and graph neural network techniques that identifies positive and negative ties in a hierarchical network. The proposed model uses human resource attributes as node information and web-logged work conversation data as link information. Our findings suggest that the presence of conversation data improves the tie valence classification by 8.91% compared to employing user attributes alone. This gain came from accurately distinguishing positive ties, particularly for male, non-minority, and older employee groups. We also show a substantial difference in conversation patterns for positive and negative ties with positive ties being associated with more messages exchanged on weekends, and lower use of words related to anger and sadness. These findings have broad implications for facilitating collaboration and managing conflict within organizational and other social networks.
Karandeep Singh, SeungEon Lee 0001, Giuseppe (Joe) Labianca, Jesse Michael Fagan, Meeyoung Cha
ACM Trans. Knowl. Discov. Data1
2023 Active Learning for Human-in-the-Loop Customs Inspection
abstract
We study the human-in-the-loop customs inspection scenario, where an AI-assisted algorithm supports customs officers by recommending a set of imported goods to be inspected. If the inspected items are fraudulent, the officers can levy extra duties. These logs are then used as additional training data for the next iterations. Choosing to inspect suspicious items first leads to an immediate gain in customs revenue, yet such inspections may not bring new insights for learning dynamic traffic patterns. On the other hand, inspecting uncertain items can help acquire new knowledge, which will be used as a supplementary training resource to update the selection systems. Based on multiyear customs datasets from three countries, we demonstrate that some degree of exploration is necessary to cope with domain shifts in the trade data. The results show that a hybrid strategy of selecting likely fraudulent and uncertain items will eventually outperform the exploitation-only strategy.
Sundong Kim, Tung-Duong Mai, Sungwon Han 0001, Sungwon Park 0001, Thi Nguyen Duc Khanh, Jaechan So, Karandeep Singh, Meeyoung Cha
IEEE Trans. Knowl. Data Eng.7
2022 Downscaling Earth System Models with Deep Learning
abstract
Modern climate models offer simulation results that provide unprecedented details at the local level. However, even with powerful supercomputing facilities, their computational complexity and associated costs pose a limit on simulation resolution that is needed for agile planning of resource allocation, parameter calibration, and model reproduction. As regional information is vital for policymakers, data from coarse-grained resolution simulations undergo the process of "statistical downscaling" to generate higher-resolution projection at a local level. We present a new method for downscaling climate simulations called GINE (Geospatial INformation Encoded statistical downscaling). To preserve the characteristics of climate simulation data during this process, our model applies the latest computer vision techniques over topography-driven spatial and local-level information. The comprehensive evaluations on 2x, 4x, and 8x resolution factors show that our model substantially improves performance in terms of RMSE and the visual quality of downscaled data.
Sungwon Park 0001, Karandeep Singh, Arjun Nellikkattil, Elke Zeller, Tung-Duong Mai, Meeyoung Cha
KDD2
2021 Making Health AI Work in the Real World: Strategies, innovations, and best practices for using AI to improve care delivery
Suchi Saria, Marzyeh Ghassemi, Ziad Obermeyer, Karandeep Singh, Pei-Yun S. Hsueh, Eric J. Topol
AMIA4
2020 DATE: Dual Attentive Tree-aware Embedding for Customs Fraud Detection
abstract
Intentional manipulation of invoices that lead to undervaluation of trade goods is the most common type of customs fraud to avoid ad valorem duties and taxes. To secure government revenue without interrupting legitimate trade flows, customs administrations around the world strive to develop ways to detect illicit trades. This paper proposes DATE, a model of Dual-task Attentive Tree-aware Embedding, to classify and rank illegal trade flows that contribute the most to the overall customs revenue when caught. The strength of DATE comes from combining a tree-based model for interpretability and transaction-level embeddings with dual attention mechanisms. To accurately identify illicit transactions and predict tax revenue, DATE learns simultaneously from illicitness and surtax of each transaction. With a five-year amount of customs import data with a test illicit ratio of 2.24%, DATE shows a remarkable precision of 92.7% on illegal cases and a recall of 49.3% on revenue after inspecting only 1% of all trade flows. We also discuss issues on deploying DATE in Nigeria Customs Service, in collaboration with the World Customs Organization.
Sundong Kim, Yu-Che Tsai, Karandeep Singh, Yeonsoo Choi, Etim Ibok, Cheng-Te Li, Meeyoung Cha
KDD3
2018 Towards a Learning Health System to Reduce Emergency Department Visits at a Population Level
Elliott Brannon, Jeremy Lapedis, Paul Valenstein, Michael S. Klinkman, Ellen Bunting, Alice Stanulis, Karandeep Singh
AMIA8
2018 A Micro-Level Data-Calibrated Agent-Based Model: The Synergy between Microsimulation and Agent-Based Modeling
abstract
Artificial life (ALife) examines systems related to natural life, its processes, and its evolution, using simulations with computer models, robotics, and biochemistry. In this article, we focus on the computer modeling, or "soft," aspects of ALife and prepare a framework for scientists and modelers to be able to support such experiments. The framework is designed and built to be a parallel as well as distributed agent-based modeling environment, and does not require end users to have expertise in parallel or distributed computing. Furthermore, we use this framework to implement a hybrid model using microsimulation and agent-based modeling techniques to generate an artificial society. We leverage this artificial society to simulate and analyze population dynamics using Korean population census data. The agents in this model derive their decisional behaviors from real data (microsimulation feature) and interact among themselves (agent-based modeling feature) to proceed in the simulation. The behaviors, interactions, and social scenarios of the agents are varied to perform an analysis of population dynamics. We also estimate the future cost of pension policies based on the future population structure of the artificial society. The proposed framework and model demonstrates how ALife techniques can be used by researchers in relation to social issues and policies.
Karandeep Singh, Chang-Won Ahn, Euihyun Paik, Jang Won Bae, Chun-Hee Lee
Artif. Life1
2016 Mobile Apps for Vulnerable Populations Study
Urmimala Sarkar, Gato Gourley, Courtney R. Lyles, Lina Tieu, Cassidy Clarity, Lisa P. Newmark, Karandeep Singh, David W. Bates
AMIA7
2014 An Electronic Health Record for Google Glass: A System Demonstration
Karandeep Singh
AMIA1
2014 Developing an Electronic Health Record for Google Glass: Challenges and Use Cases
Karandeep Singh, Adam B. Landman, Joseph V. Bonventre, Adam Wright
AMIA1
2011 Heterogeneous Cloud Computing
abstract
Current cloud computing infrastructure typically assumes a homogeneous collection of commodity hardware, with details about hardware variation intentionally hidden from users. In this paper, we present our approach for extending the traditional notions of cloud computing to provide a cloud-based access model to clusters that contain a heterogeneous architectures and accelerators. We describe our ongoing work extending the Open Stack cloud computing stack to support heterogeneous architectures and accelerators, and our experiences running Open Stack on our local heterogeneous cluster testbed.
Stephen P. Crago, Kyle Dunn, Patrick Eads, Lorin Hochstein, Dong-In Kang 0001, Mikyung Kang, Devendra Modium, Karandeep Singh, Jinwoo Suh, John Paul Walters
CLUSTER8