Ankur Teredesai

dblp:t/AnkurTeredesai · DBLP profile ↗
← Back
27ranked-venue papers in the field
6as first author
7since 2021 · last 2024
0000-0002-2112-5895ORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 9 (4 first)Database Systems & Data Management · 8Big Data, Cloud & Distributed Data Systems · 7 (1 first)Other / Interdisciplinary · 2 (1 first)Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2024 NudgeRank: Digital Algorithmic Nudging for Personalized Health
abstract
In this paper we describe NudgeRankTM, an innovative digital algorithmic nudging system designed to foster positive health behaviors on a population-wide scale. Utilizing a novel combination of Graph Neural Networks augmented with an extensible Knowledge Graph, this Recommender System is operational in production, delivering personalized and context-aware nudges to over 1.1 million care recipients daily. This enterprise deployment marks one of the largest AI-driven health behavior change initiatives, accommodating diverse health conditions and wearable devices. Rigorous evaluation reveals statistically significant improvements in health outcomes, including a 6.17% increase in daily steps and 7.61% more exercise minutes. Moreover, user engagement and program enrollment surged, with a 13.1% open rate compared to baseline systems' 4%. Demonstrating scalability and reliability, NudgeRankTM operates efficiently on commodity compute resources while maintaining automation and observability standards essential for production systems.
Jodi Chiam, Aloysius Lim, Ankur Teredesai
KDD3
2024 The 2nd International Workshop: From Innovation to Scale (I2S) - Successfully Build, Commercialize, and Scale AI Innovations
abstract
In recent years, there have been exciting and accelerated developments in AI with novel developments in foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, the pace of adoption of these innovations driven by both academic and industry research labs has sped up with both big tech companies and startups looking to deliver value-differentiated products and services. With Generative AI (GenAI) garnering significant attention, the second edition of the I2S workshop focuses on two aspects: First, bringing together AI thought leaders from academia, big tech, and startups to discuss the opportunities, use-case themes, challenges, and risks of GenAI in various business verticals; and Second, bringing together startup founders to share experiences and lessons learned in commercializing GenAI innovations into successful enterprises highlighting challenges through the entire commercial journey - from productization to acquiring customers, building a team, and securing funding.
Ankur Teredesai, Michael Zeller, Mohak Shah, Shenghua Bao, Wee Hyong Tok, Linsey Pang
KDD1
2023 NPRL: Nightly Profile Representation Learning for Early Sepsis Onset Prediction in ICU Trauma Patients
abstract
Sepsis is a syndrome that develops in the body in response to the presence of an infection. Characterized by severe organ dysfunction, sepsis is one of the leading causes of mortality in Intensive Care Units (ICUs) worldwide. These complications can be reduced through early application of antibiotics. Hence, the ability to anticipate the onset of sepsis early is crucial to the survival and well-being of patients. Current machine learning algorithms deployed inside medical infrastructures have demonstrated poor performance and are insufficient for anticipating sepsis onset early. Recently, deep learning methodologies have been proposed to predict sepsis, but some fail to capture the time of onset (e.g., classifying patients’ entire visits as developing sepsis or not) and others are unrealistic for deployment in clinical settings (e.g., creating training instances using a fixed time to onset, where the time of onset needs to be known apriori). In this paper, we first propose a novel but realistic prediction framework that predicts each morning whether sepsis onset will occur within the next 24 hours using the most recent data collected the previous night, when patient-provider ratios are higher due to cross-coverage resulting in limited observation to each patient. However, as we increase the prediction rate into daily, the number of negative instances will increase, while that of positive instances remain the same. This causes a severe class imbalance problem making it hard to capture these rare sepsis cases. To address this, we propose a nightly profile representation learning (NPRL) approach. We prove that NPRL can theoretically alleviate the rare event problem and our empirical study using data from a level-1 trauma center demonstrates the effectiveness of our proposal.
Tucker Stewart, Katherine Stern, Grant E. O'Keefe, Ankur Teredesai, Juhua Hu
IEEE Big Data4
2023 From Innovation to Scale (I2S) - Discuss and Learn How to Successfully Build, Commercialize, and Scale AI Innovations in Challenging Market Conditions
abstract
In recent years, the AI community has witnessed an exciting acceleration in innovation across foundation models, deep learning, new AI applications across numerous verticals, and more. In addition, AI innovations driven by both academic and industry research labs have rapidly been adopted by big tech companies and startups to deliver value-differentiated products and services.
Ankur Teredesai, Michael Zeller, Shenghua Bao, Wee Hyong Tok, Linsey Pang
KDD1
2023 Social Public Health Infrastructure for a Smart City Citizen Patient: Advances and Opportunities for AI Driven Disruptive Innovation
abstract
Promoting health, preventing disease, and prolonging life are central to the success of any smart city initiative. Today, wireless communication, data infrastructure, and low-cost sensors such as lifestyle and activity trackers are making it increasingly possible for cities to collect, collate, and innovate on developing a smart infrastructure. Combining this with AI driven disruptions for human behavior change can fundamentally transform delivery of public health for the citizen patient[4]. Most urban development government bodies consider such infrastructure to be a distributed ecosystem consisting of physical infrastructure, institutional infrastructure, social infrastructure and economic infrastructure[1]. In this talk we will focus mostly on the social health infrastructure component and first showcase some of the recent initiatives created with the purpose of addressing health which is a key social goal, as a smart city goal. Specifically we will discuss how technical advances in IoT, recommendation systems, geospatial computing, and digital health therapeutics are creating a new future. Yet, such advances are not addressing the issues of making healthy behavior change sustainable with broad health equity for those that need it most[3]. Overwhelming nature of siloed apps and digital health solutions which often leave the citizens overwhelmed and those most in need underserved[2]. This talk highlights the advances and opportunities created when behavioral economics and public health combine with AI and cloud infrastructure to make smart public health initiatives personalized to each individual citizen patient.
Ankur Teredesai
WSDM1
2022 Sub-Sequence Graph Representation Learning on High Variability Data for Dynamic Risk Prediction in Critical Care
abstract
Sepsis is an extreme inflammatory response of the body to an infection. It is one of the leading causes of death in critical care and ICUs worldwide, resulting in approximately 25% mortality in critically ill populations. Early identification and intervention are crucial to reducing sepsis-associated mortality and improving patient prognosis because severe sepsis cases can lead to organ failure and other life-threatening complications. Diagnosis of Sepsis is challenging in terms of diagnostic accuracy and timeliness due to ambiguous symptoms and individual differences which are captured in heterogeneous varied data sources with irregular time sequences. In recent years, numerous efforts have been made using machine learning methods for sepsis prediction. However, there are still very limited successful industry scale implementations due to limitations of consistency of input data, which often does not fit the characteristics of uncertain time intervals and large number of missing values in real-world critical care settings. In this study, we demonstrate an innovative approach to predict sepsis occurrence in real-time, on heterogeneous sets of variables with multiple time-granularity, using a flexible graph structure to model patient health records. Our modeling task is to predict future sepsis risk at any time after the first 48 hours of admission using observations data from any time window within the past 12 hours. To the best of our knowledge, our proposed approach is the first ever implementation that is dynamic, uses a graph representation to overcome the problem of irregular input features, and offers continuous risk prediction, thereby improving the compatibility of the model in clinical settings. Unlike prior efforts that report results on de-identified curated public datasets with synthetic data filters, our experiments and results are validated by clinical experts, and are based on a large level-1 trauma center’s multi-year real-world longitudinal data.
Ankur Teredesai, Sijin Huang, Tucker Stewart, Juhua Hu, Armaan Thakker, Katherine Stern, Grant E. O'Keefe
IEEE Big Data1
2021 Software as a Medical Device: Regulating AI in Healthcare via Responsible AI
abstract
With the increased adoption of AI in healthcare, there is a growing recognition and demand to regulate AI in healthcare to avoid potential harm and unfair bias against vulnerable populations. Around a hundred governmental bodies and commissions as well as leaders in the tech sector have proposed principles to create responsible AI systems. However, most of these proposals are short on specifics which has led to charges of ethics washing. In this tutorial we offer a guide to help navigate through complex governmental regulations and explain the various constituent practical elements of a responsible AI system in healthcare in the light of proposed regulations. Additionally, we breakdown and emphasize that the recommendations from regulatory bodies like FDA or the EU are necessary but not sufficient elements of creating a responsible AI system. We elucidate how regulations and guidelines often focus on epistemic concerns to the detriment of practical concerns e.g., requirement for fairness without explicating what fairness constitutes for a use case. FDA's Software as a medical device document and EU's GDPR among other AI governance documents talk about the need for implementing sufficiently good machine learning practices. In this tutorial we elucidate what that would mean from a practical perspective for real world use cases in healthcare throughout the machine learning cycle i.e., Data Management, Data Specification, Feature Engineering, Model Evaluation, Model Specification, Model Explainability, Model Fairness, Reproducibility, checks for data leakage and model leakage. We note that conceptualizing responsible AI as a process rather than an end goal accords well with how AI systems are used in practice. We also discuss how a domain centric stakeholder perspective translates into balancing requirements for multiple competing optimization criteria.
Muhammad Aurangzeb Ahmad, Steve Overman, Christine Allen, Ankur Teredesai, Carly Eckert
KDD5
2020 Fairness in Machine Learning for Healthcare
abstract
The issue of bias and fairness in healthcare has been around for centuries. With the integration of AI in healthcare the potential to discriminate and perpetuate unfair and biased practices in healthcare increases many folds The tutorial focuses on the challenges, requirements and opportunities in the area of fairness in healthcare AI and the various nuances associated with it. The problem healthcare as a multi-faceted systems level problem that necessitates careful of different notions of fairness in healthcare to corresponding concepts in machine learning is elucidated via different real world examples.
Muhammad Aurangzeb Ahmad, Arpit Patel, Carly Eckert, Ankur Teredesai
KDD5
2018 Privacy-Preserving Scoring of Tree Ensembles: A Novel Framework for AI in Healthcare
abstract
Machine Learning (ML) techniques now impact a wide variety of domains. Highly regulated industries such as healthcare and finance have stringent compliance and data governance policies around data sharing. Advances in secure multiparty computation (SMC) for privacy-preserving machine learning (PPML) can help transform these regulated industries by allowing ML computations over encrypted data with personally identifiable information (PII). Yet very little of SMC-based PPML has been put into practice so far. In this paper we present the very first framework for privacy-preserving classification of tree ensembles with application in healthcare. We first describe the underlying cryptographic protocols that enable a healthcare organization to send encrypted data securely to a ML scoring service and obtain encrypted class labels without the scoring service actually seeing that input in the clear. We then describe the deployment challenges we solved to integrate these protocols in a cloud based scalable risk-prediction platform with multiple ML models for healthcare AI. Included are system internals, and evaluations of our deployment for supporting physicians to drive better clinical outcomes in an accurate, scalable, and provably secure manner. To the best of our knowledge, this is the first such applied framework with SMC-based privacy-preserving machine learning for healthcare.
Kyle Fritchman, Keerthanaa Saminathan, Rafael Dowsley, Tyler Hughes, Martine De Cock, Anderson C. A. Nascimento, Ankur Teredesai
IEEE BigData7
2018 Distributed NoSQL Data Stores: Performance Analysis and a Case Study
abstract
NoSQL data-stores are commonly used to provide flexibility and availability for big data handling. However, there is a lack of comprehensive studies about which NoSQL data-store performs the best from the two scalability aspects, (scale-up, and scale-out), in a distributed and parallel processing environment. This paper compares the popular NoSQL data-stores (Cassandra, HBase, and MongoDB) and analyzes the resulting performance. Our experiments measure throughput, latency, and run-time of the evaluated data-stores on a big data set that consist of standard benchmarking workloads. Our results provide that the performance of each NoSQL data-store varies according to two main factors, (a) the type of executed operation, (read, scan, update, write, and insert), and (b) the level of distribution.
Abdeltawab M. Hendawi, Jayant Gupta, Jiayi Liu 0002, Ankur Teredesai, Naveen Ramakrishnan, Mohak Shah, Mohamed H. Ali
IEEE BigData4
2017 Smart Personalized Routing for Smart Cities
abstract
In smart cities, commuters have the opportunities for smart routing that may enable selecting a route with less car accidents, or one that is more scenic, or perhaps a straight and flat route. Such smart personalization requires a data management framework that goes beyond a static road network graph. This paper introduces PreGo, a novel system developed to provide real time personalized routing. The recommended routes by PreGo are smart and personalized in the sense of being (1) adjustable to individual users preferences, (2) subjective to the trip start time, and (3) sensitive to changes of the road conditions. Extensive experimental evaluation using real and synthetic data demonstrates the efficiency of the PreGo system.
Abdeltawab M. Hendawi, Aqeel Rustum, Amr A. Ahmadain, David Hazel, Ankur Teredesai, Dev Oliver, Mohamed H. Ali, John A. Stankovic
ICDE5
2016 Dynamic and Personalized Routing in PreGo
abstract
Existing routing services calculate the best route from source to destination over a road network graph. Most commercial routing services offer the best route in terms of either the shortest travel distance or the shortest travel time (with or without considering current traffic conditions). While travel distance and travel time are crucial route preferences for the commuter, other preferences are equally, or even more, important. Examples of other route preferences include fuel consumption, gas emissions, road safety, points of interest along the route, construction activities, open shops and restaurants. While some route preferences are static (e.g., travel distance and points of interests), other route preferences are dynamic and vary according to the time of the day (e.g., traffic-dependent travel time and the number of open shops/restaurants). Volunteered Geographic information (VGI) has been proposed as an approach to collect massive amounts of route information and, more specifically, the time varying parameters. This demo presents PreGo, a time-dependent multi-preference routing engine. During the demo, audience would interact with the PreGo routing engine to (1) find the optimal route w.r.t. The user's personal preferences for a given start time, (2) dynamically obtain the best start time for a trip given a set of preferences, (3) feed the system with VGI and examine their effect on the chosen route at real time, and (4) examine the correctness and efficiency of the PreGo selected routes compared to routes chosen by other commercial systems.
Abdeltawab M. Hendawi, Aqeel Rustum, Amr A. Ahmadain, Dev Oliver, David Hazel, Ankur Teredesai, Mohamed H. Ali
MDM6
2015 Scalable adaptive label propagation in Grappa
abstract
Nodes of a social graph often represent entities with specific labels, denoting properties such as age-group or gender. Design of algorithms to assign labels to unlabeled nodes by leveraging node-proximity and a-priori labels of seed nodes is of significant interest. A semi-supervised approach to solve this problem is termed "LPA-Label Propagation Algorithm" where labels of a subset of nodes are iteratively propagated through the network to infer yet unknown node labels. While LPA for node labelling is extremely fast and simple, it works well only with an assumption of node-homophily — connected nodes are connected because they must deserve a similar label — which can often be a misnomer. In this paper we propose a novel algorithm "Adaptive Label Propagation" that dynamically adapts to the underlying characteristics of homophily, heterophily, or otherwise, of the connections of the network, and applies suitable label propagation strategies accordingly. Moreover, our adaptive label propagation approach is scalable as demonstrated by its implementation in Grappa, a distributed shared-memory system. Our experiments on social graphs from Facebook, YouTube, Live Journal, Orkut and Netlog demonstrate that our approach not only improves the labelling accuracy but also computes results for millions of users within a few seconds.
Golnoosh Farnadi, Zeinab Mahdavifar, Ivan Keller, Jacob Nelson 0001, Ankur Teredesai, Marie-Francine Moens, Martine De Cock
IEEE BigData5
2015 Dynamic Hierarchical Classification for Patient Risk-of-Readmission
abstract
Congestive Heart Failure (CHF) is a serious chronic condition often leading to 50% mortality within 5 years. Improper treatment and post-discharge care of CHF patients leads to repeat frequent hospitalizations (i.e., readmissions). Accurately predicting patient's risk-of-readmission enables care-providers to plan resources, perform factor analysis, and improve patient quality of life. In this paper, we describe a supervised learning framework, Dynamic Hierarchical Classification (DHC) for patient's risk-of-readmission prediction. Learning the hierarchy of classifiers is often the most challenging component of such classification schemes. The novelty of our approach is to algorithmically generate various layers and combine them to predict overall 30-day risk-of-readmission. While the components of DHC are generic, in this work, we focus on congestive heart failure (CHF), a pressing chronic condition. Since healthcare data is diverse and rich and each source and feature-subset provides different insights into a complex problem, our DHC based prediction approach intelligently leverages each source and feature-subset to optimize different objectives (such as, Recall or AUC) for CHF risk-of-readmission. DHC's algorithmic layering capability is trained and tested over two real world datasets and is currently integrated into the clinical decision support tools at MultiCare Health System (MHS), a major provider of healthcare services in the northwestern US. It is integrated into a QlikView App (with EMR integration planned for Q2) and currently scores patients everyday, helping to mitigate readmissions and improve quality of care, leading to healthier outcomes and cost savings.
Senjuti Basu Roy, Ankur Teredesai, Kiyana Zolfaghar, Rui Liu 0014, David Hazel, Stacey Newman, Albert Marinez
KDD2
2015 COMA: Road Network Compression for Map-Matching
abstract
Road-network data compression reduces the size of the network to occupy lesser storage with the aim to fit small form-factor routing devices, mobile devices, or embedded systems. Compression (1) reduces the storage cost of memory and disks, and (2) reduces the I/O and communication overhead. There are several road network compression techniques proposed in literature. These techniques are evaluated by their compression ratios. However, none of these techniques takes into consideration the possibility that the generated compressed data can be used directly in map-matching. Map-matching is an essential component of routing services that matches a measured latitude and longitude of an object to an edge in the road network graph. In this paper, we propose a novel compression technique, named COMA, that significantly reduces the size of a given road network data. Another advantage of the proposed technique is that it enables the generated compressed road network graph to be used directly in map-matching without a need to decompress it beforehand. COMA smartly deletes those nodes and edges that will not affect neither the graph connectivity nor the accuracy of map-matching objects' location. COMA is equipped with an adjustable parameter, termed conflict factor C, by which location-based services can achieve a trade-off between the compression gain and map-matching accuracy. Extensive experimental evaluation on real road network data demonstrates competitive performance on compression-ratio and the high mapmatching accuracy achieved by the proposed technique.
Abdeltawab M. Hendawi, Amruta Khot, Aqeel Rustum, Anas Basalamah, Ankur Teredesai, Mohamed H. Ali
MDM (1)5
2015 A Map-Matching Aware Framework for Road Network Compression
abstract
We demonstrate a novel location aware services framework termed COMA for efficient compression and map-matching of road-network graph data. Key innovations include working demonstration of a new compression algorithm to eliminate nodes and edges that do not affect graph connectivity while ensuring object location map-matching accuracy. The demonstration features: (1) Algorithm to leverage compressed versions of road-network graphs for map-matching of objects locations to correct road edges without decompression, (2) Upwards of 75% compression ratio implying significant savings for road network data transmission costs, a key constraint for internet of things (IOT) devices, and (3) Use of a new controllable parameter, termed conflict factor C, whereby location aware services can trade the compression efficiency with map-matching accuracy at varying granularity. In addition to above features the demonstration features an extensible framework that enables experimentation and comparison between various compression and map-matching algorithms in a rich interactive interface. In this paper we outline data management challenges for location aware services using various scenarios for compression of a real road-network map of a large region of United States, along with both real and synthetic moving object trajectories distributed over this map. We describe the COMA framework through its map-based Graphical User Interface, ability to select the area of interest from the road-network map, and submit a compression request. COMA can export and save the compact versions of the selected area in different formats and plot the compressed graph over the original map for visual inspection to study the differences between compressed and uncompressed versions and various related statistics.
Abdeltawab M. Hendawi, Amruta Khot, Aqeel Rustum, Anas Basalamah, Ankur Teredesai, Mohamed H. Ali
MDM (1)5
2014 Computing fuzzy rough approximations in large scale information systems
abstract
Rough set theory is a popular and powerful machine learning tool. It is especially suitable for dealing with information systems that exhibit inconsistencies, i.e. objects that have the same values for the conditional attributes but a different value for the decision attribute. In line with the emerging granular computing paradigm, rough set theory groups objects together based on the indiscernibility of their attribute values. Fuzzy rough set theory extends rough set theory to data with continuous attributes, and detects degrees of inconsistency in the data. Key to this is turning the indiscernibility relation into a gradual relation, acknowledging that objects can be similar to a certain extent. In very large datasets with millions of objects, computing the gradual indiscernibility relation (or in other words, the soft granules) is very demanding, both in terms of runtime and in terms of memory. It is however required for the computation of the lower and upper approximations of concepts in the fuzzy rough set analysis pipeline. Current non-distributed implementations in R are limited by memory capacity. For example, we found that a state of the art non-distributed implementation in R could not handle 30,000 rows and 10 attributes on a node with 62GB of memory. This is clearly insufficient to scale fuzzy rough set analysis to massive datasets. In this paper we present a parallel and distributed solution based on Message Passing Interface (MPI) to compute fuzzy rough approximations in very large information systems. Our results show that our parallel approach scales with problem size to information systems with millions of objects. To the best of our knowledge, no other parallel and distributed solutions have been proposed so far in the literature for this problem.
Hasan Asfoor, Rajagopalan Srinivasan, Gayathri Vasudevan, Nele Verbiest, Chris Cornelis, Matthew E. Tolentino, Ankur Teredesai, Martine De Cock
IEEE BigData7
2014 Routing service with real world severe weather
abstract
Traditional routing services aim to save driving time by recommending the shortest path, in terms of distance or time, to travel from a start location to a given destination. However, these methods are relatively static and to a certain extent rely on traffic patterns under relatively normal conditions to calculate and recommend an appropriate route. As such, they do not necessarily translate effectively during severe weather events such as tornadoes. In these scenarios, the guiding principal is not, optimize for travel time, but rather, optimize for survivability of the event, i.e., can we recommend an evacuation route to those users inside the hazardous areas. In this demo, we present a framework for routing services for evacuating and avoiding real world severe weather threats that is able to: (1) Identify the users inside the dangerous region of a severe weather event (2) Recommend an evacuation route to guide the users out to a safe destination or shelter (3) Assure the recommended route to be one of the shortest paths after excluding the risky area (4) Maintain the flow of traffic by normalizing the evacuation on the possible safe routes. During the demo, attendees will be able to use the system interactively through its graphical user interface within a number of different scenarios. They will be able to locate the severe weather events on real time basis in any area in USA and examine detailed information about each event, to issue an evacuation query from an existing dangerous area by identifying a destination location and receiving the routing direction on their mobile devices, to issue an avoidance routing query to ask for a shortest path that avoids the dangerous region, to have an inside look into the internal system components and finally, to evaluate the overall system performance.
YiRu Li, Sarah George, Craig Apfelbeck, Abdeltawab M. Hendawi, David Hazel, Ankur Teredesai, Mohamed H. Ali
SIGSPATIAL/GIS6
2014 A Framework to Recommend Interventions for 30-Day Heart Failure Readmission Risk
abstract
In this paper, we describe a novel framework to recommend personalized intervention strategies to minimize 30-day readmission risk for heart failure (HF) patients, as they move through the provider's cardiac care protocol. We design principled solutions by learning the structure and parameters of a multi-layer hierarchical Bayesian network from underlying high-dimensional patient data. Next, we generate and summarize the rules leading to personalized interventions which can be applied to individual patients as they progress from admit to discharge. We present comprehensive experimental results as well as interesting case studies to demonstrate the effectiveness of our proposed framework using large real-world patient datasets on Microsoft Azure for Research platform.
Rui Liu 0014, Kiyana Zolfaghar, Si-Chi Chin, Senjuti Basu Roy, Ankur Teredesai
ICDM5
2013 Big data solutions for predicting risk-of-readmission for congestive heart failure patients
abstract
Developing holistic predictive modeling solutions for risk prediction is extremely challenging in healthcare informatics. Risk prediction involves integration of clinical factors with socio-demographic factors, health conditions, disease parameters, hospital care quality parameters, and a variety of variables specific to each health care provider making the task increasingly complex. Unsurprisingly, many of such factors need to be extracted independently from different sources, and integrated back to improve the quality of predictive modeling. Such sources are typically voluminous, diverse, and vary significantly over the time. Therefore, distributed and parallel computing tools collectively termed big data have to be developed. In this work, we study big data driven solutions to predict the 30-day risk of readmission for congestive heart failure (CHF) incidents. First, we extract useful factors from National Inpatient Dataset (NIS) and augment it with our patient dataset from Multicare Health System (MHS). Then, we develop scalable data mining models to predict risk of readmission using the integrated dataset. We demonstrate the effectiveness and efficiency of the open-source predictive modeling framework we used, describe the results from various modeling algorithms we tested, and compare the performance against baseline non-distributed, non-parallel, non-integrated small data results previously published to demonstrate comparable accuracy over millions of records.
Kiyana Zolfaghar, Naren Meadem, Ankur Teredesai, Senjuti Basu Roy, Si-Chi Chin, Brian Muckian
IEEE BigData3
2012 ACM SIGSPATIAL GIS Cup 2012
abstract
The 20th ACM SIGSPATIAL Conference on Advances in Geographic Information Systems (GIS) was held in November of 2012. In conjunction with this conference, we organized the conference's first competition, called the SIGSPATIAL GIS Cup 2012. The subject of the competition was map matching, which is the problem of correctly matching a sequence of noisy GPS points to roads. We describe the details of the contest, the results of the competition, and the lessons we learned in running a contest like this.
Mohamed H. Ali, John Krumm, Travis Rautman, Ankur Teredesai
SIGSPATIAL/GIS4
2011 An Extensibility Approach for Spatio-temporal Stream Processing Using Microsoft StreamInsight
Jeremiah Miller, Miles Raymond, Josh Archer, Seid Adem, Leo Hansel, Sushma Konda, Malik Luti, Ankur Teredesai, Mohamed H. Ali
SSTD9
2010 Ranking Approaches for Microblog Search
abstract
Ranking microblogs, such as tweets, as search results for a query is challenging, among other things because of the sheer amount of microblogs that are being generated in real time, as well as the short length of each individual microblog. In this paper, we describe several new strategies for ranking microblogs in a real-time search engine. Evaluating these ranking strategies is non-trivial due to the lack of a publicly available ground truth validation dataset. We have therefore developed a framework to obtain such validation data, as well as evaluation measures to assess the accuracy of the proposed ranking strategies. Our experiments demonstrate that it is beneficial for microblog search engines to take into account social network properties of the authors of microblogs in addition to properties of the microblog itself.
Rinkesh Nagmoti, Ankur Teredesai, Martine De Cock
Web Intelligence2
2009 Multi-dimensional phenomenon-aware stream query processing
abstract
Geographically co-located sensors tend to participate in the same environmental phenomena. Phenomenon-aware stream query processing improves scalability by subscribing each query only to a subset of sensors that participate in the phenomena of interest to that query. In the case of sensors that generate readings with a multi-attribute schema, phenomena may develop across the values of one or more attributes. However tracking and detecting phenomena across all attributes does not scale well as the dimensions increase. As the size of sensor network increases, and as the number of attributes being tracked by a sensor increases this becomes a major bottleneck. In this paper, we present a novel n-dimensional Phenomenon Detection and Tracking mechanism (termed as nd-PDT) over n-ary sensor readings. We reduce the number of dimensions to be tracked by first dropping dimensions without any meaningful phenomena, and then we further reduce the dimensionality by continuously detecting and updating various forms of functional dependencies amongst the phenomenon dimensions.
Ashish Bindra, Ankur Teredesai, Mohamed H. Ali, Walid G. Aref
GIS2
2009 A Comparative Analysis of Trust-Enhanced Recommenders for Controversial Items
Patricia Victor, Chris Cornelis, Martine De Cock, Ankur Teredesai
ICWSM4
2006 CoMMA: a framework for integrated multimedia mining using multi-relational associations
Ankur Teredesai, Muhammad Aurangzeb Ahmad, Juveria Kanodia, Roger S. Gaborski
Knowl. Inf. Syst.1
2001 Active Digit Classifiers: A Separability Optimization Approach to Emulate Cognition
abstract
Given sufficient resources, any classification task is possible with a high accuracy, but to achieve a particular task given finite resources, the problem is to utilize these resources intelligently. Cognitive studies in human vision associate multi-resolution features with high recognition accuracy. We show that classifier development using separability optimization is very similar to emulation of human cognition. The identification of key features leads to optimal resource utilization by the classifier. Evolving such classifiers is the focus of the paper. The resources required for classification can be identified in terms of amount of time required to develop a recognizer amount of processing power required and the number and kind of features extracted. Our digit recognition method strives not only to report high accuracy but also targets generation of simple solutions. The simplicity of a solution can be a measure of the resources utilized. Our methodology is termed as active based on the premise that once the complexity of a classification task is known an intelligent recognizer should incrementally increase the resources needed for classification.
Ankur Teredesai, Venu Govindaraju
ICDAR1