EDBT 2026 Demo / reviewers in the wild / expert
Rayid Ghani
dblp:19/6687
· DBLP profile ↗
45ranked-venue papers
15as first author
6since 2021 · last 2025
0000-0003-0235-1843ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 12 first-author · 5 since 2021Databases, data management, data science and information retrieval · 30 · 12 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Trustworthy machine learning · 77% Reinforcement learning · 6% Planning, search and constraint satisfaction · 6% | |
| Interdisciplinary, comprehensive, and emerging computing
16 papers |
Computational social science and digital humanities · 52% Medical and health informatics · 21% Computing education · 16% | |
| Databases, data mining, and information retrieval
14 papers |
Information retrieval · 46% Data mining · 28% Machine learning and data management · 17% | |
| Software engineering, system software, and programming languages
3 papers |
Empirical software engineering · 77% Software testing · 13% Requirements engineering and software design · 10% |
Topics — the 30 heaviest of 58, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
2.4 | 5 | 2024 | Aequitas Flow: Streamlining Fair ML Experimentation · J. Mach. Learn. Res. 2024 Addressing Bias and Fairness in Machine Learning: A Practical Guide and Hands-on Tutorial · KDD 2023 Dealing with Bias and Fairness in Data Science Systems: A Practical Hands-on Tutorial · KDD 2020 |
Machine learning › Trustworthy machine learning › fairness
bias mitigation |
1.1 | 2 | 2023 | Addressing Bias and Fairness in Machine Learning: A Practical Guide and Hands-on Tutorial · KDD 2023 Dealing with Bias and Fairness in Data Science Systems: A Practical Hands-on Tutorial · KDD 2020 |
Machine learning › Trustworthy machine learning › fairness
algorithmic fairness |
0.8 | 1 | 2024 | Aequitas Flow: Streamlining Fair ML Experimentation · J. Mach. Learn. Res. 2024 |
Machine learning › Trustworthy machine learning › interpretability
explainable AI |
0.8 | 1 | 2024 | On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML Methods · AAAI 2024 |
Machine learning › Trustworthy machine learning
interpretability |
0.8 | 1 | 2024 | On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML Methods · AAAI 2024 |
Empirical software engineering
experimental methodology |
0.8 | 1 | 2024 | On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML Methods · AAAI 2024 |
Machine learning › Trustworthy machine learning › fairness
fairness auditing |
0.7 | 1 | 2023 | Addressing Bias and Fairness in Machine Learning: A Practical Guide and Hands-on Tutorial · KDD 2023 |
Computational social science and digital humanities
public policy |
0.6 | 2 | 2018 | Deploying Machine Learning Models for Public Policy: A Framework · KDD 2018 Designing Policy Recommendations to Reduce Home Abandonment in Mexico · KDD 2016 |
Machine learning › Reinforcement learning › bandit
bandit learning |
0.6 | 1 | 2022 | Bandit Data-Driven Optimization for Crowdsourcing Food Rescue Platforms · AAAI 2022 |
Machine learning and data management
predictive maintenance |
0.3 | 1 | 2018 | Using Machine Learning to Assess the Risk of and Prevent Water Main Breaks · KDD 2018 |
Natural language and speech › Information extraction and text analysis › document understanding › table recognition
table parsing |
0.2 | 1 | 2016 | Identifying Earmarks in Congressional Bills · KDD 2016 |
Information retrieval › document processing › document analysis › text reuse detection
plagiarism detection |
0.2 | 1 | 2016 | The Legislative Influence Detector: Finding Text Reuse in State Legislation · KDD 2016 |
Information retrieval › document processing › document analysis
text reuse detection |
0.2 | 1 | 2016 | The Legislative Influence Detector: Finding Text Reuse in State Legislation · KDD 2016 |
Machine learning › Trustworthy machine learning › fairness
fair resource allocation |
0.2 | 1 | 2024 | Preventing Eviction-Caused Homelessness through ML-Informed Distribution of Rental Assistance · AAAI 2024 |
Computing education › student performance prediction
at-risk student identification |
0.2 | 1 | 2015 | A Machine Learning Framework to Identify Students at Risk of Adverse Academic Outcomes · KDD 2015 |
Medical and health informatics
clinical prediction |
0.2 | 1 | 2015 | Early Prediction of Cardiac Arrest (Code Blue) using Electronic Medical Records · KDD 2015 |
Medical and health informatics › clinical prediction
early warning |
0.2 | 1 | 2015 | Early Prediction of Cardiac Arrest (Code Blue) using Electronic Medical Records · KDD 2015 |
Medical and health informatics
public health |
0.2 | 1 | 2015 | Predictive Modeling for Public Health: Preventing Childhood Lead Poisoning · KDD 2015 |
Computing education
student performance prediction |
0.2 | 1 | 2015 | A Machine Learning Framework to Identify Students at Risk of Adverse Academic Outcomes · KDD 2015 |
Computing education › STEM education
data science education |
0.2 | 1 | 2023 | Addressing Bias and Fairness in Machine Learning: A Practical Guide and Hands-on Tutorial · KDD 2023 |
Machine learning › Efficient and distributed learning
active learning |
0.2 | 1 | 2013 | Online Active Learning with Imbalanced Classes · ICDM 2013 |
Machine learning › Efficient and distributed learning › data selection
instance selection |
0.2 | 1 | 2013 | Online Active Learning with Imbalanced Classes · ICDM 2013 |
Computational social science and digital humanities › causal inference
randomized controlled trial |
0.2 | 1 | 2013 | Targeting and influencing at scale: from presidential elections to social good · KDD 2013 |
Computational social science and digital humanities › legal informatics
legislative text analysis |
0.1 | 2 | 2016 | Identifying Earmarks in Congressional Bills · KDD 2016 The Legislative Influence Detector: Finding Text Reuse in State Legislation · KDD 2016 |
Machine learning › Trustworthy machine learning › fairness
bias evaluation |
0.1 | 1 | 2020 | Dealing with Bias and Fairness in Data Science Systems: A Practical Hands-on Tutorial · KDD 2020 |
Natural language and speech › Information extraction and text analysis
text classification |
0.1 | 4 | 2002 | Combining Labeled and Unlabeled Data for MultiClass Text Categorization · ICML 2002 Hypertext Categorization using Hyperlink Patterns and Meta Data · ICML 2001 Combining Labeled and Unlabeled Data for Text Classification with a Large Number of Categories · ICDM 2001 |
Data mining › anomaly detection
fraud detection |
0.1 | 1 | 2011 | Interactive learning for efficiently detecting errors in insurance claims · KDD 2011 |
Privacy and data protection
anonymization |
0.1 | 1 | 2011 | Testing software in age of data privacy: a balancing act · SIGSOFT FSE 2011 |
Data mining
anomaly detection |
0.1 | 1 | 2010 | Data mining to predict and prevent errors in health insurance claims processing · KDD 2010 |
Data stream processing › evolving data
concept drift |
0.1 | 1 | 2010 | Data mining to predict and prevent errors in health insurance claims processing · KDD 2010 |
Methods — techniques the papers use, named apart from their topics
machine learning · 4.8bias audit frameworks · 1.3aequitas toolkit · 1.3online bandit learning · 1.1offline predictive analytics · 1.1hyperparameter optimization · 0.8predictive modeling · 0.5early intervention system · 0.5fairness metrics · 0.4bias mitigation frameworks · 0.4gradient-boosted decision tree · 0.3gradient boosted decision trees · 0.3table-parsing algorithm · 0.2program analysis · 0.2machine learning classifier · 0.2data privacy framework · 0.2unsupervised performance score · 0.2randomized experiment · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | "You're in a Ferrari. I'm Waiting for the Bus": Confronting Tensions in Community-University PartnershipsabstractThere have been increasing calls within HCI to build sustained partnerships with communities that go beyond surface-level engagement. However, little is known about how communities view such partnerships and their outcomes. In collaboration with a community-based organization, we co-analyzed a series of interviews to understand the impacts of university-led research initiatives and publicly deployed technologies on local communities, and to explore strategies for more equitable community-university partnerships. Our findings reveal that local communities often perceive technology companies and academic institutions as potential threats due to their shared role in a series of projects, including predictive policing, surveillance, and broader concerns on technological bias and exclusion against minoritized groups. While interviewees named material benefits, sustained relationships, and meaningful accountability as desirable from universities, they pointed to academia's institutional priorities that pose barriers to forming effective partnerships. Drawing from la paperson's concept of a Third University, we argue that researchers and academic institutions must contend with these complexities, while taking a decolonizing approach to community-university partnerships through the lens of revestment. Cella Monet Sum, Jiayin Zhi, Amil N. T. Cook, Patrick James Cooper, Arturo Lozano, Tj Johnson, Jason Perez, Rayid Ghani, Michael Skirpan, Motahhare Eslami, Hong Shen 0004, Sarah E. Fox |
Proc. ACM Hum. Comput. Interact. | 8 |
| 2024 | On the Importance of Application-Grounded Experimental Design for Evaluating Explainable ML MethodsabstractMost existing evaluations of explainable machine learning (ML) methods rely on simplifying assumptions or proxies that do not reflect real-world use cases; the handful of more robust evaluations on real-world settings have shortcomings in their design, generally leading to overestimation of methods' real-world utility. In this work, we seek to address this by conducting a study that evaluates post-hoc explainable ML methods in a setting consistent with the application context and provide a template for future evaluation studies. We modify and improve a prior study on e-commerce fraud detection by relaxing the original work's simplifying assumptions that departed from the deployment context. Our study finds no evidence for the utility of the tested explainable ML methods in the context, which is a drastically different conclusion from the earlier work. This highlights how seemingly trivial experimental design choices can yield misleading conclusions about method utility. In addition, our work carries lessons about the necessity of not only evaluating explainable ML methods using tasks, data, users, and metrics grounded in the intended application context but also developing methods tailored to specific applications, moving beyond general-purpose explainable ML methods. Kasun Amarasinghe, Kit T. Rodolfa, Sérgio M. Jesus, Valerie Chen, Vladimir Balayan, Pedro Saleiro, Pedro Bizarro, Ameet Talwalkar, Rayid Ghani |
AAAI | 9 |
| 2024 | Preventing Eviction-Caused Homelessness through ML-Informed Distribution of Rental AssistanceabstractRental assistance programs provide individuals with financial assistance to prevent housing instabilities caused by evictions and avert homelessness. Since these programs operate under resource constraints, they must decide who to prioritize. Typically, funding is distributed by a reactive allocation process that does not systematically consider risk of future homelessness. We partnered with Anonymous County (PA) to explore a proactive and preventative allocation approach that prioritizes individuals facing eviction based on their risk of future homelessness. Our ML models, trained on state and county administrative data accurately identify at-risk individuals, outperforming simpler prioritization approaches by at least 20% while meeting our equity and fairness goals across race and gender. Furthermore, our approach would reach 28% of individuals who are overlooked by the current process and end up homeless. Beyond improvements to the rental assistance program in Anonymous County, this study can inform the development of evidence-based decision support tools in similar contexts, including lessons about data needs, model design, evaluation, and field validation. Catalina Vajiac, Arun Frey, Joachim Baumann 0002, Abigail Smith, Kasun Amarasinghe, Alice Lai, Kit T. Rodolfa, Rayid Ghani |
AAAI | 8 |
| 2024 | Aequitas Flow: Streamlining Fair ML ExperimentationabstractAequitas Flow is an open-source framework and toolkit for end-to-end Fair Machine Learning (ML) experimentation, and benchmarking in Python. This package fills integration gaps that exist in other fair ML packages. In addition to the existing audit capabilities in Aequitas, the Aequitas Flow module provides a pipeline for fairness-aware model training, hyperparameter optimization, and evaluation, enabling easy-to-use and rapid experiments and analysis of results. Aimed at ML practitioners and researchers, the framework offers implementations of methods, datasets, metrics, and standard interfaces for these components to improve extensibility. By facilitating the development of fair ML practices, Aequitas Flow hopes to enhance the incorporation of fairness concepts in AI systems making AI systems more robust and fair. Sérgio M. Jesus, Pedro Saleiro, Inês Oliveira e Silva, Beatriz M. Jorge, Rita P. Ribeiro, João Gama 0001, Pedro Bizarro, Rayid Ghani |
J. Mach. Learn. Res. | 8 |
| 2023 | Addressing Bias and Fairness in Machine Learning: A Practical Guide and Hands-on TutorialabstractAs data science and machine learning (ML) increasingly shape our society, the importance of developing fair algorithmic decision-making systems becomes paramount. There is a pressing need to train data scientists and practitioners on handling bias and fairness in real-world scenarios, from early stages of a data science project to maintaining ML systems in production. Existing resources are mostly academic and cover the ML training and optimization aspects of bias mitigation, leaving practitioners without comprehensive frameworks for making decisions throughout a real-world project lifecycle. This tutorial aims to bridge the gap between research and practice, providing an in-depth exploration of algorithmic fairness, encompassing metrics and definitions, practical case studies, data bias understanding, bias mitigation and model fairness audits using the Aequitas toolkit. Participants will be equipped to engage in conversations about bias, assist decision-makers in understanding options and trade-offs, evaluate project scoping aspects influencing fairness outcomes, and define actions and interventions based on model predictions. They will also learn to identify cohorts, target variables, evaluation metrics, and establish bias and fairness goals for different groups. Moreover, participants will gain insights into auditing and mitigating model bias, and implementing continuous monitoring to assess retraining needs. The tutorial addresses the current lack of practical training materials, methodologies, and tools for researchers and developers working on real-world algorithmic decision-making systems. By the conclusion of this hands-on tutorial, attendees will be well-versed in navigating bias-related issues, selecting appropriate metrics, and applying bias audit and mitigation frameworks and tools for informed design decisions in real-world data science systems. Rayid Ghani, Kit T. Rodolfa, Pedro Saleiro, Sérgio M. Jesus |
KDD | 1 |
| 2022 | Bandit Data-Driven Optimization for Crowdsourcing Food Rescue PlatformsabstractFood waste and insecurity are two societal challenges that coexist in many parts of the world. A prominent force to combat these issues, food rescue platforms match food donations to organizations that serve underprivileged communities, and then rely on external volunteers to transport the food. Previous work has developed machine learning models for food rescue volunteer engagement. However, having long worked with domain practitioners to deploy AI tools to help with food rescues, we understand that there are four main pain points that keep such a machine learning model from being actually useful in practice: small data, data collected only under the default intervention, unmodeled objectives due to communication gap, and unforeseen consequences of the intervention. In this paper, we introduce bandit data-driven optimization which not only helps address these pain points in food rescue, but also is applicable to other nonprofit domains that share similar challenges. Bandit data-driven optimization combines the advantages of online bandit learning and offline predictive analytics in an integrated framework. We propose PROOF, a novel algorithm for this framework and formally prove that it has no-regret. We show that PROOF performs better than existing baseline on food rescue volunteer recommendation. Zheyuan Shi, Steven Z. Wu, Rayid Ghani, Fei Fang 0001 |
AAAI | 3 |
| 2020 | Dealing with Bias and Fairness in Data Science Systems: A Practical Hands-on TutorialabstractTackling issues of bias and fairness when building and deploying data science systems has received increased attention from the research community in recent years, yet a lot of the research has focused on theoretical aspects and very limited set of application areas and data sets. There is a lack of 1) practical training materials, 2) methodologies, and 3) tools for researchers and developers working on real-world algorithmic decision making system to deal with issues of bias and fairness. Today, treating bias and fairness as primary metrics of interest, and building, selecting, and validating models using those metrics is not standard practice for data scientists. In this hands-on tutorial we will try to bridge the gap between research and practice, by deep diving into algorithmic fairness, from metrics and definitions to practical case studies, including bias audits using the Aequitas toolkit (http://github.com/dssg/aequitas). By the end of this hands-on tutorial, the audience will be familiar with bias mitigation frameworks and tools to help them making decisions during a project based on intervention and deployment contexts in which their system will be used. Pedro Saleiro, Kit T. Rodolfa, Rayid Ghani |
KDD | 3 |
| 2019 | Using machine learning to help vulnerable tenants in New York cityabstractTo keep housing affordable, the City of New York has implemented rent-stabilization policies to restrict the rate at which the rent of certain units can be increased every year. However, some landlords of these rent-stabilized units try to illegally force their tenants out in order to circumvent rent-stabilization laws and greatly increase the rent they can charge. To identify and help tenants who are vulnerable to such landlord harassment, the New York City Public Engagement Unit (NYC PEU) conducts targeted outreach to tenants to inform them of their rights and to assist them with serious housing challenges. In this paper, we1 collaborated with NYC PEU to develop machine learning models to better prioritize outreach and help to vulnerable tenants. Our best-performing model can potentially help TSU find 59% more buildings where tenants face landlord harassment than the current outreach method using the same resources. The results also highlight the factors that help predict the risk of experiencing tenant harassment, and provide a data-driven and comprehensive approach to improve the city's policy of proactive outreach to vulnerable tenants. Teng Ye, Rebecca Ann Johnson, Samantha Fu, Jerica Copeny, Bridgit Donnelly, Alex Freeman, Mirian Lima, Joe Walsh, Rayid Ghani |
COMPASS | 9 |
| 2018 | Reducing Incarceration through Prioritized InterventionsabstractThe most vulnerable individuals in society often struggle with long-lasting, multi-faceted challenges like mental illness, substance abuse, chronic health conditions, and homelessness. Individuals experiencing these difficulties tend to interact with public services and departments frequently, but many communities are struggling to identify those individuals, let alone meet their needs in meaningful and cost-effective ways. In this paper, we describe our work with Johnson County, Kansas, that uses machine learning to prioritize outreach to individuals most at risk of being booked into jail within the next year. For the first time, we brought together Johnson County's jail, emergency medical, and mental health data, identified individuals who touch multiple systems, and built a model to predict individual jail bookings. Our system significantly outperformed both a random baseline and several simple heuristics that domain experts are likely to use and implement. By focusing on 200 individuals (which is the intervention capacity of Johnson County) who had interacted with both mental health services and the criminal justice system, we predicted jail bookings in the following year with 51% precision, which outperforms a baseline heuristic model by 1.5 times, and is 4.6 times better than a random baseline. This work provides a framework and prototype system for Johnson County as well as many other jurisdictions that are part of the Data Driven Justice Initiative as they develop intervention models to proactively connect social and mental health workers with individuals in need of care to avoid incarceration. Matthew J. Bauman, Kate S. Boxer, Tzu-Yun Lin, Erika Salomon, Hareem Naveed, Lauren Haynes, Joe Walsh, Jennifer Helsby, Steve Yoder, Robert Sullivan, Chris Schneweis, Rayid Ghani |
COMPASS | 12 |
| 2018 | Improving Government Response to Citizen Requests OnlineabstractThe Mexican constitution guarantees its citizens the right to submit individual requests to the government. Public officials are obligated to read and respond to citizen requests in a timely manner. Each request goes through three processing steps during which human employees read, analyze, and route requests to the appropriate federal agency depending on its content. The Mexican government recently created a centralized online submission system. In the year following the release of the online system, the number of submitted requests doubled. With limited resources to manually process each request, the Sistema Atención Ciudadana (SAC) office in charge of handling requests has struggled to keep up with the increasing volume, resulting in longer processing time. Our goal is to build a machine learning system to process requests in order to allow the government to respond to citizen requests more efficiently. Garren Gaut, Andrea Navarrete Rivera, Laila Wahedi, Paul van der Boor, Adolfo De Unánue, Jorge Díaz, Eduardo Clark, Rayid Ghani |
COMPASS | 8 |
| 2018 | Deploying Machine Learning Models for Public Policy: A FrameworkabstractMachine learning research typically focuses on optimization and testing on a few criteria, but deployment in a public policy setting requires more. Technical and non-technical deployment issues get relatively little attention. However, for machine learning models to have real-world benefit and impact, effective deployment is crucial. In this case study, we describe our implementation of a machine learning early intervention system (EIS) for police officers in the Charlotte-Mecklenburg (North Carolina) and Metropolitan Nashville (Tennessee) Police Departments. The EIS identifies officers at high risk of having an adverse incident, such as an unjustified use of force or sustained complaint. We deployed the same code base at both departments, which have different underlying data sources and data structures. Deployment required us to solve several new problems, covering technical implementation, governance of the system, the cost to use the system, and trust in the system. In this paper we describe how we addressed and solved several of these challenges and provide guidance and a framework of important issues to consider for future deployments. Klaus Ackermann, Joe Walsh, Adolfo De Unánue, Hareem Naveed, Andrea Navarrete Rivera, Sun-Joo Lee, Jason Bennett, Michael Defoe, Crystal Cody, Lauren Haynes, Rayid Ghani |
KDD | 11 |
| 2018 | Using Machine Learning to Assess the Risk of and Prevent Water Main BreaksabstractWater infrastructure in the United States is beginning to show its age, particularly through water main breaks. Main breaks cause major disruptions in everyday life for residents and businesses. Water main failures in Syracuse, N.Y. (as in most cities) are handled reactively rather than proactively. A barrier to proactive maintenance with limited resources is the city's inability to properly prioritize the allocation of its resources. We built a Machine Learning system to assess the risk of a water mains breaking. Using historical data on which mains have failed, descriptors of pipes, and other data sources, we evaluated several models' abilities to predict breaks three years into the future. Our results show that our system using gradient boosted decision trees performed best out of several algorithms and expert heuristics, achieving precision at 1% ([email protected]) of 0.62. Our model outperforms a random baseline ([email protected] of 0.08) and expert heuristics such as water main age ([email protected] of 0.10) and history of past main breaks ([email protected] of 0.48). The model is currently deployed in the City of Syracuse. We are conducting a pilot by calculating the risk of failure for each city block over the period 2016-2018 using data up to the end of 2015 and, as of the end of 2017, there have been 42 breaks on our riskiest 52 mains. This has been a successful initiative for the city of Syracuse in improving its infrastructure and we believe this approach can be applied to other cities. Avishek Kumar, Syed Ali Asad Rizvi, Benjamin Brooks, Ali Vanderveld, Kevin H. Wilson, Chad Kenney, Sam Edelstein, Adria Finch, Andrew Maxwell, Joe Zuckerbraun, Rayid Ghani |
KDD | 11 |
| 2018 | Data Science for Social Good and Public Policy: Examples, Opportunities, and ChallengesabstractCan data science help reduce police violence and misconduct? Can it help increase retention of patients in care? Can it help prevent children from getting lead poisoning? Can it help cities better target limited resources to improve lives of citizens? We're all aware of the hype around data science and related buzzwords right now but turning this hype into social impact takes cross-disciplinary training, teams, and methods. In this talk, I'll discuss lessons learned from our work at University of Chicago while working on dozens of data science projects over the past few years with non-profits and governments on high-impact public policy and social challenges in criminal justice, public health, education, economic development, public safety, workforce training, and urban infrastructure. I'll highlight opportunities for IR researchers to get involved in these efforts as well as information retrieval challenges that are open research problems that need to be solved in order to increase the effectiveness of today's machine learning and data science algorithms in order to have social and policy impact in a fair and equitable manner. Rayid Ghani |
SIGIR | 1 |
| 2016 | Detecting fraud, corruption, and collusion in international development contracts: The design of a proof-of-concept automated systemabstractInternational development banks provide low-interest loans to developing countries in an effort to stimulate social and economic development. These loans support key infrastructure projects including the building of roads, schools, and hospitals. However, despite the best efforts of development banks, these loan funds are often lost to fraud, corruption, and collusion. In an effort to sanction and deter this wrongdoing and to ensure proper use of funds, development banks conduct extensive, costly investigations that can take over a year to complete. This paper describes a proof-of-concept of a fully automated fraud, corruption, and collusion classification system for identifying risk in international development contracts. We developed this system in conjunction with the World Bank Group - the largest international development bank - to improve the time and cost efficiency of their investigation process. Using historical monetary award data and past investigation outcomes, our classifier assigns a “risk score” to World Bank contracts. This risk score is designed to enable World Bank investigators to identify the contracts most likely to lead to a substantiated investigation. If implemented, our automated system is predicted to successfully identify fraud, corruption, and collusion in 70% of cases. Emily Grace, Ankit Rai, Elissa M. Redmiles, Rayid Ghani |
IEEE BigData | 4 |
| 2016 | Designing Policy Recommendations to Reduce Home Abandonment in MexicoabstractInfonavit, the largest provider of mortgages in Mexico, assists working families to obtain low-interest rate housing solutions. An increasingly prevalent problem is home abandonment: when a homeowner decides to leave their property and forego their investment. A major causal factor of this outcome is a mismatch between the homeowner's needs, in terms of access to services and employment, and the location characteristics of the home. This paper describes our collaboration with Infonavit to reduce home abandonment at two levels: develop policy recommendations for targeted improvements in location characteristics, and develop a decision-support tool to assist the homeowner in the home location decision. Using 20 years of mortgage history data combined with surveys, census, and location information, we develop a model to predict the probability of home abandonment based on both individual and location characteristics. The model is used to develop a tool that provides Infonavit the ability to give advice to Mexican workers when they apply for a loan, evaluate and improve the locations of new housing developments, and provide data-driven recommendations to the federal government to influence local development initiatives and infrastructure investments. The result is improving economic outcomes for the citizens of Mexico by pre-emptively identifying at-risk home mortgages, thereby allowing them to be altered or remedied before they result in abandonment. Klaus Ackermann, Eduardo Blancas Reyes, Sue He, Thomas Anderson Keller, Paul van der Boor, Romana Khan, Rayid Ghani, José Carlos González |
KDD | 7 |
| 2016 | The Legislative Influence Detector: Finding Text Reuse in State LegislationabstractState legislatures introduce at least 45,000 bills each year. However, we lack a clear understanding of who is actually writing those bills. As legislators often lack the time and staff to draft each bill, they frequently copy text written by other states or interest groups. However, existing approaches to detect text reuse are slow, biased, and incomplete. Journalists or researchers who want to know where a particular bill originated must perform a largely manual search. Watchdog organizations even hire armies of volunteers to monitor legislation for matches. Given the time-consuming nature of the analysis, journalists and researchers tend to limit their analysis to a subset of topics (e.g. abortion or gun control) or a few interest groups. Matthew Burgess, Eugenia Giraudy, Julian Katz-Samuels, Joe Walsh, Derek Willis, Lauren Haynes, Rayid Ghani |
KDD | 7 |
| 2016 | Identifying Police Officers at Risk of Adverse EventsabstractAdverse events between police and the public, such as deadly shootings or instances of racial profiling, can cause serious or deadly harm, damage police legitimacy, and result in costly litigation. Evidence suggests these events can be prevented by targeting interventions based on an Early Intervention System (EIS) that flags police officers who are at a high risk for involvement in such adverse events. Today's EIS are not data-driven and typically rely on simple thresholds based entirely on expert intuition. In this paper, we describe our work with the Charlotte-Mecklenburg Police Department (CMPD) to develop a machine learning model to predict which officers are at risk for an adverse event. Our approach significantly outperforms CMPD's existing EIS, increasing true positives by ~12% and decreasing false positives by ~32%. Our work also sheds light on features related to officer characteristics, situational factors, and neighborhood factors that are predictive of adverse events. This work provides a starting point for police departments to take a comprehensive, data-driven approach to improve policing and reduce harm to both officers and members of the public. Samuel Carton, Jennifer Helsby, Kenneth Joseph, Ayesha Mahmud, Joe Walsh, Crystal Cody, C. P. T. Estella Patterson, Lauren Haynes, Rayid Ghani |
KDD | 10 |
| 2016 | Identifying Earmarks in Congressional BillsabstractEarmarks are legislative provisions that direct federal funds to specific projects, circumventing the competitive grant-making process of federal agencies. Identifying and cataloging earmarks is a tedious, time-consuming process carried out by experts from public interest groups. In this paper, we present a machine learning system for automatically extracting earmarks from congressional bills and reports. We first describe a table-parsing algorithm for extracting budget allocations from appropriations tables in congressional bills. We then use machine learning classifiers to identify budget allocations as earmarked objects with an out of sample ROC AUC score of 0.89. Using this system, we construct the first publicly available database of earmarks dating back to 1995. Our machine learning approach adds transparency, accuracy, and speed to the congressional appropriations process. Ellery Wulczyn, Madian Khabsa, Vrushank Vora, Matthew Heston, Joe Walsh, Christopher Berry, Rayid Ghani |
KDD | 7 |
| 2015 | A Machine Learning Framework to Identify Students at Risk of Adverse Academic OutcomesabstractMany school districts have developed successful intervention programs to help students graduate high school on time. However, identifying and prioritizing students who need those interventions the most remains challenging. This paper describes a machine learning framework to identify such students, discusses features that are useful for this task, applies several classification algorithms, and evaluates them using metrics important to school administrators. To help test this framework and make it practically useful, we partnered with two U.S. school districts with a combined enrollment of approximately 200,000 students. We together designed several evaluation metrics to assess the goodness of machine learning algorithms from an educator's perspective. This paper focuses on students at risk of not finishing high school on time, but our framework lays a strong foundation for future work on other adverse academic outcomes. Himabindu Lakkaraju, Everaldo Aguiar, Carl Shan, Nasir Bhanpuri, Rayid Ghani, Kecia L. Addison |
KDD | 6 |
| 2015 | Predictive Modeling for Public Health: Preventing Childhood Lead PoisoningabstractLead poisoning is a major public health problem that affects hundreds of thousands of children in the United States every year. A common approach to identifying lead hazards is to test all children for elevated blood lead levels and then investigate and remediate the homes of children with elevated tests. This can prevent exposure to lead of future residents, but only after a child has been poisoned. This paper describes joint work with the Chicago Department of Public Health (CDPH) in which we build a model that predicts the risk of a child to being poisoned so that an intervention can take place before that happens. Using two decades of blood lead level tests, home lead inspections, property value assessments, and census data, our model allows inspectors to prioritize houses on an intractably long list of potential hazards and identify children who are at the highest risk. This work has been described by CDPH as pioneering in the use of machine learning and predictive analytics in public health and has the potential to have a significant impact on both health and economic outcomes for communities across the US. Eric Potash, Joe Brew, Alexander Loewi, Subhabrata Majumdar, Andrew Reece, Joe Walsh, Eric William Davis, Emile Jorgenson, Raed Mansour, Rayid Ghani |
KDD | 10 |
| 2015 | Early Prediction of Cardiac Arrest (Code Blue) using Electronic Medical RecordsabstractCode Blue is an emergency code used in hospitals to indicate when a patient goes into cardiac arrest and needs resuscitation. When Code Blue is called, an on-call medical team staffed by physicians and nurses is paged and rushes in to try to save the patient's life. It is an intense, chaotic, and resource-intensive process, and despite the considerable effort, survival rates are still less than 20% [4]. Research indicates that patients actually start showing clinical signs of deterioration some time before going into cardiac arrest [1][2[][3], making early prediction, and possibly intervention, feasible. In this paper, we describe our work, in partnership with NorthShore University HealthSystem, that preemptively flags patients who are likely to go into cardiac arrest, using signals extracted from demographic information, hospitalization history, vitals and laboratory measurements in patient-level electronic medical records. We find that early prediction of Code Blue is possible and when compared with state of the art existing method used by hospitals (MEWS - Modified Early Warning Score)[4], our methods perform significantly better. Based on these results, this system is now being considered for deployment in hospital settings. Sriram Somanchi, Samrachana Adhikari, Allen Lin, Elena Eneva, Rayid Ghani |
KDD | 5 |
| 2013 | Online Active Learning with Imbalanced ClassesabstractThis paper proposes an online algorithm for active learning that switches between different candidate instance selection strategies (ISS) for classification in imbalanced data sets. This is important for two reasons: 1) many real-world problems have imbalanced class distributions and 2) there is no ISS that always outperforms all the other techniques. We first empirically compare the performance of existing techniques on imbalanced data sets and show that different strategies work better on different data sets and some techniques even hurt compared to random selection. We then propose an unsupervised score to track and predict the performance of individual instance selection techniques, allowing us to select an effective technique without using a holdout set and wasting valuable labeled data. This score is used in a simple online learning approach that switches between different ISS at each iteration. The proposed approach performs better than the best individual strategy available to the online algorithm over data sets in this paper and provides a way to build practical and effective active learning system for imbalanced data sets. Zahra Ferdowsi, Rayid Ghani, Raffaella Settimi |
ICDM | 2 |
| 2013 | Targeting and influencing at scale: from presidential elections to social goodabstractIf you're still recovering from the barrage of ads, news, emails, Facebook posts, and newspaper articles that were giving you the latest poll numbers, asking you to volunteer, donate money, and vote, this talk will give you a look behind the scenes on why you were seeing what you were seeing. I will talk about how machine learning and data mining along with randomized experiments were used to target and influence tens of millions of people. Beyond the presidential elections, these methodologies for targeting and influence have the power to solve big problems in education, healthcare, energy, transportation, and related areas. I will talk about some recent work we're doing at the University of Chicago Data Science for Social Good summer fellowship program working with non-profits and government organizations to tackle some of these challenges. Rayid Ghani |
KDD | 1 |
| 2011 | A Machine Learning Based System for Semi-Automatically Redacting DocumentsabstractRedacting text documents has traditionally been a mostly manual activity, making it expensive and prone to disclosure risks. This paper describes a semi-automated system to en- sure a specified level of privacy in text data sets. Recent work has attempted to quantify the likelihood of privacy breaches for text data. We build on these notions to provide a means of obstructing such breaches by framing it as a multi-class classification problem. Our system gives users fine-grained control over the level of privacy needed to obstruct sensi- tive concepts present in that data. Additionally, our system is designed to respect a user-defined utility metric on the data (such as disclosure of a particular concept), which our methods try to maximize while anonymizing. We describe our redaction framework, algorithms, as well as a prototype tool built in to Microsoft Word that allows enterprise users to redact documents before sharing them internally and obscure client specific information. In addition we show experimen- tal evaluation using publicly available data sets that show the effectiveness of our approach against both automated attack- ers and human subjects.The results show that we are able to preserve the utility of a text corpus while reducing disclosure risk of the sensitive concept. Chad M. Cumby, Rayid Ghani |
IAAI | 2 |
| 2011 | Interactive learning for efficiently detecting errors in insurance claimsabstractMany practical data mining systems such as those for fraud detection and surveillance deal with building classifiers that are not autonomous but part of a larger interactive system with an expert in the loop. The goal of these systems is not just to maximize the performance of the classifier but to make the experts more efficient at performing their task, thus maximizing the overall Return on Investment of the system. This paper describes an interactive system for detecting payment errors in insurance claims with claim auditors in the loop. We describe an interactive claims prioritization component that uses an online cost-sensitive learning approach (more-like-this) to make the system efficient. Our interactive prioritization component is built on top of a batch classifier that has been trained to detect payment errors in health insurance claims and optimizes the interaction between the classifier and the domain experts who are consuming the results of this system. The goal is to make these auditors more efficient and effective as well as improving the classification performance of the system. The result is both a reduction in time it takes for the auditors to review and label claims as well as improving the precision of the system in finding payment errors. We show results obtained from applying this system at two major US health insurance companies indicating significant reduction in claim audit costs and potential savings of $20-$26 million/year making the insurance providers more efficient and lowering their operating costs. Our system reduces the money being wasted by providers and insurers dealing with incorrectly processed claims and makes the healthcare system more efficient. Rayid Ghani |
KDD | 1 |
| 2011 | Testing software in age of data privacy: a balancing actabstractDatabase-centric applications (DCAs) are common in enterprise computing, and they use nontrivial databases. Testing of DCAs is increasingly outsourced to test centers in order to achieve lower cost and higher quality. When proprietary DCAs are released, their databases should also be made available to test engineers. However, different data privacy laws prevent organizations from sharing this data with test centers because databases contain sensitive information. Currently, testing is performed with anonymized data, which often leads to worse test coverage (such as code coverage) and fewer uncovered faults, thereby reducing the quality of DCAs and obliterating benefits of test outsourcing. To address this issue, we offer a novel approach that combines program analysis with a new data privacy framework that we design to address constraints of software testing. With our approach, organizations can balance the level of privacy with needs of testing. We have built a tool for our approach and applied it to nontrivial Java DCAs. Our results show that test coverage can be preserved at a higher level by anonymizing data based on their effect on corresponding DCAs. Kunal Taneja, Mark Grechanik, Rayid Ghani, Tao Xie 0001 |
SIGSOFT FSE | 3 |
| 2010 | Data mining to predict and prevent errors in health insurance claims processingabstractHealth insurance costs across the world have increased alarmingly in recent years. A major cause of this increase are payment errors made by the insurance companies while processing claims. These errors often result in extra administrative effort to re-process (or rework) the claim which accounts for up to 30% of the administrative staff in a typical health insurer. We describe a system that helps reduce these errors using machine learning techniques by predicting claims that will need to be reworked, generating explanations to help the auditors correct these claims, and experiment with feature selection, concept drift, and active learning to collect feedback from the auditors to improve over time. We describe our framework, problem formulation, evaluation metrics, and experimental results on claims data from a large US health insurer. We show that our system results in an order of magnitude better precision (hit rate) over existing approaches which is accurate enough to potentially result in over $15-25 million in savings for a typical insurer. We also describe interesting research problems in this domain as well as design choices made to make the system easily deployable across health insurance companies. Rayid Ghani, Zhu-Song Mei |
KDD | 2 |
| 2009 | Toward Optimal Ordering of Prediction TasksabstractMany applications involve a set of prediction tasks that must be accomplished sequentially through user interaction.If the tasks are interdependent, the order in which they are performed may have a significant impact on the overall performance of the prediction systems.However, manual specification of an optimal order may be difficult when the interdependencies are complex, especially if the number of tasks is large, making exhaustive search intractable.This paper presents the first attempt at solving the optimal task ordering problem using an approximate formulation in terms of pairwise task order preferences, reducing the problem to the well-known Linear Ordering Problem.We propose two approaches for inducing the pairwise task order preferences -1) a classifier-agnostic approach based on conditional entropy that determines the prediction tasks whose correct labels lead to the least uncertainty for the remaining predictions, and 2) a classifier-dependent approach that empirically determines which tasks are favored before others for better predictive performance.We apply the proposed solutions to two practical applications that involve computer-assisted trouble report generation and document annotation, respectively.In both applications, the user fills up a series of fields and at each step, the system is expected to provide useful suggestions, which comprise the prediction (i.e.classification and ranking) tasks.Our experiments show encouraging improvements in predictive performance, as compared to approaches that do not take task dependencies into account. Abhimanyu Lad, Yiming Yang 0002, Rayid Ghani, Bryan Kisiel |
SDM | 3 |
| 2007 | Towards 'Interactive' Active Learning in Multi-view Feature Sets for Information Extraction
Katharina Probst, Rayid Ghani |
ECML | 2 |
| 2007 | Semi-Supervised Learning of Attribute-Value Pairs from Product Descriptions
Katharina Probst, Rayid Ghani, Marko Krema, Andrew E. Fano, Yan Liu 0002 |
IJCAI | 2 |
| 2005 | Building intelligent shopping assistants using individual consumer modelsabstractThis paper describes an Intelligent Shopping Assistant designed for a shopping cart mounted tablet PC that enables individual interactions with customers. We use machine learning algorithms to predict a shopping list for the customer's current trip and present this list on the device. As they navigate through the store, personalized promotions are presented using consumer models derived from loyalty card data for each inidvidual. In order for shopping assistant devices to be effective, we believe that they have to be powered by algorithms that are tuned for individual customers and can make accurate predictions about an individual's actions. We formally frame the shopping list prediction as a classification problem, describe the algorithms and methodology behind our system, and show that shopping list prediction can be done with high levels of accuracy, precision, and recall. Beyond the prediction of shopping lists we briefly introduce other aspects of the shopping assistant project, such as the use of consumer models to select appropriate promotional tactics, and the development of promotion planning simulation tools to enable retailers to plan personalized promotions delivered through such a shopping assistant. Chad M. Cumby, Andrew E. Fano, Rayid Ghani, Marko Krema |
IUI | 3 |
| 2005 | Price prediction and insurance for online auctionsabstractOnline auctions are generating a new class of fine-grained data about online transactions. This data lends itself to a variety of applications and services that can be provided to both buyers and sellers in online marketplaces. We collect data from online auctions and use several classification algorithms to predict the probable-end prices of online auction items. This paper describes the feature extraction and selection process, and several machine learning formulations of the price prediction problem. As a prototype application, we developed Auction Price Insurance that uses the predicted end-price to offer price insurance to sellers in online auctions. We define Price Insurance as a service that offers insurance to auction sellers that guarantees a price for their goods, for an appropriate premium. If the item sells for less than the insured price, the seller is reimbursed for the difference. We show that our price prediction techniques are accurate enough to offer price insurance as a profitable business. While this paper deals specifically with online auctions, we believe that this is an interesting case study that applies to dynamic markets where the price of the goods is variable and is affected by both internal and external factors that change over time. Rayid Ghani |
KDD | 1 |
| 2005 | Building Minority Language Corpora by Learning to Generate Web Search Queries
Rayid Ghani, Rosie Jones, Dunja Mladenic |
Knowl. Inf. Syst. | 1 |
| 2004 | Predicting customer shopping lists from point-of-sale purchase dataabstractThis paper describes a prototype that predicts the shopping lists for customers in a retail store. The shopping list prediction is one aspect of a larger system we have developed for retailers to provide individual and personalized interactions with customers as they navigate through the retail store. Instead of using traditional personalization approaches, such as clustering or segmentation, we learn separate classifiers for each customer from historical transactional data. This allows us to make very fine-grained and accurate predictions about what items a particular individual customer will buy on a given shopping trip.We formally frame the shopping list prediction as a classification problem, describe the algorithms and methodology behind our system, its impact on the business case in which we frame it, and explore some of the properties of the data source that make it an interesting testbed for KDD algorithms. Our results show that we can predict a shopper's shopping list with high levels of accuracy, precision, and recall. We believe that this work impacts both the data mining and the retail business community. The formulation of shopping list prediction as a machine learning problem results in algorithms that should be useful beyond retail shopping list prediction. For retailers, the result is not only a practical system that increases revenues by up to 11%, but also enhances customer experience and loyalty by giving them the tools to individually interact with customers and anticipate their needs. Chad M. Cumby, Andrew E. Fano, Rayid Ghani, Marko Krema |
KDD | 3 |
| 2002 | Using Text Mining to Infer Semantic Attributes for Retail Data MiningabstractCurrent data mining techniques usually do not have a mechanism to automatically infer semantic features inherent in the data being "mined". The semantics are either injected in the initial stages (by feature construction) or by interpreting the results produced by the algorithms. Both of these techniques have proved effective but require a lot of human effort. In many domains, semantic information is implicitly available and can be extracted automatically to improve data mining systems. In this paper we present a case study of a system that is trained to extract semantic features for apparel products and populate a knowledge base with these products and features. We show that semantic features of these items can be successfully extracted by applying text learning techniques to the descriptions obtained from websites of retailers. We also describe several applications of such a knowledge base of product semantics that we have built including recommender systems and competitive intelligence tools and provide evidence that our approach can successfully build a knowledge base with accurate facts which can then be used to create profiles of individual customers, groups of customers, or entire retail stores. Rayid Ghani, Andrew E. Fano |
ICDM | 1 |
| 2002 | Combining Labeled and Unlabeled Data for MultiClass Text Categorization
Rayid Ghani |
ICML | 1 |
| 2002 | A Study of Approaches to Hypertext Categorization
Yiming Yang 0002, Seán Slattery, Rayid Ghani |
J. Intell. Inf. Syst. | 3 |
| 2001 | Mining the Web to Create Minority Language CorporaabstractThe Web is a valuable source of language specific resources but the process of collecting, organizing and utilizing these resources is difficult. We describe CorpusBuilder, an approach for automatically generating Web-search queries for collecting documents in a minority language. It differs from pseudo-relevance feedback in that retrieved documents are labeled by an automatic language classifier as relevant or irrelevant, and this feedback is used to generate new queries. We experiment with various query-generation methods and query-lengths to find inclusion/exclusion terms that are helpful for retrieving documents in the target language and find that using odds-ratio scores calculated over the documents acquired so far was one of the most consistently accurate query-generation methods. We also describe experiments using a handful of words elicited from a user instead of initial documents and show that the methods perform similarly. Experiments applying the same approach to multiple languages are also presented showing that our approach generalizes to a variety of languages. Rayid Ghani, Rosie Jones, Dunja Mladenic |
CIKM | 1 |
| 2001 | Combining Labeled and Unlabeled Data for Text Classification with a Large Number of CategoriesabstractWe develop a framework to incorporate unlabeled data in the error-correcting output coding (ECOC) setup by decomposing multiclass problems into multiple binary problems and then use co-training to learn the individual binary classification problems. We show that our method is especially useful for classification tasks involving a large number of categories where co-training doesn't perform very well by itself and when combined with ECOC, outperforms several other algorithms that combine labeled and unlabeled data for text classification in terms of accuracy, precision-recall tradeoff, and efficiency. Rayid Ghani |
ICDM | 1 |
| 2001 | Hypertext Categorization using Hyperlink Patterns and Meta Data
Rayid Ghani, Seán Slattery, Yiming Yang 0002 |
ICML | 1 |
| 2001 | Automatic Web Search Query Generation to Create Minority Language CorporaabstractThe Web is a valuable source of language specific resources but collecting, organizing and utilizing this information is difficult. We describe CorpusBuilder, an approach for automatically generating Web-search queries to collect documents in a minority language. It differs from pseudo-relevance feedback in that retrieved documents are labeled by an automatic language classifier as relevant or irrelevant and a subset of documents is used to generate new queries. We experiment with various query-generation methods and query-lengths to find inclusion/exclusion terms that are helpful for finding documents in the target language and find that using odds-ratio scores calculated over the documents acquired so far was one of the most consistently accurate query-generation methods. We also describe experiments using a handful of words elicited from a user instead of initial documents and show that the methods perform similarly. Applying the same approach to multiple languages show that our system generalizes to a variety of languages. Rayid Ghani, Rosie Jones, Dunja Mladenic |
SIGIR | 1 |
| 2001 | Online Learning for Web Query Generation: Finding Documents Matching a Minority Concept on the Web
Rayid Ghani, Rosie Jones, Dunja Mladenic |
Web Intelligence | 1 |
| 2000 | Learning a Monolingual Language Model from a Multilingual Text Databaseabstract9999 9999 9999 9999 9999 9999 9999 9999 9999 + ,\t,+- , "#$\t!\t! %&\t' !($(*) 29039-47500 4* $%!3 29039-46450 &\t' !($(*) 29039-47500 '"6= 9 !$:\t! ,"<; 29659-45410 9039-47500 G"+\t\t!G"H '$/ !$"= 9039-47500 (/ ,"./ J+; 28939-43300 / !$"= 9039-47500 1\t'H+""Q%G1:\t'\t,H$%(DEH1:\t\t,R:\t' !$F\t@(-(/("/ $%\t !(+\t,+=(J""%\t!\t, ,R:\t' !$F\t@(-(/("/ 2 +\tV\t!F("$"#W\t!X\t/$(YV$%(+ (/("/ 2 $+\t,"$%/\t- '\tJ8\t!("""E(/(/ ("/ 28800 <\\?-+$"""/(:+(1/ ""E(/(/ ("/ 28800 1\t'] $%/\t>\t,6; 28969-37040 ""E(/(/ ("... Rayid Ghani, Rosie Jones |
CIKM | 1 |
| 2000 | Analyzing the Effectiveness and Applicability of Co-trainingabstractRecently there has been signi cant i n terest in supervised learning algorithms that combine labeled and unlabeled data for text learning tasks.The co-training setting [1] applies to datasets that have a natural separation of their features into two disjoint sets.We demonstrate that when learning from labeled and unlabeled data, algorithms explicitly leveraging a natural independent split of the features outperform algorithms that do not.When a natural split does not exist, co-training algorithms that manufacture a feature split may out-perform algorithms not using a split.These results help explain why co-training algorithms are both discriminative in nature and robust to the assumptions of their embedded classi ers. Kamal Nigam, Rayid Ghani |
CIKM | 2 |
| 2000 | Using Error-Correcting Codes for Text Classification
Rayid Ghani |
ICML | 1 |