EDBT 2026 Demo / reviewers in the wild / expert
Nitesh V. Chawla
dblp:c/NiteshVChawla · also Nitesh Vinay Chawla
· DBLP profile ↗
134ranked-venue papers in the field
5as first author
45since 2021 · last 2026
0000-0003-3932-5956ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 95 (3 first)Information Retrieval & Web Search · 25 (1 first)Database Systems & Data Management · 8Big Data, Cloud & Distributed Data Systems · 5 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CORE: Data Augmentation for Link Prediction via Information BottleneckabstractLink Prediction (LP) is a fundamental task in graph representation learning, with numerous applications in diverse domains. However, the generalizability of LP models is often compromised due to the presence of noisy or spurious information in graphs and the inherent incompleteness of graph data. To address these challenges, we draw inspiration from the Information Bottleneck principle and propose a novel data augmentation method, COmplete and REduce (CORE) to learn compact and predictive augmentations for LP models. In particular, CORE aims to recover missing edges in graphs while simultaneously removing noise from the graph structures, thereby enhancing the model’s robustness and performance. Extensive experiments on multiple benchmark datasets demonstrate the applicability and superiority of CORE over state-of-the-art methods, showcasing its potential as a leading approach for robust LP in graph representation learning. Kaiwen Dong, Zhichun Guo, Nitesh V. Chawla |
ACM Trans. Knowl. Discov. Data | 3 |
| 2026 | Generation of Loss Functions from Matrix-Based Binary Classification MetricsabstractMost evaluation metrics for binary classification are derived from the confusion matrix, which is inherently non-differentiable because it relies on discrete predictions. This limits their direct use as loss functions in gradient-based learning, creating a mismatch between training objectives and evaluation criteria. To that end, we offer a general-purpose approach, AnyLoss , that transforms any confusion-matrix-based metric into a differentiable loss function. AnyLoss employs a distinct approximation strategy to estimate a specific, targeted metric score for the prediction model. This is followed by a theoretical and practical analysis of the method, which involves conducting extensive experiments with neural network architectures ranging from simple to advanced across diverse data modalities, including tabular, image, and text. The experimental results demonstrate the generality of our new method, which can target any evaluation metrics derived from a confusion matrix, and highlight that it excels at handling imbalanced datasets. Do Heon Han, Nuno Moniz, Nitesh V. Chawla |
ACM Trans. Knowl. Discov. Data | 3 |
| 2026 | Safety in Graph Machine Learning: Threats and Safeguards
Song Wang 0013, Yushun Dong, Binchi Zhang, Zihan Chen 0002, Xingbo Fu, Yinhan He, Cong Shen 0001, Chuxu Zhang, Nitesh V. Chawla, Jundong Li |
IEEE Trans. Knowl. Data Eng. | 9 |
| 2025 | When Data & AI Converge for Good, Societal Impact AcceleratesabstractAs data and AI increasingly converge to drive societal impact, their true potential emerges at the intersection of innovation and translational research. In this keynote, I will present our group's work spanning the journey from data to algorithms to real-world translation – advancing methods for learning on graphs, addressing imbalanced data, and developing large language models, all with an eye toward meaningful impact across domains such as healthcare, sciences, and peace processes. I will also examine the “last-mile” challenge: bridging the gap between algorithmic advances and practical deployment, including issues of data quality, bias mitigation, governance, and equitable access, and discuss how closing this gap can accelerate inclusive and responsible progress. Nitesh V. Chawla |
IEEE Big Data | 1 |
| 2025 | ReactionTeam: Teaming Experts for Divergent Thinking Beyond Typical Reaction Patterns
Taicheng Guo, Changsheng Ma, Xiuying Chen, Bozhao Nan, Kehan Guo, Shichao Pei, Olaf Wiest, Nitesh V. Chawla, Xiangliang Zhang 0001 |
IEEE Big Data | 8 |
| 2025 | Socially Responsible and Trustworthy Generative Foundation Models: Principles, Challenges, and PracticesabstractGenerative foundation models (GenFMs), including large language and multimodal models, are transforming information retrieval and knowledge management. However, their rapid adoption raises urgent concerns about social responsibility, trustworthiness, and governance. This tutorial offers a comprehensive, hands-on overview of recent advances in responsible GenFMs, covering foundational concepts, multi-dimensional risk taxonomies (including safety, privacy, robustness, truthfulness, fairness, and machine ethics), state-of-the-art evaluation benchmarks, and effective mitigation strategies. We integrate real-world case studies and practical exercises using open-source tools, and present key perspectives from both policy and industry, including recent regulatory developments and enterprise practices. The session concludes with a discussion of open challenges, providing actionable guidance for the CIKM community. Yue Huang 0001, Canyu Chen, Lu Cheng 0001, Bhavya Kailkhura, Nitesh V. Chawla, Xiangliang Zhang 0001 |
CIKM | 5 |
| 2025 | Think it Image by Image: Multi-Image Moral Reasoning of Large Vision-Language ModelsabstractVision Language Models (VLMs) have demonstrated remarkable success in downstream applications, yet they often exhibit biases, raising ethical concerns. While previous efforts have aimed to evaluate and improve the moral reasoning capabilities of VLMs, existing approaches are limited by simplified, unimodal settings or overly static visual scenarios. We propose a novel multi-image-based dataset pipeline MIST (Moral Inference through Storytelling with Text and Images) designed to assess moral reasoning in complex, dynamic scenarios to address these limitations. To ensure better alignment between these modalities, we introduce the concept of ''text-image flow,'' which seamlessly integrates visual and textual information across complex scenarios. Using this dataset, we evaluate seven widely used VLMs, offering critical insights into their performance in moral reasoning tasks. Chujie Gao, Yue Huang 0001, Xiangqi Wang, Siyuan Wu 0001, Nitesh V. Chawla, Xiangliang Zhang 0001 |
CIKM | 5 |
| 2025 | Proto-Yield: An Uncertainty-Aware Prototype Network for Yield Prediction in Real-world Chemical ReactionsabstractReaction yield prediction underpins computer-aided synthesis prediction (CASP). Formulated as a regression problem that takes both reactants and products as input, this task has been extensively studied using machine learning methods, based on handcrafted fingerprint features, SMILES encoded by Transformers, and molecular graphs encoded by Graph Neural Networks. However, a major limitation of these methods is their inability to effectively capture and model the underlying uncertainties, arising both from the inherently stochastic nature of chemical reaction processes and from inconsistencies or noise in how yields are measured and reported. What makes this seemingly simple regression problem even more challenging is the lack of any principled way to account for the underlying uncertainties, due to missing or unrecorded experimental process (commonly happens in chemical labs). Kehan Guo, Zhen Liu 0069, Zhichun Guo, Bozhao Nan, Olexandr Isayev, Nitesh V. Chawla, Olaf Wiest, Xiangliang Zhang 0001 |
CIKM | 6 |
| 2025 | 8th Workshop on Machine Learning in FinanceabstractThe financial industry leverages machine learning in more ways than just finding the right alpha signal. It grapples with supply chains, business processes, marketing, churn, fraud, and money laundering, all while maintaining compliance with the various regulatory frameworks it is beholden to. Due to the sheer volume of wealth being handled by the financial industry and its critical role in everyday life, it has been a lucrative target for a wide spectrum of ever-evolving bad actors. With each successive iteration of this workshop, we have attempted to capture the breadth of these actors - fraudsters, money launderers, market manipulators, and potentially nation-state-level risks. The emerging advances in Generative AI make this a particularly exciting time to host this workshop. GenAI offers groundbreaking approaches to handling the various data types prevalent in the financial sector. From a security point of view, bad actors are actively using Generative AI creatively to thwart conventional defenses (e.g. voice cloning, better synthetic identities), and this workshop's audience would benefit from commonly applicable defenses & best practices against such threats. Last but not the least, there is now an increasing willingness from the financial industry towards deeper engagement and data sharing with academia. Saurabh Nagrecha, Isha Chaturvedi, Senthil Kumar, Nitesh V. Chawla, Mahashweta Das, Daksha Yadav, José A. Rodríguez-Serrano, Eren Kurshan |
KDD (2) | 4 |
| 2025 | Graph Foundation Models: Challenges, Methods, and Open QuestionsabstractFoundation models have revolutionized machine learning by enabling general-purpose reasoning across diverse tasks and domains. These models, pretrained on large-scale data, demonstrate strong adaptability with minimal task-specific supervision, leading to breakthroughs in natural language processing and computer vision. Inspired by this paradigm, Graph Foundation Models (GFMs) have emerged to extend the benefits of foundation models to graph-structured data, which is pretrained on massive graphs and can be fast adapted to different downstream tasks. In this paper, we provide a comprehensive survey of the state-of-the-art techniques of graph foundation models. In particular, we (1) formally categorize the challenges in designing graph foundation models; (2) comprehensively review the existing and recent advances of graph foundation models; (3) extend the graph foundation models in real-world problems; and (4) elucidate open questions and future research directions. Our systematic review summarizes representative models, highlights key design principles, and provides comparative analyses. This survey introduces major topics within foundation models and offers a guide to a new frontier of graph learning. Our extended survey is available at https://arxiv.org/abs/2505.15116. Zehong Wang, Chuxu Zhang, Jundong Li, Nitesh V. Chawla, Yanfang Ye 0001 |
KDD (2) | 4 |
| 2025 | MOPI-HFRS: A Multi-objective Personalized Health-aware Food Recommendation System with LLM-enhanced InterpretationabstractThe prevalence of unhealthy eating habits has become a growing concern in the United States. However, popular food recommendation platforms, such as Yelp, tend to prioritize users' dietary preferences over the healthiness of their choices. While some efforts have focused on developing health-aware food recommendation systems, personalization based on specific health conditions remains underexplored. Additionally, the lack of interpretability in these systems prevents users from evaluating the reliability of recommendations, limiting their practical adoption. To address these issues, we introduce two large-scale personalized health-aware food recommendation benchmarks at the first attempt. Building on this, we propose a novel framework called the Multi-Objective Personalized Interpretable Health-aware Food Recommendation System (MOPI-HFRS). This system generates food recommendations by jointly optimizing three objectives: user preference, personalized healthiness, and nutritional diversity. It also incorporates a reasoning module enhanced by large language models (LLMs) to provide interpretable recommendations that promote healthy dietary knowledge. The framework integrates descriptive features and health data using two structure learning and pooling modules within a graph learning framework. Pareto optimization is applied to balance the multi-faceted objectives. To further enhance healthy dietary knowledge, the system leverages LLMs by infusing knowledge from the recommendation model, generating meaningful interpretations for the recommendations. Extensive experiments on the proposed benchmarks demonstrate that MOPI-HFRS outperforms state-of-the-art methods by delivering diverse, healthy food recommendations alongside reliable explanations. Zheyuan Zhang 0008, Zehong Wang, Varun Sameer Taneja, Sofia Nelson, Nhi Ha Lan Le, Keerthiram Murugesan, Mingxuan Ju, Nitesh V. Chawla, Chuxu Zhang, Yanfang Ye 0001 |
KDD (1) | 9 |
| 2025 | Transaction Categorization with Relational Deep Learning in QuickBooks
Kaiwen Dong, Padmaja Jonnalagedda, Xiang Gao 0011, Ayan Acharya, Maria Kissa, Mauricio Flores, Nitesh V. Chawla, Kamalika Das |
ECML/PKDD (9) | 7 |
| 2025 | Beyond Answers: Transferring Reasoning Capabilities to Smaller LLMs Using Multi-Teacher Knowledge DistillationabstractTransferring the reasoning capability from stronger large language models (LLMs) to smaller ones has been quite appealing, as smaller LLMs are more flexible to deploy with less expense. Among the existing solutions, knowledge distillation stands out due to its outstanding efficiency and generalization. However, existing methods suffer from several drawbacks, including limited knowledge diversity and the lack of rich contextual information. To solve the problems and facilitate the learning of compact language models, we propose TinyLLM, a new knowledge distillation paradigm to learn a small student LLM from multiple large teacher LLMs. In particular, we encourage the student LLM to not only generate the correct answers but also understand the rationales behind these answers. Given that different LLMs possess diverse reasoning skills, we guide the student model to assimilate knowledge from various teacher LLMs. We further introduce an in-context example generator and a teacher-forcing Chain-of-Thought strategy to ensure that the rationales are accurate and grounded in contextually appropriate scenarios. Extensive experiments on six datasets across two reasoning tasks demonstrate the superiority of our method. Results show that TinyLLM can outperform large teacher LLMs significantly, despite a considerably smaller model size. The source code is available at: https://github.com/YikunHan42/TinyLLM. Yijun Tian 0001, Yikun Han, Xiusi Chen, Wei Wang 0010, Nitesh V. Chawla |
WSDM | 5 |
| 2025 | Ventana a la Verdad (Window to the Truth): A Chatbot Application for Navigating The Colombian Truth Commission's ArchivesabstractWe present Ventana a la Verdad, a chatbot designed to make the Clar- ification Archive and the reports of the Colombian Truth Commis- sion [6] more accessible to a wider audience. These archives contain a wealth of documents, interviews, and testimonies from Colom- bia's armed conflict, but navigating them can be challenging due to their volume and complexity. Using existing large language models (LLMs) and natural language processing techniques, our chatbot allows users to interact with the archives through natural language queries, receiving relevant and contextually appropriate responses. In the sensitive context of peace and reconciliation, where misin- formation or hallucinations can have significant adverse effects, ensuring the accuracy and reliability of information is paramount. This tool aims to facilitate better understanding and engagement with historical content, supporting educational and research efforts. We discuss the development of the chatbot, the challenges encoun- tered, and its potential impact on making the Colombian Truth Commission's archives more accessible. The chatbot is available by link here: http://ventanaverdad.lucyapps.net:1337/ Anna Sokol, Matthew L. Sisk, Josefina Echavarría Alvarez, Nitesh V. Chawla |
WSDM | 4 |
| 2025 | WildlifeLookup: A Chatbot Facilitating Wildlife Management with Accessible Data and InsightsabstractWildlife management is increasingly reliant on data-driven insights to address the impacts of climate change on species and ecosystems. However, the complexity of accessing and querying large, multimodal datasets often limits the ability of non-technical users, such as wildlife managers and conservationists, to make informed decisions. To address this challenge, we present WildlifeLookup, a public accessible, intelligent chatbot designed to facilitate natural language interaction with a novel knowledge graph (KN-Wildlife) that houses critical wildlife and environmental data. WildlifeLookup simplifies access to species distributions, habitat interactions, and climate-related events by converting user queries into precise graph queries, reducing the technical barriers for end users. The chatbot WildlifeLookup is available at https://oknbot.ngrok.dev/ Xiangqi Wang, Jason R. Rohr, Brett Scheffers, Nitesh V. Chawla, Xiangliang Zhang 0001 |
WSDM | 5 |
| 2024 | Traversing the Journey of Data and AI: From Convergence to TranslationabstractIn this talk, I will present our work on fundamental advances in AI, inspired by interdisciplinary problem statements and societal challenges. I will also highlight our innovation journey that encapsulates both the opportunities and challenges inherent in harnessing the full potential of AI for societal benefit, in particular highlighting the realization of societal impact through translational work and partnerships. Additionally, I will highlight our educational endeavors, emphasizing experiential learning and interdisciplinary approaches as fundamental elements of the student experience. Nitesh V. Chawla |
CIKM | 1 |
| 2024 | Application of Large Language Models in Chemistry Reaction Data Extraction and CleaningabstractChemical reaction data has existed and still largely exists in unstructured forms. But curating such information into datasets suitable for tasks such as yield and reaction outcome prediction is impractical via manual curation and not possible to automate through programmatic means alone. Large language models (LLMs) have emerged as potent tools, showcasing remarkable capabilities in processing textual information and therefore could be extremely useful in automating this process. To address the challenge of unstructured data, we manually curated a dataset of structured chemical reaction data to fine-tune and evaluate LLMs. We propose a paradigm that leverages prompt-tuning, fine-tuning techniques, and a verifier to check the extracted information. We evaluate the capabilities of various LLMs, including LLAMA-2 and GPT models with different parameter counts, on the data extraction task. Our results show that prompt tuning of GPT-4 yields the best accuracy and evaluation results. Fine-tuning LLAMA-2 models with hundreds of samples does enable them and organize scientific material according to user-defined schemas better though. This workflow shows an adaptable approach for chemical reaction data extraction but also highlights the challenges associated with nuance in chemical information. We open-sourced our code at https://github.com/joker-bruce/LLM_Extraction_Chem. Xiaobao Huang, Mihir Surve, Yuhan Liu 0010, Tengfei Luo, Olaf Wiest, Xiangliang Zhang 0001, Nitesh V. Chawla |
CIKM | 7 |
| 2024 | ChefFusion: Multimodal Foundation Model Integrating Recipe and Food Image GenerationabstractSignificant work has been conducted in the domain of food computing, yet these studies typically focus on single tasks such as t2t (instruction generation from food titles and ingredients), i2t (recipe generation from food images), or t2i (food image generation from recipes). None of these approaches integrate all modalities simultaneously. To address this gap, we introduce a novel food computing foundation model that achieves true multimodality, encompassing tasks such as t2t, t2i, i2t, it2t, and t2ti. By leveraging large language models (LLMs) and pre-trained image encoder and decoder models, our model can perform a diverse array of food computing-related tasks, including food understanding, food recognition, recipe generation, and food image generation. Compared to previous models, our foundation model demonstrates a significantly broader range of capabilities and exhibits superior performance, particularly in food image generation and recipe generation tasks. We open-sourced ChefFusion at https://github.com/Peiyu-Georgia-Li/ChefFusion-Multimodal-Foundation-Model-Integrating-Recipe-and-Food-Image-Generation.git. Xiaobao Huang, Yijun Tian 0001, Nitesh V. Chawla |
CIKM | 4 |
| 2024 | Data Augmentation's Effect on Machine Learning Models when Learning with Imbalanced DataabstractReal-world data is often imbalanced, such that the number of training instances varies by class. Data augmentation (DA) of under-represented classes is commonly used to improve model generalization in the face of class imbalance. Despite its ubiquity, the impact of data augmentation on machine learning (ML) models is not clearly understood. Here, we undertake a holistic examination of the effect of DA on under-represented classes. Unlike other studies, which focus on a single ML model type, we examine three different classifier families: convolutional neural networks, support vector machines, and logistic regression models; five different DA techniques and two different data modalities - image and tabular. Our research indicates that DA, when applied to imbalanced data, produces substantial changes in model weights, support vectors and front-end feature selection. These changes occur with respect to all classes, not just the ones that DA is applied to. Further, our empirical analysis shows that data augmentation's positive influence on generalization does not necessarily occur as a result of reducing weight norms. Rather, weight and support vector specialization play important roles in generalization. The specialization process may be a form of memorization that is spawned by variances introduced by augmented data. We investigate the seeming contradiction between improved generalization versus weight and support vector specialization. Damien Dablain, Nitesh V. Chawla |
DSAA | 2 |
| 2024 | Machine Learning in FinanceabstractThis workshop aims to explore the intersection of Generative AI with the rich tapestry of financial data types, seeking to uncover new methodologies and techniques that can enhance predictive analytics, fraud detection, and customer insights across the sector. By harnessing these advancements in AI, we can pave the way to not only understand customer behavior but also anticipate their needs more effectively, leading to superior customer outcomes and more personalized services. Our objective is to shed light on the challenges and opportunities presented by the diverse data formats in finance. We aim to bridge the gap between the dominance of traditional models for tabular data analysis and the emerging potential of Generative AI to revolutionize the treatment of time series, click streams, and other unstructured data forms. Leman Akoglu, Nitesh V. Chawla, Josep Domingo-Ferrer, Eren Kurshan, Senthil Kumar, Vidyut M. Naware, José A. Rodríguez-Serrano, Isha Chaturvedi, Saurabh Nagrecha, Mahashweta Das, Tanveer A. Faruquie |
KDD | 2 |
| 2024 | AnyLoss: Transforming Classification Metrics into Loss FunctionsabstractMany evaluation metrics can be used to assess the performance of models in binary classification tasks. However, most of them are derived from a confusion matrix in a non-differentiable form, making it very difficult to generate a differentiable loss function that could directly optimize them. The lack of solutions to bridge this challenge not only hinders our ability to solve difficult tasks, such as imbalanced learning, but also requires the deployment of computationally expensive hyperparameter search processes in model selection. In this paper, we propose a general-purpose approach that transforms any confusion matrix-based metric into a loss function, AnyLoss, that is available in optimization processes. To this end, we use an approximation function to make a confusion matrix represented in a differentiable form, and this approach enables any confusion matrix-based metric to be directly used as a loss function. The mechanism of the approximation function is provided to ensure its operability and the differentiability of our loss functions is proved by suggesting their derivatives. We conduct extensive experiments under diverse neural networks with many datasets, and we demonstrate their general availability to target any confusion matrix-based metrics. Our method, especially, shows outstanding achievements in dealing with imbalanced datasets, and its competitive learning speed, compared to multiple baseline models, underscores its efficiency. Do Heon Han, Nuno Moniz, Nitesh V. Chawla |
KDD | 3 |
| 2024 | A Survey of Large Language Models for GraphsabstractGraphs are an essential data structure utilized to represent relationships in real-world scenarios. Prior research has established that Graph Neural Networks (GNNs) deliver impressive outcomes in graph-centric tasks, such as link prediction and node classification. Despite these advancements, challenges like data sparsity and limited generalization capabilities continue to persist. Recently, Large Language Models (LLMs) have gained attention in natural language processing. They excel in language comprehension and summarization. Integrating LLMs with graph learning techniques has attracted interest as a way to enhance performance in graph learning tasks. In this survey, we conduct an in-depth review of the latest state-of-the-art LLMs applied in graph learning and introduce a novel taxonomy to categorize existing methods based on their framework design. We detail four unique designs: i) GNNs as Prefix, ii) LLMs as Prefix, iii) LLMs-Graphs Integration, and iv) LLMs-Only, highlighting key methodologies within each category. We explore the strengths and limitations of each framework, and emphasize potential avenues for future research, including overcoming current integration challenges between LLMs and graph learning techniques, and venturing into new application areas. This survey aims to serve as a valuable resource for researchers and practitioners eager to leverage large language models in graph learning, and to inspire continued progress in this dynamic field. We consistently maintain the related open-source materials at \url{https://github.com/HKUDS/Awesome-LLM4Graph-Papers}. Xubin Ren, Jiabin Tang, Dawei Yin 0001, Nitesh V. Chawla, Chao Huang 0001 |
KDD | 4 |
| 2024 | Graph Cross Supervised Learning via Generalized KnowledgeabstractThe success of GNNs highly relies on the accurate labeling of data. Existing methods of ensuring accurate labels, such as weakly-supervised learning, mainly focus on the existing nodes in the graphs. However, in reality, new nodes always continuously emerge on dynamic graphs, with different categories and even label noises. To this end, we formulate a new problem, Graph Cross-Supervised Learning, or Graph Weak-Shot Learning, that describes the challenges of modeling new nodes with novel classes and potential label noises. To solve this problem, we propose Lipshitz-regularized Mixture-of-Experts similarity network (LIME), a novel framework to encode new nodes and handle label noises. Specifically, we first design a node similarity network to capture the knowledge from the original classes, aiming to obtain insights for the emerging novel classes. Then, to enhance the similarity network's generalization to new nodes that could have a distribution shift, we employ the Mixture-of-Experts technique to increase the generalization of knowledge learned by the similarity network. To further avoid losing generalization ability during training, we introduce the Lipschitz bound to stabilize model output and alleviate the distribution shift issue. Empirical experiments validate LIME's effectiveness: we observe a substantial enhancement of up to 11.34% in node classification accuracy compared to the backbone model when subjected to the challenges of label noise on novel classes across five benchmark datasets. The code can be accessed through https://github.com/xiangchi-yuan/Graph-Cross-Supervised-Learning. Xiangchi Yuan, Yijun Tian 0001, Yanfang Ye 0001, Nitesh V. Chawla, Chuxu Zhang |
KDD | 5 |
| 2024 | Diet-ODIN: A Novel Framework for Opioid Misuse Detection with Interpretable Dietary PatternsabstractThe opioid crisis has been one of the most critical society concerns in the United States. Although the medication assisted treatment (MAT) is recognized as the most effective treatment for opioid misuse and addiction, the various side effects can trigger opioid relapse. In addition to MAT, the dietary nutrition intervention has been demonstrated its importance in opioid misuse prevention and recovery. However, research on the alarming connections between dietary patterns and opioid misuse remain under-explored. In response to this gap, in this paper, we first establish a large-scale multifaceted dietary benchmark dataset related to opioid users at the first attempt and then develop a novel framework - i.e., namely Opioid Misuse Detection with INterpretable Dietary Patterns (Diet-ODIN) - to bridge heterogeneous graph (HG) and large language model (LLM) for the identification of users with opioid misuse and the interpretation of their associated dietary patterns. Specifically, in Diet-ODIN, we first construct an HG to comprehensively incorporate both dietary and health-related information, and then we devise a holistic graph learning framework with noise reduction to fully capitalize both users' individual dietary habits and shared dietary patterns for the detection of users with opioid misuse. To further delve into the intricate correlations between dietary patterns and opioid misuse, we exploit an LLM by utilizing the knowledge obtained from the graph learning model for interpretation. The extensive experimental results based on our established benchmark with quantitative and qualitative measures demonstrate the outstanding performance of Diet-ODIN on exploring the complex interplay between opioid misuse and dietary patterns, by comparison with state-of-the-art baseline methods. Our code, built benchmark and system demo are available at https://github.com/JasonZhangzy1757/Diet-ODIN. Zheyuan Zhang 0008, Zehong Wang, Shifu Hou, Evan Hall, Landon Bachman, Jasmine White, Vincent Galassi, Nitesh V. Chawla, Chuxu Zhang, Yanfang Ye 0001 |
KDD | 8 |
| 2024 | RelKD 2024: The Second International Workshop on Resource-Efficient Learning for Knowledge DiscoveryabstractModern machine learning techniques, particularly deep learning, have showcased remarkable efficacy across numerous knowledge discovery and data mining applications. However, the advancement of many of these methods is frequently impeded by resource constraint challenges in many scenarios, such as limited labeled data (data-level), small model size requirements in real-world computing platforms (model-level), and efficient mapping of the computations to heterogeneous target hardware (system-level). Addressing all these factors is crucial for effectively and efficiently deploying developed models across a broad spectrum of real-world systems, including large-scale social network analysis, recommendation systems, and real-time anomaly detection. Therefore, there is a critical need to develop efficient learning techniques to address the challenges posed by resource limitations, whether from data, model/algorithm, or system/hardware perspectives. The proposed international workshop on "Resource-Efficient Learning for Knowledge Discovery (RelKD 2024)" will provide a great venue for academic researchers and industrial practitioners to share challenges, solutions, and future opportunities of resource-efficient learning. Chuxu Zhang, Dongkuan Xu, Kaize Ding, Jundong Li, Mojan Javaheripi, Subhabrata Mukherjee, Nitesh V. Chawla, Huan Liu 0001 |
KDD | 7 |
| 2024 | HetGPT: Harnessing the Power of Prompt Tuning in Pre-Trained Heterogeneous Graph Neural Networks
Yihong Ma, Ning Yan 0007, Jiayu Li 0002, Masood S. Mortazavi, Nitesh V. Chawla |
WWW | 5 |
| 2023 | Efficient Augmentation for Imbalanced Deep LearningabstractDeep learning models may not effectively generalize across under-represented or minority classes. We empirically study a convolutional neural network’s (CNN) internal representation of imbalanced image data and measure the generalization gap between a model’s feature embeddings in the training and test sets, showing that the gap is wider for minority classes. This insight enables us to design an efficient three-phase CNN training framework for imbalanced data. The framework involves training the network end-to-end on imbalanced data to learn feature embeddings, performing data augmentation in the learned embedding space to balance the training data distribution, and fine-tuning the classifier head on the embedded balanced training data. We develop Expansive Over-Sampling (EOS) as a data augmentation technique to utilize in the training framework. EOS forms synthetic training instances as convex combinations between the minority class samples and their nearest adversaries in the embedding space to reduce the generalization gap. The proposed framework improves the accuracy over leading cost-sensitive and resampling methods commonly used in imbalanced learning. Moreover, it is more computationally efficient than standard data pre-processing methods, such as SMOTE and GAN-based over-sampling, as it requires fewer parameters and less training time. The source code for the proposed framework is available at: https://github.com/dd1github/EOS. Damien Dablain, Colin Bellinger, Bartosz Krawczyk, Nitesh V. Chawla |
ICDE | 4 |
| 2023 | KDD Workshop on Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There is also the newly emerging threat of fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including self-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We plan to invite regular papers, positional papers and extended abstracts of work in progress. We will also encourage short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. Leman Akoglu, Nitesh V. Chawla, Senthil Kumar, Saurabh Nagrecha, Mahashweta Das, Vidyut M. Naware, Tanveer A. Faruquie |
KDD | 2 |
| 2023 | Foundations and Applications in Large-scale AI Models: Pre-training, Fine-tuning, and Prompt-based LearningabstractDeep learning techniques have advanced rapidly in recent years, leading to significant progress in pre-trained and fine-tuned large-scale AI models. For example, in the natural language processing domain, the traditional "pre-train, fine-tune" paradigm is shifting towards the "pre-train, prompt, and predict" paradigm, which has achieved great success on many tasks across different application domains such as ChatGPT/BARD for Conversational AI and P5 for a unified recommendation system. Moreover, there has been a growing interest in models that combine vision and language modalities (vision-language models) which are applied to tasks like Visual Captioning/Generation. Considering the recent technological revolution, it is essential to emphasize these paradigm shifts and highlight the paradigms with the potential to solve different tasks. We thus provide a platform for academic and industrial researchers to showcase their latest work, share research ideas, discuss various challenges, and identify areas where further research is needed in pre-training, fine-tuning, and prompt-learning methods for large-scale AI models. We foster the development of a strong research community focused on solving challenges related to large-scale AI models, providing superior and impactful strategies that can change people's lives in the future. Zhiyuan Cheng 0002, Dhaval Patel 0002, Linsey Pang, Sameep Mehta, Kexin Xie, Ed H. Chi, Wei Liu 0007, Nitesh V. Chawla, James Bailey 0001 |
KDD | 8 |
| 2023 | Modeling Co-Evolution of Attributed and Structural Information in Graph SequenceabstractMost graph neural network models learn embeddings of nodes in static attributed graphs for predictive analysis. Recent attempts have been made to learn temporal proximity of the nodes. We find that real dynamic attributed graphs exhibit complex phenomenon of co-evolution between node attributes and graph structure. Learning node embeddings for forecasting change of node attributes and evolution of graph structure over time remains an open problem. In this work, we present a novel framework called CoEvoGNN for modeling dynamic attributed graph sequence. It preserves the impact of earlier graphs on the current graph by embedding generation through the sequence of attributed graphs. It has a temporal self-attention architecture to model long-range dependencies in the evolution. Moreover, CoEvoGNN optimizes model parameters jointly on two dynamic tasks, attribute inference and link prediction over time. So the model can capture the co-evolutionary patterns of attribute change and link formation. This framework can adapt to any graph neural algorithms so we implemented and investigated three methods based on it: CoEvoGCN, CoEvoGAT, and CoEvoSAGE. Experiments demonstrate the framework (and its methods) outperforms strong baseline methods on predicting an entire unseen graph snapshot of personal attributes and interpersonal links in dynamic social graphs and financial graphs. Daheng Wang, Zhihan Zhang 0001, Yihong Ma, Tong Zhao 0003, Tianwen Jiang, Nitesh V. Chawla, Meng Jiang 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Hierarchical Spatio-Temporal Graph Neural Networks for Pandemic ForecastingabstractThe spread of COVID-19 throughout the world has led to cataclysmic consequences on the global community, which poses an urgent need to accurately understand and predict the trajectories of the pandemic. Existing research has relied on graph-structured human mobility data for the task of pandemic forecasting. To perform pandemic forecasting of COVID-19 in the United States, we curate Large-MG, a large-scale mobility dataset that contains 66 dynamic mobility graphs, with each graph having over 3k nodes and an average of 540k edges. One drawback with existing Graph Neural Networks (GNNs) for pandemic forecasting is that they generally perform information propagation in a flat way and thus ignore the inherent community structure in a mobility graph. To bridge this gap, we propose a Hierarchical Spatio-Temporal Graph Neural Network (HiSTGNN) to perform pandemic forecasting, which learns both spatial and temporal information from a sequence of dynamic mobility graphs. HiSTGNN consists of two network architectures. One is a hierarchical graph neural network (HiGNN) that constructs a two-level neural architecture: county-level and region-level, and performs information propagation in a hierarchical way. The other network architecture is a Transformer-based model that captures the temporal dynamics among the sequence of learned node representations from HiGNN. Additionally, we introduce a joint learning objective to further optimize HiSTGNN. Extensive experiments have demonstrated HiSTGNN's superior predictive power of COVID-19 new case/death counts compared with state-of-the-art baselines. Yihong Ma, Patrick Gérard, Yijun Tian 0001, Zhichun Guo, Nitesh V. Chawla |
CIKM | 5 |
| 2022 | Malicious Repositories Detection with Adversarial Heterogeneous Graph Contrastive LearningabstractGitHub, as the largest social coding platform, has attracted an increasing number of cybercriminals to disseminate malware by posting malicious code repositories. To address the imminent problem, some tools were developed to detect malicious repositories based on the code content. However, most of them ignore the rich relational information among repositories and usually require abundant labeled data to train the model. To this end, one effective way is to exploit unlabeled data to pre-train a model which considers both structural relation and code content of repositories, and further transfer the pre-trained model to the downstream tasks with labeled repository data. In this paper, we propose a novel model adversarial contrastive learning on heterogeneous graph (CLA-HG) to detect malicious repository in GitHub. First of all, CLA-HG builds a heterogeneous graph (HG) to comprehensively model repository data. Afterwards, to exploit unlabeled information in HG, CLA-HG introduces a dual-stream graph contrastive learning mechanism that distinguishes both adversarial subgraph pairs and standard subgraph pairs to pre-train graph neural networks using unlabeled data. Finally, the pre-trained model is fine-tuned to the downstream malicious repository detection task enhanced by a knowledge distillation (KD) module. Extensive experiments on two collected datasets from GitHub demonstrate the effectiveness of CLA-HG in comparison with state-of-the-art methods and popular commercial anti-malware products. Yiyue Qian, Yiming Zhang 0002, Nitesh V. Chawla, Yanfang Ye 0001, Chuxu Zhang |
CIKM | 3 |
| 2022 | Toward Graph Minimally-Supervised LearningabstractTo model graph-structured data, graph learning, in particular deep graph learning with graph neural networks, has drawn much attention in both academic and industrial communities lately. The effectiveness of prevailing graph learning methods usually rely on abundant labeled data for model training. However, it is common that graphs are scarcely labeled since data annotation and labeling on graphs is always time and resource-consuming. Therefore, it is imperative to investigate graph learning with minimal human supervision for the low-resource settings where limited or even no labeled data is available. In this tutorial, we will focus on the state-of-the-art techniques of Graph Minimally-Supervised Learning, in particular a series of weakly-supervised learning, few-shot learning, and self-supervised learning methods on graph-structured data as well as their real-world applications. The objectives of this tutorial are to: (1) formally categorize the problems in graph minimally-supervised learning and discuss the challenges under different learning scenarios; (2) comprehensively review the existing and recent advances of graph minimally-supervised learning; and (3) elucidate open questions and future research directions. This tutorial introduces major topics within minimally-supervised learning and offers a guide to a new frontier of graph learning. Kaize Ding, Chuxu Zhang, Jie Tang 0001, Nitesh V. Chawla, Huan Liu 0001 |
KDD | 4 |
| 2022 | KDD Workshop on Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There is also the newly emerging threat of fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including semi-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We plan to invite regular papers, positional papers and extended abstracts of work in progress. We will also encourage short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. Senthil Kumar, Leman Akoglu, Nitesh V. Chawla, Saurabh Nagrecha, Vidyut M. Naware, Tanveer A. Faruquie, Hays 'Skip' McCormick |
KDD | 3 |
| 2022 | Graph Minimally-supervised LearningabstractGraphs are widely used for abstracting complex systems of interacting objects, such as social networks, knowledge graphs, and traffic networks, as well as for modeling molecules, manifolds, and source code. To model such graph-structured data, graph learning, in particular deep graph learning with graph neural networks, has drawn much attention in both academic and industrial communities lately. Prevailing graph learning methods usually rely on learning from "big'' data, requiring a large amount of labeled data for model training. However, it is common that graphs are associated with "small'' labeled data as data annotation and labeling on graphs is always time and resource-consuming. Therefore, it is imperative to investigate graph learning with minimal human supervision for the low-resource settings where limited or even no labeled data is available. In this tutorial, we will focus on the state-of-the-art techniques of Graph Minimally-supervised Learning, in particular a series of weakly-supervised learning, few-shot learning, and self-supervised learning methods on graph-structured data as well as their real-world applications. The objectives of this tutorial are to: (1) formally categorize the problems in graph minimally-supervised learning and discuss the challenges under different learning scenarios; (2) comprehensively review the existing and recent advances of graph minimally-supervised learning; and (3) elucidate open questions and future research directions. This tutorial introduces major topics within minimally-supervised learning and offers a guide to a new frontier of graph learning. We believe this tutorial is beneficial to researchers and practitioners, allowing them to collaborate on graph learning. Kaize Ding, Jundong Li, Nitesh V. Chawla, Huan Liu 0001 |
WSDM | 3 |
| 2022 | AttrE2vec: Unsupervised attributed edge representation learning
Piotr Bielak, Tomasz Kajdanowicz, Nitesh V. Chawla |
Inf. Sci. | 3 |
| 2022 | Representation Learning on Variable Length and Incomplete Wearable-Sensory Time SeriesabstractThe prevalence of wearable sensors (e.g., smart wristband) is creating unprecedented opportunities to not only inform health and wellness states of individuals, but also assess and infer personal attributes, including demographic and personality attributes. However, the data captured from wearables, such as heart rate or number of steps, present two key challenges: (1) the time series is often of variable length and incomplete due to different data collection periods (e.g., wearing behavior varies by person); and (2) there is inter-individual variability to external factors like stress and environment. This article addresses these challenges and brings us closer to the potential of personalized insights about an individual, taking the leap from quantified self to qualified self. Specifically, HeartSpace proposed in this article learns embedding of the time-series data with variable length and missing values via the integration of a time-series encoding module and a pattern aggregation network. Additionally, HeartSpace implements a Siamese-triplet network to optimize representations by jointly capturing intra- and inter-series correlations during the embedding learning process. The empirical evaluation over two different real-world data presents significant performance gains over state-of-the-art baselines in a variety of applications, including user identification, personality prediction, demographics inference, job performance prediction, and sleep duration estimation. Xian Wu 0003, Chao Huang 0001, Pablo Robles-Granda, Nitesh V. Chawla |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2021 | Recipe Representation Learning with NetworksabstractLearning effective representations for recipes is essential in food studies for recommendation, classification, and other applications. Unlike what has been developed for learning textual or cross-modal embeddings for recipes, the structural relationship among recipes and food items are less explored. In this paper, we formalize the problem recipe representation learning with networks to involve both the textual feature and the structural relational feature into recipe representations. Specifically, we first present RecipeNet, a new and large-scale corpus of recipe data to facilitate network based food studies and recipe representation learning research. We then propose a novel heterogeneous recipe network embedding model, rn2vec, to learn recipe representations. The proposed model is able to capture textual, structural, and nutritional information through several neural network modules, including textual CNN, inner-ingredients transformer, and a graph neural network with hierarchical attention. We further design a combined objective function of node classification and link prediction to jointly optimize the model. The extensive experiments show that our model outperforms state-of-the-art baselines on two classic food study tasks. Dataset and codes are available at https://github.com/meettyj/rn2vec. Yijun Tian 0001, Chuxu Zhang, Ronald A. Metoyer, Nitesh V. Chawla |
CIKM | 4 |
| 2021 | motif2vec: Semantic-aware Representation Learning for Wearables' Time Series DataabstractThe proliferation of wearable sensors allows for the continuous collection of temporal characterization of an individual's physical activity and physiological data. This is enabling an unprecedented opportunity to delve into a deeper analysis of the underlying patterns of such temporal data and to infer attributes associated with health, behaviors, and well-being. However, there remain several challenges to fully discover both structural and temporal patterns (motifs) in these data streams and to leverage the semantic relationship among these motifs. These include: i) the temporal data of variable length and high resolution leads to the motifs of various sizes; ii) periodic occurrences and hierarchical overlaps of these motifs further challenge the modeling of their complex structural and semantic relations. We propose a semantic-aware unsupervised representation learning model, motif2vec, to learn the latent representation of time series data collected from wearable sensors. The motif2vec consists of three major components: 1) transforming the time series into a set of variable-length motif sequences; 2) formalizing random walks to construct the neighborhood of motifs and thus to extract structural and semantic relationship among motifs; 3) learning time series latent features to capture the motif neighborhood structure with a skip-gram model. Experiments on two real-world datasets, derived from two different wearables and population groups, show motif2vec outperforms six state-of-the-art benchmarks on various tasks. Suwen Lin, Xian Wu 0003, Nitesh V. Chawla |
DSAA | 3 |
| 2021 | Dynamic Attributed Graph Prediction with Conditional Normalizing FlowsabstractGraph representation learning aims at preserving structural and attributed information in latent representations. It has been studied mostly in the setting of static graph. In this work, we propose a novel approach for representation learning over dynamic attributed graph using the tool of normalizing flows for exact density estimation. Our approach has three components: (1) a time-aware graph neural component for aggregating graph information at each time step, (2) an adapted graph recurrent component for updating graph temporal contexts, and (3) a conditional normalizing flows component for capturing the evolution of node representations in latent space along time. Particularly, the third component has two sub-models of normalizing flows. One is used to capture the distribution of node representations of arbitrary complexity by considering graph temporal contexts as conditions. It learns invertible transformations to map node representations into simple priors conditioning on temporal contexts. The other one is dedicated to capture the evolutionary patterns of prior distributions. Extensive experiments demonstrate the proposed approach can outperform competitive baselines by a significant margin for dynamic link prediction on future graphs. Daheng Wang, Tong Zhao 0003, Nitesh V. Chawla, Meng Jiang 0001 |
ICDM | 3 |
| 2021 | Machine Learning in FinanceabstractThe finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There are also newly emerging threats such as fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including semi-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We have invited regular papers, positional papers and extended abstracts of work in progress. We have also encouraged short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. This event is the fourth in a sequence of finance related workshops we have organized at KDD since 2017. Senthil Kumar, Leman Akoglu, Nitesh V. Chawla, José A. Rodríguez-Serrano, Tanveer A. Faruquie, Saurabh Nagrecha |
KDD | 3 |
| 2021 | An Optimized NL2SQL System for Enterprise Data Mart
Kaiwen Dong, David A. Cieslak, Nitesh V. Chawla |
ECML/PKDD (5) | 5 |
| 2021 | Few-Shot Graph Learning for Molecular Property PredictionabstractThe recent success of graph neural networks has significantly boosted molecular property prediction, advancing activities such as drug discovery. The existing deep neural network methods usually require large training dataset for each property, impairing their performance in cases (especially for new molecular properties) with a limited amount of experimental data, which are common in real situations. To this end, we propose Meta-MGNN, a novel model for few-shot molecular property prediction. Meta-MGNN applies molecular graph neural network to learn molecular representations and builds a meta-learning framework for model optimization. To exploit unlabeled molecular information and address task heterogeneity of different molecular properties, Meta-MGNN further incorporates molecular structures, attribute based self-supervised modules and self-attentive task weights into the former framework, strengthening the whole learning model. Extensive experiments on two public multi-property datasets demonstrate that Meta-MGNN outperforms a variety of state-of-the-art methods. Zhichun Guo, Chuxu Zhang, Wenhao Yu 0002, John Herr, Olaf Wiest, Meng Jiang 0001, Nitesh V. Chawla |
WWW | 7 |
| 2021 | Modeling Complementarity in Behavior Data with Multi-Type Itemset EmbeddingabstractPeople are looking for complementary contexts, such as team members of complementary skills for project team building and/or reading materials of complementary knowledge for effective student learning, to make their behaviors more likely to be successful. Complementarity has been revealed by behavioral sciences as one of the most important factors in decision making. Existing computational models that learn low-dimensional context representations from behavior data have poor scalability and recent network embedding methods only focus on preserving the similarity between the contexts. In this work, we formulate a behavior entry as a set of context items and propose a novel representation learning method, Multi-type Itemset Embedding , to learn the context representations preserving the itemset structures. We propose a measurement of complementarity between context items in the embedding space. Experiments demonstrate both effectiveness and efficiency of the proposed method over the state-of-the-art methods on behavior prediction and context recommendation. We discover that the complementary contexts and similar contexts are significantly different in human behaviors. Daheng Wang, Qingkai Zeng 0001, Nitesh V. Chawla, Meng Jiang 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2021 | Inductive Contextual Relation Learning for PersonalizationabstractWeb personalization, e.g., recommendation or relevance search, tailoring a service/product to accommodate specific online users, is becoming increasingly important. Inductive personalization aims to infer the relations between existing entities and unseen new ones, e.g., searching relevant authors for new papers or recommending new items to users. This problem, however, is challenging since most of recent studies focus on transductive problem for existing entities. In addition, despite some inductive learning approaches have been introduced recently, their performance is sub-optimal due to relatively simple and inflexible architectures for aggregating entity’s content. To this end, we propose the inductive contextual personalization (ICP) framework through contextual relation learning. Specifically, we first formulate the pairwise relations between entities with a ranking optimization scheme that employs neural aggregator to fuse entity’s heterogeneous contents. Next, we introduce a node embedding term to capture entity’s contextual relations, as a smoothness constraint over the prior ranking objective. Finally, the gradient descent procedure with adaptive negative sampling is employed to learn the model parameters. The learned model is capable of inferring the relations between existing entities and inductive ones. Thorough experiments demonstrate that ICP outperforms numerous baseline methods for two different applications, i.e., relevant author search and new item recommendation. Chuxu Zhang, Huaxiu Yao, Lu Yu 0006, Chao Huang 0001, Dongjin Song, Meng Jiang 0001, Nitesh V. Chawla |
ACM Trans. Inf. Syst. | 8 |
| 2020 | Overcoming Data Sparsity in Predicting User Characteristics from Behavior through Graph EmbeddingsabstractUnderstanding user characteristics such as demographic information is useful for the personalization of online content promoted to users. However, it is difficult to obtain such data for each user visiting the website. Since demographic data for some users can be collected, their behavior can be used to predict the attributes of unknown users. Through online news consumption, we can infer the attributes of users from the articles they view. Most existing models take a supervised learning approach to this modeling task. However, by representing the user-URL interactions with a network, we can convert it to a semi-supervised learning problem and learn embeddings for users. Graph embeddings have become popular in recent years, with research mainly focusing on algorithmic developments. However, while we have an intuitive understanding of the problems they may overcome, such as data sparsity, this problem remains unexplored in the domain of demographic prediction using behavior. In this paper, we first investigate the effectiveness of using user embeddings generated from network representation learning for prediction by comparing its performance with other traditional feature sets, including content and item-based features. We find that the embeddings can represent a user generally on two prediction tasks, (1) gender prediction (classification) and (2) age prediction (regression). Second, we explore the advantages of using these embeddings over the other methods in two cases of data sparsity, where (1) the training and testing sets of users are temporally split and (2) the user labels are imbalanced. In both these cases, the embeddings outperform the baseline. Munira Syed, Daheng Wang, Meng Jiang 0001, Oliver Conway, Vishal Juneja, Sriram Subramanian, Nitesh V. Chawla |
ASONAM | 7 |
| 2020 | MBead: Semi-supervised Multilabel Behaviour Anomaly Detection on Multivariate Temporal Sensory DataabstractHuman abnormal physical and psychological behaviors, such as high level of stress, may result in negative impacts on work and life, if not handled efficiently. However, the continuous collection of behavioral data from questionnaires is not feasible, as is often the case for the natural downside of survey data gathering. Thanks to the proliferation of mobile sensors, it brings compelling opportunities for us to more deeply analyze human behavior. In this work, we ask the question of detecting anomalies in human physical and psychological behaviors from multivariate temporal data from multi-modal sensors. In the past decades, many efforts have been made in developing anomaly detection methods, but there remain several challenges in this specific domain problem: 1) data contains missing values at random positions. 2) data from multiple sensors is of multi-resolution and multivariate. 3) human behaviors are correlated to each other, thus it poses a multi-label problem. 4) the available labeled instances are limited, which requires the semi-supervised learning setting. 5) the frequency of anomaly occurrence is much smaller than that of normal instances, leading to imbalance problems. We propose a novel framework MBead to resolve these concerns. MBead consists of three key components: reweighted autoencoder to capture the dependency across temporal domain and multiple modalities, relevance learning module to learn the pairwise relations among labeled instances, and temporal prediction module to detect the anomalies while trained in semi-supervised settings. Extensive experiments show our MBead outperforms seven state-of-art baselines on three tasks of behavior anomaly detection: stress, affect, and work performance. Suwen Lin, Louis Faust, Sidney K. D'Mello, Gonzalo J. Martínez, Nitesh V. Chawla |
IEEE BigData | 5 |
| 2020 | GraSeq: Graph and Sequence Fusion Learning for Molecular Property PredictionabstractWith the recent advancement of deep learning, molecular representation learning -- automating the discovery of feature representation of molecular structure, has attracted significant attention from both chemists and machine learning researchers. Deep learning can facilitate a variety of downstream applications, including bio-property prediction, chemical reaction prediction, etc. Despite the fact that current SMILES string or molecular graph molecular representation learning algorithms (via sequence modeling and graph neural networks, respectively) have achieved promising results, there is no work to integrate the capabilities of both approaches in preserving molecular characteristics (e.g, atomic cluster, chemical bond) for further improvement. In this paper, we propose GraSeq, a joint graph and sequence representation learning model for molecular property prediction. Specifically, GraSeq makes a complementary combination of graph neural networks and recurrent neural networks for modeling two types of molecular inputs, respectively. In addition, it is trained by the multitask loss of unsupervised reconstruction and various downstream tasks, using limited size of labeled datasets. In a variety of chemical property prediction tests, we demonstrate that our GraSeq model achieves better performance than state-of-the-art approaches. Zhichun Guo, Wenhao Yu 0002, Chuxu Zhang, Meng Jiang 0001, Nitesh V. Chawla |
CIKM | 5 |
| 2020 | Personalized Imputation on Wearable-Sensory Time Series via Knowledge TransferabstractThe analysis of wearable-sensory time series data (e.g., heart rate records) benefits many applications (e.g., activity recognition, disease diagnosis). However, sensor measurements usually contain missing values due to various factors (e.g., user behavior, lack of charging), which may degrade the performance of downstream analytical tasks (e.g., regression, prediction). Thus, time series imputation is desired, which is capable of making sensory time series complete. Existing time series imputation methods generally employ various deep neural network models (e.g., GRU and GAN) to fill missing values by leveraging temporal patterns extracted from the contextual observations. Despite their effectiveness, we argue that most existing models can only achieve sub-optimal imputation performance due to the fact that they are inherently limited in sharing only one single set of model parameters to perform imputation on all individuals. Relying on one set of parameters limits the expressiveness of the imputation model as such models are bound to fail in capturing various complex personal characteristics. Therefore, most existing models tend to achieve inferior imputation performance, especially when a long duration of missing values, i.e., a large gap, is observed in the time series data. To address the limitation, this work develops a new imputation framework--Personalized Wearable-Sensory Time Series Imputation framework (PTSI) to provide a fully personalized treatment for time series imputation via effective knowledge transfer. In particular, PTSI first leverages a meta-learning paradigm to learn a well-generalized initialization to facilitate the adaption process for each user. To make the time series imputation be reflective of an individual's unique characteristics, we further endow PTSI with the capability of learning personalized model parameters, which is achieved by designing a parameter initialization modulating component. Extensive experiments on real-world human heart rate datasets demonstrate that our PTSI framework outperforms various state-of-the-art methods by a large margin consistently. Xian Wu 0003, Stephen M. Mattingly, Shayan Mirjafari, Chao Huang 0001, Nitesh V. Chawla |
CIKM | 5 |
| 2020 | Fighting a Pandemic: Convergence of Expertise, Data Science and PolicyabstractThis panel will address the challenges and opportunities of using data science to fight a pandemic. Of particular interest are real-world cases where using data science helped the fight against the pandemic and cautionary tales of when it hindered that fight. Tina Eliassi-Rad, Nitesh V. Chawla, Vittoria Colizza, Lauren Gardner, Marcel Salathé, Samuel V. Scarpino, Joseph T. Wu |
KDD | 2 |
| 2020 | Calendar Graph Neural Networks for Modeling Time Structures in Spatiotemporal User BehaviorsabstractUser behavior modeling is important for industrial applications such as demographic attribute prediction, content recommendation, and target advertising. Existing methods represent behavior log as a sequence of adopted items and find sequential patterns; however, concrete location and time information in the behavior log, reflecting dynamic and periodic patterns, joint with the spatial dimension, can be useful for modeling users and predicting their characteristics. In this work, we propose a novel model based on graph neural networks for learning user representations from spatiotemporal behavior data. Our model's architecture incorporates two networked structures. One is a tripartite network of items, sessions, and locations. The other is a hierarchical calendar network of hour, week, and weekday nodes. It first aggregates embeddings of location and items into session embeddings via the tripartite network, and then generates user embeddings from the session embeddings via the calendar structure. The user embeddings preserve spatial patterns and temporal patterns of a variety of periodicity (e.g., hourly, weekly, and weekday patterns). It adopts the attention mechanism to model complex interactions among the multiple patterns in user behaviors. Experiments on real datasets (i.e., clicks on news articles in a mobile app) show our approach outperforms strong baselines for predicting missing demographic attributes. Daheng Wang, Meng Jiang 0001, Munira Syed, Oliver Conway, Vishal Juneja, Sriram Subramanian, Nitesh V. Chawla |
KDD | 7 |
| 2020 | Multi-modal Network Representation LearningabstractIn today's information and computational society, complex systems are often modeled as multi-modal networks associated with heterogeneous structural relation, unstructured attribute/content, temporal context, or their combinations. The abundant information in multi-modal network requires both a domain understanding and large exploratory search space when doing feature engineering for building customized intelligent solutions in response to different purposes. Therefore, automating the feature discovery through representation learning in multi-modal networks has become essential for many applications. In this tutorial, we systematically review the area of multi-modal network representation learning, including a series of recent methods and applications. These methods will be categorized and introduced in the perspectives of unsupervised, semi-supervised and supervised learning, with corresponding real applications respectively. In the end, we conclude the tutorial and raise open discussions. The authors of this tutorial are active and productive researchers in this area. Chuxu Zhang, Meng Jiang 0001, Xiangliang Zhang 0001, Yanfang Ye 0001, Nitesh V. Chawla |
KDD | 5 |
| 2020 | Filling Missing Values on Wearable-Sensory Time Series DataabstractMissing data points is a common problem associated with data collected from wearables. This problem is particularly compounded if different subjects have different aspects of missingness associated with them – that is varying degrees of compliance behavior of individuals (participants) with respect to wearables as well as personal changes in lifestyle and health impacting heart rate. Moreover, despite the varying degree of compliance behavior, the wearable in itself might have glitches that lead to observations being dropped. Thus, any missing value imputation in such data has to not only generalize to the wearable behavior but also to the participant behavior. In this paper, we present a deep learning based approach for imputing missing values in heart rate time series data collected from a participant's wearable. In particular, for each participant, we first leverage his/her historical heart rate records as a reference set to extract the underlying personalized characteristics, and then impute the missing heart rate values by considering both contextual information of the current observations and the user's features learned from previous records. Adversarial training is applied to guide the learning process, which imputed more reasonable heart rate series with the consideration of human health conditions, e.g., heart rate fluctuations. Extensive experiments are conducted on two real-world data to show the superiority of our proposed method over state-of-the-art baselines. Suwen Lin, Xian Wu 0003, Gonzalo J. Martínez, Nitesh V. Chawla |
SDM | 4 |
| 2020 | Learning from Cross-Modal Behavior Dynamics with Graph-Regularized Neural Contextual BanditabstractContextual multi-armed bandit algorithms have received significant attention in modeling users’ preferences for online personalized recommender systems in a timely manner. While significant progress has been made along this direction, a few major challenges have not been well addressed yet: (i) a vast majority of the literature is based on linear models that cannot capture complex non-linear inter-dependencies of user-item interactions; (ii) existing literature mainly ignores the latent relations among users and non-recommended items: hence may not properly reflect users’ preferences in the real-world; (iii) current solutions are mainly based on historical data and are prone to cold-start problems for new users who have no interaction history. Xian Wu 0003, Suleyman Cetintas, Deguang Kong, Miao Lu, Jian Yang 0002, Nitesh V. Chawla |
WWW | 6 |
| 2020 | Hierarchically Structured Transformer Networks for Fine-Grained Spatial Event ForecastingabstractSpatial event forecasting is challenging and crucial for urban sensing scenarios, which is beneficial for a wide spectrum of spatial-temporal mining applications, ranging from traffic management, public safety, to environment policy making. In spite of significant progress has been made to solve spatial-temporal prediction problem, most existing deep learning based methods based on a coarse-grained spatial setting and the success of such methods largely relies on data sufficiency. In many real-world applications, predicting events with a fine-grained spatial resolution do play a critical role to provide high discernibility of spatial-temporal data distributions. However, in such cases, applying existing methods will result in weak performance since they may not well capture the quality spatial-temporal representations when training triple instances are highly imbalanced across locations and time. Xian Wu 0003, Chao Huang 0001, Chuxu Zhang, Nitesh V. Chawla |
WWW | 4 |
| 2019 | Similarity-Aware Network Embedding with Self-Paced LearningabstractNetwork embedding, which aims to learn low-dimensional vector representations for nodes in a network, has shown promising performance for many real-world applications, such as node classification and clustering. While various embedding methods have been developed for network data, they are limited in their assumption that nodes are correlated with their neighboring nodes with the same similarity degree. As such, these methods can be suboptimal for embedding network data. In this paper, we propose a new method named SANE, short for Similarity-Aware Network Embedding, to learn node representations by explicitly considering different similarity degrees between connected nodes in a network. In particular, we develop a new framework based on self-paced learning by accounting for both the explicit relations (i.e., observed links) and implicit relations (i.e., unobserved node similarities) in network representation learning. To justify our proposed model, we perform experiments on two real-world network data. Experiments results show that SNAE outperforms state-of-the-art embedding models on the tasks of node classification and node clustering. Chao Huang 0001, Baoxu Shi, Xuchao Zhang, Xian Wu 0003, Nitesh V. Chawla |
CIKM | 5 |
| 2019 | Deep Prototypical Networks for Imbalanced Time Series Classification under Data ScarcityabstractWith the increase of temporal data availability, time series classification has drawn a lot of attention in the literature because of its wide spectrum of applications in diverse domains (e.g., healthcare, bioinformatics and finance), ranging from human activity recognition to financial pattern identification. While significant progress has been made to solve time series classification problem, the success of such methods relies on data sufficiency, and may not well capture the quality embeddings when training triple instances are scarce and highly imbalance across classes. To address these challenges, we propose a prototype embedding framework-Deep Prototypical Networks (DPN), which leverages a main embedding space to capture the discrepancies of difference time series classes for alleviating data scarcity. In addition, we further augment DPN framework with a relationship-dependent masking module to automatically fuse relevant information with a distance metric learning process, which addresses the data imbalance issue and performs robust time series classification. Experimental results show significant and consistent improvements compared to state-of-the-art techniques. Chao Huang 0001, Xian Wu 0003, Xuchao Zhang, Suwen Lin, Nitesh V. Chawla |
CIKM | 5 |
| 2019 | Online Purchase Prediction via Multi-Scale Modeling of Behavior DynamicsabstractOnline purchase forecasting is of great importance in e-commerce platforms, which is the basis of how to present personalized interesting product lists to individual customers. However, predicting online purchases is not trivial as it is influenced by many factors including: (i) the complex temporal pattern with hierarchical inter-correlations; (ii) arbitrary category dependencies. To address these factors, we develop a Graph Multi-Scale Pyramid Networks (GMP) framework to fully exploit users' latent behavioral patterns with both multi-scale temporal dynamics and arbitrary inter-dependencies among product categories. In GMP, we first design a multi-scale pyramid modulation network architecture which seamlessly preserves the underlying hierarchical temporal factors--governing users' purchase behaviors. Then, we employ convolution recurrent neural network to encode the categorical temporal pattern at each scale. After that, we develop a resolution-wise recalibration gating mechanism to automatically re-weight the importance of each scale-view representations. Finally, a context-graph neural network module is proposed to adaptively uncover complex dependencies among category-specific purchases. Extensive experiments on real-world e-commerce datasets demonstrate the superior performance of our method over state-of-the-art baselines across various settings. Chao Huang 0001, Xian Wu 0003, Xuchao Zhang, Chuxu Zhang, Jiashu Zhao, Dawei Yin 0001, Nitesh V. Chawla |
KDD | 7 |
| 2019 | The Role of: A Novel Scientific Knowledge Graph Representation and Construction ModelabstractConditions play an essential role in scientific observations, hypotheses, and statements. Unfortunately, existing scientific knowledge graphs (SciKGs) represent factual knowledge as a flat relational network of concepts, as same as the KGs in general domain, without considering the conditions of the facts being valid, which loses important contexts for inference and exploration. In this work, we propose a novel representation of SciKG, which has three layers. The first layer has concept nodes, attribute nodes, as well as the attaching links from attribute to concept. The second layer represents both fact tuples and condition tuples. Each tuple is a node of the relation name, connecting to the subject and object that are concept or attribute nodes in the first layer. The third layer has nodes of statement sentences traceable to the original paper and authors. Each statement node connects to a set of fact tuples and/or condition tuples in the second layer. We design a semi-supervised Multi-Input Multi-Output sequence labeling model that learns complex dependencies between the sequence tags from multiple signals and generates output sequences for fact and condition tuples. It has a self-training module of multiple strategies to leverage the massive scientific data for better performance when manual annotation is limited. Experiments on a data set of 141M sentences show that our model outperforms existing methods and the SciKGs we constructed provide a good understanding of the scientific statements. Tianwen Jiang, Tong Zhao 0003, Bing Qin 0001, Ting Liu 0001, Nitesh V. Chawla, Meng Jiang 0001 |
KDD | 5 |
| 2019 | TUBE: Embedding Behavior Outcomes for Predicting SuccessabstractGiven a project plan and the goal, can we predict the plan's success rate? The key challenge is to learn the feature vectors of billions of the plan's components for effective prediction. However, existing methods did not model the behavior outcomes but component proximities. In this work, we define a measurement of behavior outcomes, which forms a test tube-shaped region to represent "success", in a vector space. We propose a novel representation learning method to learn the embeddings of behavior components (including contexts, plans, and goals) by preserving the behavior outcome information. Experiments on real datasets show that our proposed method significantly improves the performance of goal prediction as well as context recommendation over the state-of-the-art. Daheng Wang, Tianwen Jiang, Nitesh V. Chawla, Meng Jiang 0001 |
KDD | 3 |
| 2019 | Heterogeneous Graph Neural NetworkabstractRepresentation learning in heterogeneous graphs aims to pursue a meaningful vector representation for each node so as to facilitate downstream applications such as link prediction, personalized recommendation, node classification, etc. This task, however, is challenging not only because of the demand to incorporate heterogeneous structural (graph) information consisting of multiple types of nodes and edges, but also due to the need for considering heterogeneous attributes or contents (e.g., text or image) associated with each node. Despite a substantial amount of effort has been made to homogeneous (or heterogeneous) graph embedding, attributed graph embedding as well as graph neural networks, few of them can jointly consider heterogeneous structural (graph) information as well as heterogeneous contents information of each node effectively. In this paper, we propose HetGNN, a heterogeneous graph neural network model, to resolve this issue. Specifically, we first introduce a random walk with restart strategy to sample a fixed size of strongly correlated heterogeneous neighbors for each node and group them based upon node types. Next, we design a neural network architecture with two modules to aggregate feature information of those sampled neighboring nodes. The first module encodes "deep" feature interactions of heterogeneous contents and generates content embedding for each node. The second module aggregates content (attribute) embeddings of different neighboring groups (types) and further combines them by considering the impacts of different groups to obtain the ultimate node embedding. Finally, we leverage a graph context loss and a mini-batch gradient descent procedure to train the model in an end-to-end manner. Extensive experiments on several datasets demonstrate that HetGNN can outperform state-of-the-art baselines in various graph mining tasks, i.e., link prediction, recommendation, node classification & clustering and inductive node classification & clustering. Chuxu Zhang, Dongjin Song, Chao Huang 0001, Ananthram Swami, Nitesh V. Chawla |
KDD | 5 |
| 2019 | Neural Tensor Factorization for Temporal Interaction LearningabstractNeural collaborative filtering (NCF) and recurrent recommender systems (RRN) have been successful in modeling relational data (user-item interactions). However, they are also limited in their assumption of static or sequential modeling of relational data as they do not account for evolving users' preference over time as well as changes in the underlying factors that drive the change in user-item relationship over time. We address these limitations by proposing a Neural network based Tensor Factorization (NTF) model for predictive tasks on dynamic relational data. The NTF model generalizes conventional tensor factorization from two perspectives: First, it leverages the long short-term memory architecture to characterize the multi-dimensional temporal interactions on relational data. Second, it incorporates the multi-layer perceptron structure for learning the non-linearities between different latent factors. Our extensive experiments demonstrate the significant improvement in both the rating prediction and link prediction tasks on various dynamic relational data by our NTF model over both neural network based factorization models and other traditional methods. Xian Wu 0003, Baoxu Shi, Yuxiao Dong, Chao Huang 0001, Nitesh V. Chawla |
WSDM | 5 |
| 2019 | SHNE: Representation Learning for Semantic-Associated Heterogeneous NetworksabstractRepresentation learning in heterogeneous networks faces challenges due to heterogeneous structural information of multiple types of nodes and relations, and also due to the unstructured attribute or content (e.g., text) associated with some types of nodes. While many recent works have studied homogeneous, heterogeneous, and attributed networks embedding, there are few works that have collectively solved these challenges in heterogeneous networks. In this paper, we address them by developing a Semantic-aware Heterogeneous Network Embedding model (SHNE). SHNE performs joint optimization of heterogeneous SkipGram and deep semantic encoding for capturing both heterogeneous structural closeness and unstructured semantic relations among all nodes, as function of node content, that exist in the network. Extensive experiments demonstrate that SHNE outperforms state-of-the-art baselines in various heterogeneous network mining tasks, such as link prediction, document retrieval, node recommendation, relevance search, and class visualization. Chuxu Zhang, Ananthram Swami, Nitesh V. Chawla |
WSDM | 3 |
| 2019 | MiST: A Multiview and Multimodal Spatial-Temporal Learning Framework for Citywide Abnormal Event ForecastingabstractCitywide abnormal events, such as crimes and accidents, may result in loss of lives or properties if not handled efficiently. It is important for a wide spectrum of applications, ranging from public order maintaining, disaster control and people's activity modeling, if abnormal events can be automatically predicted before they occur. However, forecasting different categories of citywide abnormal events is very challenging as it is affected by many complex factors from different views: (i) dynamic intra-region temporal correlation; (ii) complex inter-region spatial correlations; (iii) latent cross-categorical correlations. In this paper, we develop a Multi-View and Multi-Modal Spatial-Temporal learning (MiST) framework to address the above challenges by promoting the collaboration of different views (spatial, temporal and semantic) and map the multi-modal units into the same latent space. Specifically, MiST can preserve the underlying structural information of multi-view abnormal event data and automatically learn the importance of view-specific representations, with the integration of a multi-modal pattern fusion module and a hierarchical recurrent framework. Extensive experiments on three real-world datasets, i.e., crime data and urban anomaly data, demonstrate the superior performance of our MiST method over the state-of-the-art baselines across various settings. Chao Huang 0001, Chuxu Zhang, Jiashu Zhao, Xian Wu 0003, Nitesh V. Chawla, Dawei Yin 0001 |
WWW | 5 |
| 2019 | Multi-Label Learning from CrowdsabstractWe consider multi-label crowdsourcing learning in two scenarios. In the first scenario, we aim at inferring instances' groundtruth given the crowds' annotations. We propose two approaches NAM/RAM (Neighborhood/Relevance Aware Multi-label crowdsourcing) modeling the crowds' expertise and label correlations from different perspectives. Extended from single-label crowdsourcing methods, NAM models the crowds' expertise on individual labels, but based on the idea that for rational workers, their annotations for instances similar in the feature space should also be similar, NAM utilizes information from the feature space and incorporates the local influence of neighborhoods' annotations. Noting that the crowds tend to act in an effort-saving manner while labeling multiple labels, i.e., rather than carefully annotating every proper label, they would prefer scanning and tagging a few most relevant labels, RAM models the crowds' expertise as their ability to distinguish the relevance between label pairs. In the second scenario, we care about cost-efficient crowdsourcing where the labeling and learning process are conducted in tandem. We extend NAM/RAM to the active paradigm and propose instance, label, and worker selection criteria such that the labeling cost is significantly saved compared to passive learning without labeling control. The proposals' effectiveness are validated on simulated and real data. Shao-Yuan Li, Yuan Jiang 0001, Nitesh V. Chawla, Zhi-Hua Zhou |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | DeepCrime: Attentive Hierarchical Recurrent Networks for Crime PredictionabstractAs urban crimes (e.g., burglary and robbery) negatively impact our everyday life and must be addressed in a timely manner, predicting crime occurrences is of great importance for public safety and urban sustainability. However, existing methods do not fully explore dynamic crime patterns as factors underlying crimes may change over time. In this paper, we develop a new crime prediction framework--DeepCrime, a deep neural network architecture that uncovers dynamic crime patterns and carefully explores the evolving inter-dependencies between crimes and other ubiquitous data in urban space. Furthermore, our DeepCrime framework is capable of automatically capturing the relevance of crime occurrences across different time periods. In particular, our DeepCrime framework enables predicting crime occurrences of different categories in each region of a city by i) jointly embedding all spatial, temporal, and categorical signals into hidden representation vectors, and ii) capturing crime dynamics with an attentive hierarchical recurrent network. Extensive experiments on real-world datasets demonstrate the superiority of our framework over many competitive baselines across various settings. Chao Huang 0001, Junbo Zhang 0004, Yu Zheng 0004, Nitesh V. Chawla |
CIKM | 4 |
| 2018 | RESTFul: Resolution-Aware Forecasting of Behavioral Time Series DataabstractLeveraging historical behavioral data (e.g., sales volume and email communication) for future prediction is of fundamental importance for practical domains ranging from sales to temporal link prediction. Current forecasting approaches often use only a single time resolution (e.g., daily or weekly), which truncates the range of observable temporal patterns. However, real-world behavioral time series typically exhibit patterns across multi-dimensional temporal patterns, yielding dependencies at each level. To fully exploit these underlying dynamics, this paper studies the forecasting problem for behavioral time series data with the consideration of multiple time resolutions and proposes a multi-resolution time series forecasting framework, RESolution-aware Time series Forecasting (RESTFul). In particular, we first develop a recurrent framework to encode the temporal patterns at each resolution. In the fusion process, a convolutional fusion framework is proposed, which is capable of learning conclusive temporal patterns for modeling behavioral time series data to predict future time steps. Our extensive experiments demonstrate that the RESTFul model significantly outperforms the state-of-the-art time series prediction techniques on both numerical and categorical behavioral time series data. Xian Wu 0003, Baoxu Shi, Yuxiao Dong, Chao Huang 0001, Louis Faust, Nitesh V. Chawla |
CIKM | 6 |
| 2018 | SMOTEBoost for Regression: Improving the Prediction of Extreme ValuesabstractSupervised learning with imbalanced domains is one of the biggest challenges in machine learning. Such tasks differ from standard learning tasks by assuming a skewed distribution of target variables, and user domain preference towards under-represented cases. Most research has focused on imbalanced classification tasks, where a wide range of solutions has been tested. Still, little work has been done concerning imbalanced regression tasks. In this paper, we propose an adaptation of the SMOTEBoost approach for the problem of imbalanced regression. Originally designed for classification tasks, it combines boosting methods and the SMOTE resampling strategy. We present four variants of SMOTEBoost and provide an experimental evaluation using 30 datasets with an extensive analysis of results in order to assess the ability of SMOTEBoost methods in predicting extreme target values, and their predictive trade-off concerning baseline boosting methods. SMOTEBoost is publicly available in a software package. Nuno Moniz, Rita P. Ribeiro, Vítor Cerqueira, Nitesh V. Chawla |
DSAA | 4 |
| 2018 | Multi-Type Itemset Embedding for Learning Behavior SuccessabstractContextual behavior modeling uses data from multiple contexts to discover patterns for predictive analysis. However, existing behavior prediction models often face difficulties when scaling for massive datasets. In this work, we formulate a behavior as a set of context items of different types (such as decision makers, operators, goals and resources), consider an observable itemset as a behavior success, and propose a novel scalable method, "multi-type itemset embedding", to learn the context items' representations preserving the success structures. Unlike most of existing embedding methods that learn pair-wise proximity from connection between a behavior and one of its items, our method learns item embeddings collectively from interaction among all multi-type items of a behavior, based on which we develop a novel framework, LearnSuc, for (1) predicting the success rate of any set of items and (2) finding complementary items which maximize the probability of success when incorporated into an itemset. Extensive experiments demonstrate both effectiveness and efficency of the proposed framework. Daheng Wang, Meng Jiang 0001, Qingkai Zeng 0001, Zachary Eberhart, Nitesh V. Chawla |
KDD | 5 |
| 2018 | ONE-M: Modeling the Co-evolution of Opinions and Network Connections
Aastha Nigam, Kijung Shin, Ashwin Bahulkar, Bryan Hooi, David Hachen, Boleslaw K. Szymanski, Christos Faloutsos, Nitesh V. Chawla |
ECML/PKDD (2) | 8 |
| 2018 | Who will Attend This Event Together? Event Attendance Prediction via Deep LSTM NetworksabstractEvent-based social network (EBSN) services have emerged as a new platform on which users can choose events of interest to attend in the physical world. Over years, there are growing research interests in predicting whether certain actors will participate in an event together. In this work, we refer to this task as the event attendance prediction problem and aim to address the predictability of individuals' event attendance. In real-world settings, the factors that influence an individual's attendance may change over time, leading to the dynamic nature of individuals' behavior. However, existing event attendance prediction methods cannot deal with such dynamic scenarios. To address this issue, we propose an end-to-end Deep Event Attendance Prediction (DEAP) framework—a three-level hierarchical LSTM architecture—to explicitly model users' multi-dimensional and evolving preferences. Extensive experiments on three real-world datasets demonstrate that DEAP significantly outperforms the state-of-the-art techniques across various settings. Xian Wu 0003, Yuxiao Dong, Baoxu Shi, Ananthram Swami, Nitesh V. Chawla |
SDM | 5 |
| 2018 | Camel: Content-Aware and Meta-path Augmented Metric Learning for Author IdentificationabstractIn this paper, we study the problem of author identification in big scholarly data, which is to effectively rank potential authors for each anonymous paper by using historical data. Most of the existing de-anonymization approaches predict relevance score of paper-author pair via feature engineering, which is not only time and storage consuming, but also introduces irrelevant and redundant features or miss important attributes. Representation learning can automate the feature generation process by learning node embeddings in academic network to infer the correlation of paper-author pair. However, the learned embeddings are often for general purpose (independent of the specific task), or based on network structure only (without considering the node content). To address these issues and make a further progress in solving the author identification problem, we propose Camel, a content-aware and meta-path augmented metric learning model. Specifically, first, the directly correlated paper-author pairs are modeled based on distance metric learning by introducing a push loss function. Next, the paper content embedding encoded by the gated recurrent neural network is integrated into the distance loss. Moreover, the historical bibliographic data of papers is utilized to construct an academic heterogeneous network, wherein a meta-path guided walk integrative learning module based on the task-dependent and content-aware Skipgram model is designed to formulate the correlations between each paper and its indirect author neighbors, and further augments the model. Extensive experiments demonstrate that Camel outperforms the state-of-the-art baselines. It achieves an average improvement of 6.3% over the best baseline method. Chuxu Zhang, Chao Huang 0001, Lu Yu 0006, Xiangliang Zhang 0001, Nitesh V. Chawla |
WWW | 5 |
| 2018 | Will Triadic Closure Strengthen Ties in Social Networks?abstractThe social triad—a group of three people—is one of the simplest and most fundamental social groups. Extensive network and social theories have been developed to understand its structure, such as triadic closure and social balance. Over the course of a triadic closure—the transition from two ties to three among three users, the strength dynamics of its social ties, however, are much less well understood. Using two dynamic networks from social media and mobile communication, we examine how the formation of the third tie in a triad affects the strength of the existing two ties. Surprisingly, we find that in about 80% social triads, the strength of the first two ties is weakened although averagely the tie strength in the two networks maintains an increasing or stable trend. We discover that (1) the decrease in tie strength among three males is more sharply than that among females, and (2) the tie strength between celebrities is more likely to be weakened as the closure of a triad than those between ordinary people. Furthermore, we formalize a triadic tie strength dynamics prediction problem to infer whether social ties of a triad will become weakened after its closure. We propose a TRIST method—a kernel density estimation (KDE)-based graphical model—to solve the problem by incorporating user demographics, temporal effects, and structural information. Extensive experiments demonstrate that TRIST offers a greater than 82% potential predictability for inferring triadic tie strength dynamics in both networks. The leveraging of the KDE and structural correlations enables TRIST to outperform baselines by up to 30% in terms of F1-score. Hong Huang 0001, Yuxiao Dong, Jie Tang 0001, Hongxia Yang, Nitesh V. Chawla, Xiaoming Fu 0001 |
ACM Trans. Knowl. Discov. Data | 5 |
| 2017 | Mining Features Associated with Effective TweetsabstractWhat tweet features are associated with higher effectiveness in tweets? Through the mining of 122 million engagements of 2.5 million original tweets, we present a systematic review of tweet time, entities, composition, and user account features. We show that the relationship between various features and tweeting effectiveness is non-linear; for example, tweets that use a few hashtags have higher effectiveness than using no or too many hashtags. This research closely relates to various industrial applications that are based on tweet features, including the analysis of advertising campaigns, the prediction of user engagement, the extraction of signals for automated trading, etc. Jian Xu 0019, Nitesh V. Chawla |
ASONAM | 2 |
| 2017 | Reliable fake review detection via modeling temporal and behavioral patternsabstractFake reviews have become a pervasive problem in online review systems, wherein fraudulent users manipulate the perception of an object (e.g., a restaurant) by fabricating fake reviews. Extensive work has been devoted to identifying fake reviews via modeling different factors separately, such as user features, object characteristics, and user-object bipartite relations. However, this problem remains challenging due to the fact that more advanced camouflage strategies are utilized by malicious users. In real-world scenarios, spammers may pretend to be normal users by giving fake reviews with the similar score distribution as normal users. To address these issues, we propose to explore the temporal patterns of users' review behavior, because spammers prefer to promote or demote the target businesses in a short period of time. In this work, we present a unified framework Reliable Fake Review Detection (RFRD) that explicitly models temporal patterns of users' review behavior into a probabilistic generative model. Moreover, the RFRD framework models users' underlying review credibility and objects' highly-skewed review distributions. We conduct experiments on two Yelp datasets, demonstrating the effectiveness of the proposed RFRD framework. Xian Wu 0003, Yuxiao Dong, Jun Tao 0002, Chao Huang 0001, Nitesh V. Chawla |
IEEE BigData | 5 |
| 2017 | ImWalkMF: Joint matrix factorization and implicit walk integrative learning for recommendationabstractData sparsity and cold-start problems are prevalent in recommender systems. To address such problems, both the observable explicit social information (e.g., user-user trust connections) and the inferable implicit correlations (e.g., implicit neighbors computed by similarity measurement) have been introduced to complement user-item ratings data for improving the performances of traditional model-based recommendation algorithms such as matrix factorization. Although effective, (1) the utilization of the explicit user-user social relationships suffers from the weakness of unavailability in real systems such as Netflix or the issue of sparse observable content like 0.03% trust density in Epinions, thus there is no or little explicit social information that can be employed to improve baseline model in real applications; (2) the current similarity measurement approaches focus on inferring implicit correlations between a user (item) and their direct neighbors or top-k similar neighbors based on user-item ratings bipartite network, so that they fail to comprehensively unfold the indirect potential relationships among users and items. To solve these issues regarding both explicit/implicit social recommendation algorithms, we design a joint model of matrix factorization and implicit walk integrative learning, i.e., ImWalkMF, which only uses explicit ratings information yet models both direct rating feedbacks and multiple direct/indirect implicit correlations among users and items from a random walk perspective. We further propose a combined strategy for training two independent components in the proposed model based on sampling. The experimental results on two real-world sparse datasets demonstrate that ImWalkMF outperforms the traditional regularized/probabilistic matrix factorization models as well as other competitive baselines that utilize explicit/implicit social information. Chuxu Zhang, Lu Yu 0006, Xiangliang Zhang 0001, Nitesh V. Chawla |
IEEE BigData | 4 |
| 2017 | Materials Science Literature-Patent Relevance Search: A Heterogeneous Network Analysis ApproachabstractIn recent decades, materials science literature and patents have grown exponentially. This has also contributed to an ever-growing challenge whether the literature is current, as there can be a gap between when the patent was filed and when it was approved. Moreover, it is difficult to ensure that a patent cites the appropriate prior art due to variety and volume of materials science data, especially when it is in two separate sources that have different curation mechanisms and purpose - publications and patents. The existing relational database schema, generally used to store publications, also presents challenges given the strict tabular schema, which may not be appropriate for organizing and querying highly interconnected information about materials in these publications and patents. For example, elements are chemically combined to form a compound, which can then be converted to other compounds via chemical reactions. Furthermore, relational database is not designed for handling combining data from multiple sources and with various formats, thus it makes discover relevance between publications and patents become difficult. In order to explore an alternative approach to represent materials data and combine data from multiple sources into the same repository, in this work, we propose a solution to integrate data from Open Quantum Materials Database (OQMD) and patent data from USPTO1 database into a network and named it heterogeneous materials information network (HMIN). We generalize prior work which based on using meta path-based topological features to explore the network, and we propose features to identify network noise and investigate relatedness between different-typed objects to meet our application needs. We built several machine learning models by using these features to explore relevance between materials science publications and patents. Experiment results show that HMIN can help researchers effectively discover related publications and patents originally kept in different sources. Our work exhibits to materials community a new way of appro-priately representing materials data and discovering connections between data from multiple sources. Pingjie Tang, Jed W. Pitera, Dmitry Zubarev, Nitesh V. Chawla |
DSAA | 4 |
| 2017 | metapath2vec: Scalable Representation Learning for Heterogeneous NetworksabstractWe study the problem of representation learning in heterogeneous networks. Its unique challenges come from the existence of multiple types of nodes and links, which limit the feasibility of the conventional network embedding techniques. We develop two scalable representation learning models, namely metapath2vec and metapath2vec++. The metapath2vec model formalizes meta-path-based random walks to construct the heterogeneous neighborhood of a node and then leverages a heterogeneous skip-gram model to perform node embeddings. The metapath2vec++ model further enables the simultaneous modeling of structural and semantic correlations in heterogeneous networks. Extensive experiments show that metapath2vec and metapath2vec++ are able to not only outperform state-of-the-art embedding models in various heterogeneous network mining tasks, such as node classification, clustering, and similarity search, but also discern the structural and semantic correlations between diverse network objects. Yuxiao Dong, Nitesh V. Chawla, Ananthram Swami |
KDD | 2 |
| 2017 | Structural Diversity and Homophily: A Study Across More Than One Hundred Big NetworksabstractA widely recognized organizing principle of networks is structural homophily, which suggests that people with more common neighbors are more likely to connect with each other. However, what influence the diverse structures embedded in common neighbors have on link formation is much less well-understood. To explore this problem, we begin by characterizing the structural diversity of common neighborhoods. Using a collection of 120 large-scale networks, we demonstrate that the impact of the common neighborhood diversity on link existence can vary substantially across networks. We find that its positive effect on Facebook and negative effect on LinkedIn suggest different underlying networking needs in these networks. We also discover striking cases where diversity violates the principle of homophily---that is, where fewer mutual connections may lead to a higher tendency to link with each other. We then leverage structural diversity to develop a common neighborhood signature (CNS), which we apply to a large set of networks to uncover unique network superfamilies not discoverable by conventional methods. Our findings shed light on the pursuit to understand the ways in which network structures are organized and formed, pointing to potential advancement in designing graph generation models and recommender systems. Yuxiao Dong, Reid A. Johnson, Jian Xu 0019, Nitesh V. Chawla |
KDD | 4 |
| 2017 | UAPD: Predicting Urban Anomalies from Spatial-Temporal Data
Xian Wu 0003, Yuxiao Dong, Chao Huang 0001, Jian Xu 0019, Dong Wang 0002, Nitesh V. Chawla |
ECML/PKDD (2) | 6 |
| 2017 | User Modeling on Demographic Attributes in Big Mobile Social NetworksabstractUsers with demographic profiles in social networks offer the potential to understand the social principles that underpin our highly connected world, from individuals, to groups, to societies. In this article, we harness the power of network and data sciences to model the interplay between user demographics and social behavior and further study to what extent users’ demographic profiles can be inferred from their mobile communication patterns. By modeling over 7 million users and 1 billion mobile communication records, we find that during the active dating period (i.e., 18--35 years old), users are active in broadening social connections with males and females alike, while after reaching 35 years of age people tend to keep small, closed, and same-gender social circles. Further, we formalize the demographic prediction problem of inferring users’ gender and age simultaneously. We propose a factor graph-based WhoAmI method to address the problem by leveraging not only the correlations between network features and users’ gender/age, but also the interrelations between gender and age. In addition, we identify a new problem—coupled network demographic prediction across multiple mobile operators—and present a coupled variant of the WhoAmI method to address its unique challenges. Our extensive experiments demonstrate the effectiveness, scalability, and applicability of the WhoAmI methods. Finally, our study finds a greater than 80% potential predictability for inferring users’ gender from phone call behavior and 73% for users’ age from text messaging interactions. Yuxiao Dong, Nitesh V. Chawla, Jie Tang 0001, Yang Yang 0009, Yang Yang 0008 |
ACM Trans. Inf. Syst. | 2 |
| 2016 | Link Prediction in a Semi-bipartite Network for Recommendation
Aastha Nigam, Nitesh V. Chawla |
ACIIDS (2) | 2 |
| 2016 | Analysis of link formation, persistence and dissolution in NetSense dataabstractWe study a unique behavioral network data set (based on periodic surveys and on electronic logs of dyadic contact via smartphones) collected at the University of Notre Dame. The participants are a sample of members of the entering class of freshmen in the fall of 2011 whose opinions on a wide variety of political and social issues and activities on campus were regularly recorded — at the beginning and end of each semester — for the first three years of their residence on campus. We create a communication activity network implied by call and text data, and a friendship network based on surveys. Both networks are limited to students participating in the NetSense surveys. We aim at finding student traits and activities on which agreements correlate well with formation and persistence of links while disagreements is highly correlated with non-existence or dissolution of links in the two social networks that we created. Using statistical analysis and machine learning, we observe several traits and activities displaying such correlations, thus being of potential use to predict social network evolution. Ashwin Bahulkar, Boleslaw K. Szymanski, Omar Lizardo, Yuxiao Dong, Yang Yang 0008, Nitesh V. Chawla |
ASONAM | 6 |
| 2016 | MedCare: Leveraging Medication Similarity for Disease PredictionabstractThe emergence of electronic health records (EHRs) has made medical history including past and current diseases, and prescribed medications easily available. This has facilitated development of personalized and population health care management systems. Contemporary disease prediction systems leverage data such as disease diagnoses codes to compute patients' similarity and predict the possible future disease risks of an individual. However, we posit that not all diseases (such as pre-existing conditions) may be represented in an EHR as a disease diagnosis code. It is likely that a patient is already taking a medication but does not have a corresponding disease in the EHR. To that end, we posit that the medication history can serve as a proxy for disease diagnoses, and ask the question whether medication and disease diagnoses combined together can improve the predictability of such systems. Building on our prior work in predicting disease risks (CARE), we develop two disease prediction systems: one using medication-based similarity (medCARE) and the other using both disease and medication-based similarity (combinedCARE). We show that combinedCARE provided a greater coverage and a higher average rank. Dipanwita Dasgupta, Nitesh V. Chawla |
DSAA | 2 |
| 2015 | Collaboration Signatures Reveal Scientific ImpactabstractCollaboration is an integral element of the scientific process that often leads to findings with significant impact. While extensive efforts have been devoted to quantifying and predicting research impact, the question of how collaborative behavior influences scientific impact remains unaddressed. In this work, we study the interplay between scientists' collaboration signatures and their scientific impact. As the basis of our study, we employ an ArnetMiner dataset with more than 1.7 million authors and 2 million papers spanning over 60 years. We formally define a scientist's collaboration signature as the distribution of collaboration strengths with each collaborator in his or her academic ego network, which is quantified by four measures: sociability, dependence, diversity, and self-collaboration. We then demonstrate that the collaboration signature allows us to effectively distinguish between researchers with dissimilar levels of scientific impact. We also discover that, even from the early stages of one's researcher career, a scientist's collaboration signature can help to reveal his or her future scientific impact. Finally, we find that as a representative group of outstanding computer scientists, Turing Award winners collectively produce distinctive collaboration signatures throughout the entirety of their careers. Our conclusions on the relationship between collaboration signatures and scientific impact give rise to important implications for researchers who wish to expand their scientific impact and more effectively stand on the shoulders of "collaborators." Yuxiao Dong, Reid A. Johnson, Yang Yang 0008, Nitesh V. Chawla |
ASONAM | 4 |
| 2015 | Recurrent Subgraph PredictionabstractInteractions in dynamic networks often transcend the dyadic barrier and emerge as subgraphs. The evolution of these subgraphs cannot be completely predicted using a pairwise link prediction analysis. We propose a novel solution to the problem---"Prediction of Recurrent Subgraphs (PReSub)" which treats subgraphs as individual entities in their own right. PReSub predicts re-occurring subgraphs using the network's vector space embedding and a set of "early warning subgraphs" which act as global and local descriptors of the subgraph's behavior. PReSub can be used as an out-of-the-box pipeline method with user-provided subgraphs or even to discover interesting subgraphs in an unsupervised manner. It can handle missing network information and is parallelizable. We show that PReSub outperforms traditional pairwise link prediction for a variety of evolving network datasets. The goal of this framework is to improve our understanding of subgraphs and provide an alternative representation in order to characterize their behavior. Saurabh Nagrecha, Nitesh V. Chawla, Horst Bunke |
ASONAM | 2 |
| 2015 | Predicting online video engagement using clickstreamsabstractAs access to broadband continues to grow along with the now almost ubiquitous availability of mobile phones, the landscape of the e-content delivery space has never been so dynamic. To establish their position in the market, businesses are beginning to realize that understanding each of their customers' likes and dislikes is perhaps as important as the offered content itself. Further, a number of companies are also delivering content, product previews, advertisements, etc. via video on their sites. The question remains - how effective are video engagement channels on sites? Can that user engagement be quantified? Clickstream data can furnish important insight into those questions using videos as a communication or messaging medium. To that end, focusing on a large set of web portals owned and managed by a private media company, we propose methods using these sites' clickstream data that can be used to provide a deeper understanding of their visitors, as well as their interests and preferences. We further expand the use of this data to show that it can be effectively used to predict user engagement to video streams, quantifying that metric by means of a survival analysis assessment. Everaldo Aguiar, Saurabh Nagrecha, Nitesh V. Chawla |
DSAA | 3 |
| 2015 | CoupledLP: Link Prediction in Coupled NetworksabstractWe study the problem of link prediction in coupled networks, where we have the structure information of one (source) network and the interactions between this network and another (target) network. The goal is to predict the missing links in the target network. The problem is extremely challenging as we do not have any information of the target network. Moreover, the source and target networks are usually heterogeneous and have different types of nodes and links. How to utilize the structure information in the source network for predicting links in the target network? How to leverage the heterogeneous interactions between the two networks for the prediction task? Yuxiao Dong, Jing Zhang 0001, Jie Tang 0001, Nitesh V. Chawla, Bai Wang 0001 |
KDD | 4 |
| 2015 | Optimizing Classifiers for Hypothetical Scenarios
Reid A. Johnson, Troy Raeder, Nitesh V. Chawla |
PAKDD (1) | 3 |
| 2015 | The Evolution of Social Relationships and Strategies Across the Lifespan
Yuxiao Dong, Nitesh V. Chawla, Jie Tang 0001, Yang Yang 0009, Yang Yang 0008 |
ECML/PKDD (3) | 2 |
| 2015 | Will This Paper Increase Your h-index?
Yuxiao Dong, Reid A. Johnson, Nitesh V. Chawla |
ECML/PKDD (3) | 3 |
| 2015 | Inferring Unusual Crowd Events from Mobile Phone Call Detail Records
Yuxiao Dong, Fabio Pinelli, Yiannis Gkoufas, Zubair Nabi, Francesco Calabrese, Nitesh V. Chawla |
ECML/PKDD (2) | 6 |
| 2015 | Will This Paper Increase Your h-index?: Scientific Impact PredictionabstractScientific impact plays a central role in the evaluation of the output of scholars, departments, and institutions. A widely used measure of scientific impact is citations, with a growing body of literature focused on predicting the number of citations obtained by any given publication. The effectiveness of such predictions, however, is fundamentally limited by the power-law distribution of citations, whereby publications with few citations are extremely common and publications with many citations are relatively rare. Given this limitation, in this work we instead address a related question asked by many academic researchers in the course of writing a paper, namely: "Will this paper increase my h-index?" Using a real academic dataset with over 1.7 million authors, 2 million papers, and 8 million citation relationships from the premier online academic service ArnetMiner, we formalize a novel scientific impact prediction problem to examine several factors that can drive a paper to increase the primary author's h-index. We find that the researcher's authority on the publication topic and the venue in which the paper is published are crucial factors to the increase of the primary author's h-index, while the topic popularity and the co-authors' h-indices are of surprisingly little relevance. By leveraging relevant factors, we find a greater than 87.5% potential predictability for whether a paper will contribute to an author's h-index within five years. As a further experiment, we generate a self-prediction for this paper, estimating that there is a 76% probability that it will contribute to the h-index of the co-author with the highest current h-index in five years. We conclude that our findings on the quantification of scientific impact can help researchers to expand their influence and more effectively leverage their position of "standing on the shoulders of giants." Yuxiao Dong, Reid A. Johnson, Nitesh V. Chawla |
WSDM | 3 |
| 2015 | Evaluating link prediction methods
Yang Yang 0008, Ryan Lichtenwalter, Nitesh V. Chawla |
Knowl. Inf. Syst. | 3 |
| 2014 | Inferring user demographics and social strategies in mobile social networksabstractDemographics are widely used in marketing to characterize different types of customers. However, in practice, demographic information such as age, gender, and location is usually unavailable due to privacy and other reasons. In this paper, we aim to harness the power of big data to automatically infer users' demographics based on their daily mobile communication patterns. Our study is based on a real-world large mobile network of more than 7,000,000 users and over 1,000,000,000 communication records (CALL and SMS). We discover several interesting social strategies that mobile users frequently use to maintain their social connections. First, young people are very active in broadening their social circles, while seniors tend to keep close but more stable connections. Second, female users put more attention on cross-generation interactions than male users, though interactions between male and female users are frequent. Third, a persistent same-gender triadic pattern over one's lifetime is discovered for the first time, while more complex opposite-gender triadic patterns are only exhibited among young people. Yuxiao Dong, Yang Yang 0008, Jie Tang 0001, Yang Yang 0009, Nitesh V. Chawla |
KDD | 5 |
| 2014 | Improving management of aquatic invasions by integrating shipping network, ecological, and environmental data: data mining for social goodabstractThe unintentional transport of invasive species (i.e., non-native and harmful species that adversely affect habitats and native species) through the Global Shipping Network (GSN) causes substantial losses to social and economic welfare (e.g., annual losses due to ship-borne invasions in the Laurentian Great Lakes is estimated to be as high as USD 800 million). Despite the huge negative impacts, management of such invasions remains challenging because of the complex processes that lead to species transport and establishment. Numerous difficulties associated with quantitative risk assessments (e.g., inadequate characterizations of invasion processes, lack of crucial data, large uncertainties associated with available data, etc.) have hampered the usefulness of such estimates in the task of supporting the authorities who are battling to manage invasions with limited resources. We present here an approach for addressing the problem at hand via creative use of computational techniques and multiple data sources, thus illustrating how data mining can be used for solving crucial, yet very complex problems towards social good. By modeling implicit species exchanges as a network that we refer to as the Species Flow Network (SFN), large-scale species flow dynamics are studied via a graph clustering approach that decomposes the SFN into clusters of ports and inter-cluster connections. We then exploit this decomposition to discover crucial knowledge on how patterns in GSN affect aquatic invasions, and then illustrate how such knowledge can be used to devise effective and economical invasive species management strategies. By experimenting on actual GSN traffic data for years 1997-2006, we have discovered crucial knowledge that can significantly aid the management authorities. Jian Xu 0019, Thanuka Wickramarathne, Nitesh V. Chawla, Erin K. Grey, Karsten Steinhaeuser, Reuben P. Keller, John M. Drake, David M. Lodge |
KDD | 3 |
| 2013 | Comparison of Gene Co-expression Networks and Bayesian Networks
Saurabh Nagrecha, Pawan Lingras, Nitesh V. Chawla |
ACIIDS (1) | 3 |
| 2013 | Link prediction in human mobility networksabstractThe understanding of how humans move is a longstanding challenge in the natural science. An important question is, to what degree is human behavior predictable? The ability to foresee the mobility of humans is crucial from predicting the spread of human to urban planning. Previous research has focused on predicting individual mobility behavior, such as the next location prediction problem. In this paper we study the human mobility behaviors from the perspective of network science. In the human mobility network, there will be a link between two humans if they are physically proximal to each other. We perform both microscopic and macroscopic explorations on the human mobility patterns. From the microscopic perspective, our objective is to answer whether two humans will be in proximity of each other or not. While from the macroscopic perspective, we are interested in whether we can infer the future topology of the human mobility network. In this paper we explore both problems by using link prediction technology, our methodology is demonstrated to have a greater degree of precision in predicting future mobility topology. Yang Yang 0008, Nitesh V. Chawla, Prithwish Basu, Bhaskar Prabhala, Thomas La Porta |
ASONAM | 2 |
| 2013 | Classifier Evaluation with Missing Negative Class Labels
Andrew K. Rider, Reid A. Johnson, Darcy A. Davis, T. Ryan Hoens, Nitesh V. Chawla |
IDA | 5 |
| 2013 | How Long Will She Call Me? Distribution, Social Theory and Duration Prediction
Yuxiao Dong, Jie Tang 0001, Tiancheng Lou, Bin Wu 0001, Nitesh V. Chawla |
ECML/PKDD (2) | 5 |
| 2013 | Reliable medical recommendation systems with patient privacyabstractOne of the concerns patients have when confronted with a medical condition is which physician to trust. Any recommendation system that seeks to answer this question must ensure that any sensitive medical information collected by the system is properly secured. In this article, we codify these privacy concerns in a privacy-friendly framework and present two architectures that realize it: the Secure Processing Architecture (SPA) and the Anonymous Contributions Architecture (ACA). In SPA, patients submit their ratings in a protected form without revealing any information about their data and the computation of recommendations proceeds over the protected data using secure multiparty computation techniques. In ACA, patients submit their ratings in the clear, but no link between a submission and patient data can be made. We discuss various aspects of both architectures, including techniques for ensuring reliability of computed recommendations and system performance, and provide their comparison. T. Ryan Hoens, Marina Blanton, Aaron Steele, Nitesh V. Chawla |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2012 | Link Prediction: Fair and Effective EvaluationabstractLink prediction is a popular area for publication. Papers appear in virtually every conference on data mining or network science with new methods. We argue that the practical performance potential of these methods is generally unknown because of challenges endemic to evaluation in many link prediction contexts. We demonstrate that current methods of evaluation are inadequate and can lead to woefully errant conclusions about practical performance potential. We argue for the use of precision-recall threshold curves and associated areas in lieu of receiver operating characteristic curves due to the extreme imbalance of the link prediction classification problem. We provide empirical examples of how current methods lead to questionable conclusions, how the fallacy of these conclusions is illuminated by methods we propose, and suggest a fair and consistent framework for link prediction evaluation for longitudinal and non-longitudinal network data sets. Ryan Lichtenwalter, Nitesh V. Chawla |
ASONAM | 2 |
| 2012 | Link Prediction and Recommendation across Heterogeneous Social NetworksabstractLink prediction and recommendation is a fundamental problem in social network analysis. The key challenge of link prediction comes from the sparsity of networks due to the strong disproportion of links that they have potential to form to links that do form. Most previous work tries to solve the problem in single network, few research focus on capturing the general principles of link formation across heterogeneous networks. In this work, we give a formal definition of link recommendation across heterogeneous networks. Then we propose a ranking factor graph model (RFG) for predicting links in social networks, which effectively improves the predictive performance. Motivated by the intuition that people make friends in different networks with similar principles, we find several social patterns that are general across heterogeneous networks. With the general social patterns, we develop a transfer-based RFG model that combines them with network structure information. This model provides us insight into fundamental principles that drive the link formation and network evolution. Finally, we verify the predictive performance of the presented transfer model on 12 pairs of transfer cases. Our experimental results demonstrate that the transfer of general social patterns indeed help the prediction of links. Yuxiao Dong, Jie Tang 0001, Sen Wu 0001, Jilei Tian, Nitesh V. Chawla, Jinghai Rao, Huanhuan Cao |
ICDM | 5 |
| 2012 | Predicting Links in Multi-relational and Heterogeneous NetworksabstractLink prediction is an important task in network analysis, benefiting researchers and organizations in a variety of fields. Many networks in the real world, for example social networks, are heterogeneous, having multiple types of links and complex dependency structures. Link prediction in such networks must model the influence propagating between heterogeneous relationships to achieve better link prediction performance than in homogeneous networks. In this paper, we introduce Multi-Relational Influence Propagation (MRIP), a novel probabilistic method for heterogeneous networks. We demonstrate that MRIP is useful for predicting links in sparse networks, which present a significant challenge due to the severe disproportion of the number of potential links to the number of real formed links. We also explore some factors that can inform the task of classification yet remain unexplored, such as temporal information. In this paper we make use of the temporal-related features by carefully investigating the issues of feasibility and generality. In accordance with our work in unsupervised learning, we further design an appropriate supervised approach in heterogeneous networks. Our experiments on co-authorship prediction demonstrate the effectiveness of our approach. Yang Yang 0008, Nitesh V. Chawla, Yizhou Sun, Jiawei Han 0001 |
ICDM | 2 |
| 2012 | Learning in non-stationary environments with class imbalanceabstractLearning in non-stationary environments is an increasingly important problem in a wide variety of real-world applications. In non-stationary environments data arrives incrementally, however the underlying generating function may change over time. In addition to the environments being non-stationary, they also often exhibit class imbalance. That is one class (the majority class) vastly outnumbers the other class (the minority class). This combination of class imbalance with non-stationary environments poses significant and interesting practical problems for classification. To overcome these issues, we introduce a novel instance selection mechanism, as well as provide a modification to the Heuristic Updatable Weighted Random Subspaces (HUWRS) method for the class imbalance problem. We then compare our modifications of HUWRS (called HUWRS.IP) to other state of the art algorithms, concluding that HUWRS. IP often achieves vastly superior performance. T. Ryan Hoens, Nitesh V. Chawla |
KDD | 2 |
| 2012 | Building Decision Trees for the Multi-class Imbalance Problem
T. Ryan Hoens, Qi Qian 0001, Nitesh V. Chawla, Zhi-Hua Zhou |
PAKDD (1) | 3 |
| 2012 | When will it happen?: relationship prediction in heterogeneous information networksabstractLink prediction, i.e., predicting links or interactions between objects in a network, is an important task in network analysis. Although the problem has attracted much attention recently, there are several challenges that have not been addressed so far. First, most existing studies focus only on link prediction in homogeneous networks, where all objects and links belong to the same type. However, in the real world, heterogeneous networks that consist of multi-typed objects and relationships are ubiquitous. Second, most current studies only concern the problem of whether a link will appear in the future but seldom pay attention to the problem of when it will happen. In this paper, we address both issues and study the problem of predicting when a certain relationship will happen in the scenario of heterogeneous networks. First, we extend the link prediction problem to the relationship prediction problem, by systematically defining both the target relation and the topological features, using a meta path-based approach. Then, we directly model the distribution of relationship building time with the use of the extracted topological features. The experiments on citation relationship prediction between authors on the DBLP network demonstrate the effectiveness of our methodology. Yizhou Sun, Jiawei Han 0001, Charu C. Aggarwal, Nitesh V. Chawla |
WSDM | 4 |
| 2012 | Vertex collocation profiles: subgraph counting for link analysis and predictionabstractWe introduce the concept of a vertex collocation profile (VCP) for the purpose of topological link analysis and prediction. VCPs provide nearly complete information about the surrounding local structure of embedded vertex pairs. The VCP approach offers a new tool for domain experts to understand the underlying growth mechanisms in their networks and to analyze link formation mechanisms in the appropriate sociological, biological, physical, or other context. The same resolution that gives VCP its analytical power also enables it to perform well when used in supervised models to discriminate potential new links. We first develop the theory, mathematics, and algorithms underlying VCPs. Then we demonstrate VCP methods performing link prediction competitively with unsupervised and supervised methods across several different network families. We conclude with timing results that introduce the comparative performance of several existing algorithms and the practicability of VCP computations on large networks. Ryan Lichtenwalter, Nitesh V. Chawla |
WWW | 2 |
| 2012 | Hellinger distance decision trees are robust and skew-insensitive
David A. Cieslak, T. Ryan Hoens, Nitesh V. Chawla, W. Philip Kegelmeyer |
Data Min. Knowl. Discov. | 3 |
| 2011 | Multi-relational Link Prediction in Heterogeneous Information NetworksabstractMany important real-world systems, modeled naturally as complex networks, have heterogeneous interactions and complicated dependency structures. Link prediction in such networks must model the influences between heterogenous relationships and distinguish the formation mechanisms of each link type, a task which is beyond the simple topological features commonly used to score potential links. In this paper, we introduce a novel probabilistically weighted extension of the Adamic/Adar measure for heterogenous information networks, which we use to demonstrate the potential benefits of diverse evidence, particularly in cases where homogeneous relationships are very sparse. However, we also expose some fundamental flaws of traditional a priori link prediction. In accordance with previous research on homogeneous networks, we further demonstrate that a supervised approach to link prediction can enhance performance and is easily extended to the heterogeneous case. Finally, we present results on three diverse, real-world heterogeneous information networks and discuss the trends and tradeoffs of supervised and unsupervised link prediction in a multi-relational setting. Darcy A. Davis, Ryan Lichtenwalter, Nitesh V. Chawla |
ASONAM | 3 |
| 2011 | DisNet: A Framework for Distributed Graph ComputationabstractWith the rise of network science as an exciting interdisciplinary research topic, efficient graph algorithms are in high demand. Problematically, many such algorithms measuring important properties of networks have asymptotic lower bounds that are quadratic, cubic, or higher in the number of vertices. For analysis of social networks, transportation networks, communication networks, and a host of others, computation is intractable. In these networks computation in serial fashion requires years or even decades. Fortunately, these same computational problems are often naturally parallel. We present here the design and implementation of a master-worker framework for easily computing such results in these circumstances. The user needs only to supply two small fragments of code describing the fundamental kernel of the computation. The framework automatically divides and distributes the workload and manages completion using an arbitrary number of heterogeneous computational resources. In practice, we have used thousands of machines and observed commensurate speedups. Writing only 31 lines of standard C++ code, we computed betweenness centrality on a network of 4.7M nodes in 25 hours. Ryan Lichtenwalter, Nitesh V. Chawla |
ASONAM | 2 |
| 2011 | Is Objective Function the Silver Bullet? A Case Study of Community Detection Algorithms on Social NetworksabstractCommunity detection or cluster detection in networks is a well-studied, albeit hard, problem. Given the scale and complexity of modern day social networks, detecting ``reasonable'' communities is an even harder problem. Since the first use of k-means algorithm in 1960s, many community detection algorithms have been invented - most of which are developed with specific goals in mind and the idea of detecting ``meaningful'' communities varies widely from one algorithm to another. With the increasing number of community detection algorithms, there has been an advent of a number of evaluation measures and objective functions such as modularity and internal density. In this paper we divide methods of measurements in to two categories, according to whether they rely on ground-truth or not. Our work is aiming to answer whether these general used objective functions are well consistent with the real performance of community detection algorithms across a number of homogeneous and heterogeneous networks. Seven representative algorithms are compared under various performance metrics, and on various real world social networks. Yang Yang 0008, Yizhou Sun, Saurav Pandit, Nitesh V. Chawla, Jiawei Han 0001 |
ASONAM | 4 |
| 2011 | Empirical comparison of correlation measures and pruning levels in complex networks representing the global climate systemabstractClimate change is an issue of growing economic, social, and political concern. Continued rise in the average temperatures of the Earth could lead to drastic climate change or an increased frequency of extreme events, which would negatively affect agriculture, population, and global health. One way of studying the dynamics of the Earth's changing climate is by attempting to identify regions that exhibit similar climatic behavior in terms of long-term variability. Climate networks have emerged as a strong analytics framework for both descriptive analysis and predictive modeling of the emergent phenomena. Previously, the networks were constructed using only one measure of similarity, namely the (linear) Pearson cross correlation, and were then clustered using a community detection algorithm. However, nonlinear dependencies are known to exist in climate, which begs the question whether more complex correlation measures are able to capture any such relationships. In this paper, we present a systematic study of different univariate measures of similarity and compare how each affects both the network structure as well as the predictive power of the clusters. Alex Pelan, Karsten Steinhaeuser, Nitesh V. Chawla, Dilkushi A. de Alwis Pitts, Auroop R. Ganguly |
CIDM | 3 |
| 2011 | Heuristic Updatable Weighted Random Subspaces for Non-stationary EnvironmentsabstractLearning in non-stationary environments is an increasingly important problem in a wide variety of real-world applications. In non-stationary environments data arrives incrementally, however the underlying generating function may change over time. While there is a variety of research into such environments, the research mainly consists of detecting concept drift (and then relearning the model), or developing classifiers which adapt to drift incrementally. We introduce Heuristic Up datable Weighted Random Subspaces (HUWRS), a new technique based on the Random Subspace Method that detects drift in individual features via the use of Hellinger distance, a distributional divergence metric. Through the use of subspaces, HUWRS allows for a more fine-grained approach to dealing with concept drift which is robust to feature drift even without class labels. We then compare our approach to two state of the art algorithms, concluding that for a wide range of datasets and window sizes HUWRS outperforms the other methods. T. Ryan Hoens, Nitesh V. Chawla, Robi Polikar |
ICDM | 2 |
| 2011 | Comparing Predictive Power in Climate Data: Clustering Matters
Karsten Steinhaeuser, Nitesh V. Chawla, Auroop R. Ganguly |
SSTD | 2 |
| 2010 | Consequences of Variability in Classifier Performance EstimatesabstractThe prevailing approach to evaluating classifiers in the machine learning community involves comparing the performance of several algorithms over a series of usually unrelated data sets. However, beyond this there are many dimensions along which methodologies vary wildly. We show that, depending on the stability and similarity of the algorithms being compared, these sometimes-arbitrary methodological choices can have a significant impact on the conclusions of any study, including the results of statistical tests. In particular, we show that performance metrics and data sets used, the type of cross-validation employed, and the number of iterations of cross-validation run have a significant, and often predictable, effect. Based on these results, we offer a series of recommendations for achieving consistent, reproducible results in classifier performance comparisons. Troy Raeder, T. Ryan Hoens, Nitesh V. Chawla |
ICDM | 3 |
| 2010 | New perspectives and methods in link predictionabstractThis paper examines important factors for link prediction in networks and provides a general, high-performance framework for the prediction task. Link prediction in sparse networks presents a significant challenge due to the inherent disproportion of links that can form to links that do form. Previous research has typically approached this as an unsupervised problem. While this is not the first work to explore supervised learning, many factors significant in influencing and guiding classification remain unexplored. In this paper, we consider these factors by first motivating the use of a supervised framework through a careful investigation of issues such as network observational period, generality of existing methods, variance reduction, topological causes and degrees of imbalance, and sampling approaches. We also present an effective flow-based predicting algorithm, offer formal bounds on imbalance in sparse network link prediction, and employ an evaluation method appropriate for the observed imbalance. Our careful consideration of the above issues ultimately leads to a completely general framework that outperforms unsupervised link prediction methods by more than 30% AUC. Ryan Lichtenwalter, Jake T. Lussier, Nitesh V. Chawla |
KDD | 3 |
| 2010 | Generating Diverse Ensembles to Counter the Problem of Class Imbalance
T. Ryan Hoens, Nitesh V. Chawla |
PAKDD (2) | 2 |
| 2010 | Privacy-Preserving Network Aggregation
Troy Raeder, Marina Blanton, Nitesh V. Chawla, Keith B. Frikken |
PAKDD (1) | 3 |
| 2010 | A Robust Decision Tree Algorithm for Imbalanced Data SetsabstractWe propose a new decision tree algorithm, Class Confidence Proportion Decision Tree (CCPDT), which is robust and insensitive to size of classes and generates rules which are statistically significant. In order to make decision trees robust, we begin by expressing Information Gain, the metric used in C4.5, in terms of confidence of a rule. This allows us to immediately explain why Information Gain, like confidence, results in rules which are biased towards the majority class. To overcome this bias, we introduce a new measure, Class Confidence Proportion (CCP), which forms the basis of CCPDT. To generate rules which are statistically significant we design a novel and efficient top-down and bottom-up approach which uses Fisher's exact test to prune branches of the tree which are not statistically significant. Together these two changes yield a classifier that performs statistically better than not only traditional decision trees but also trees learned from data that has been balanced by well known sampling techniques. Our claims are confirmed through extensive experiments and comparisons against C4.5, CART, HDDT and SPARCCC. Wei Liu 0007, Sanjay Chawla, David A. Cieslak, Nitesh V. Chawla |
SDM | 4 |
| 2010 | Time to CARE: a collaborative engine for practical disease prediction
Darcy A. Davis, Nitesh V. Chawla, Nicholas A. Christakis, Albert-László Barabási |
Data Min. Knowl. Discov. | 2 |
| 2009 | Modeling a Store's Product Space as a Social NetworkabstractA market basket is a set of products that form a single retail transaction. This purchase data of products can shed important light on how product(s) might influence sales of other product(s). Departing from the standard approach of frequent itemset mining, we posit that purchase data can be modeled as a social network. One can then discover communities of products that are bought together, which can lead to expressive exploration and discovery of a larger influence zone of product(s). We develop a novel utility measure for communities of products and show, both financially and intuitively, that community detection provides a useful complement to association rules for market basket analysis. All our conclusions are validated on real store data. Troy Raeder, Nitesh V. Chawla |
ASONAM | 2 |
| 2009 | A framework for monitoring classifiers' performance: when and why failure occurs?
David A. Cieslak, Nitesh V. Chawla |
Knowl. Inf. Syst. | 2 |
| 2008 | Predicting individual disease risk based on medical historyabstractThe monumental cost of health care, especially for chronic disease treatment, is quickly becoming unmanageable. This crisis has motivated the drive towards preventative medicine, where the primary concern is recognizing disease risk and taking action at the earliest signs. However, universal testing is neither time nor cost efficient. We propose CARE, a Collaborative Assessment and Recommendation Engine, which relies only on a patient's medical history using ICD-9-CM codes in order to predict future diseases risks. CARE uses collaborative filtering to predict each patient's greatest disease risks based on their own medical history and that of similar patients. We also describe an Iterative version, ICARE, which incorporates ensemble concepts for improved performance. These novel systems require no specialized information and provide predictions for medical conditions of all kinds in a single run. We present experimental results on a Medicare dataset, demonstrating that CARE and ICARE perform well at capturing future disease risks. Darcy A. Davis, Nitesh V. Chawla, Nicholas Blumm, Nicholas A. Christakis, Albert-László Barabási |
CIKM | 2 |
| 2008 | Start Globally, Optimize Locally, Predict Globally: Improving Performance on Imbalanced DataabstractClass imbalance is a ubiquitous problem in supervised learning and has gained wide-scale attention in the literature. Perhaps the most prevalent solution is to apply sampling to training data in order improve classifier performance. The typical approach will apply uniform levels of sampling globally. However, we believe that data is typically multi-modal, which suggests sampling should be treated locally rather than globally. It is the purpose of this paper to propose a framework which first identifies meaningful regions of data and then proceeds to find optimal sampling levels within each. This paper demonstrates that a global classifier trained on data locally sampled produces superior rank-orderings on a wide range of real-world and artificial datasets as compared to contemporary global sampling methods. David A. Cieslak, Nitesh V. Chawla |
ICDM | 2 |
| 2008 | Scaling up Classifiers to Cloud ComputersabstractAs the size of available datasets has grown from Megabytes to Gigabytes and now into Terabytes, machine learning algorithms and computing infrastructures have continuously evolved in an effort to keep pace. But at large scales, mining for useful patterns still presents challenges in terms of data management as well as computation. These issues can be addressed by dividing both data and computation to build ensembles of classifiers in a distributed fashion, but trade-offs in cost, performance, and accuracy must be considered when designing or selecting an appropriate architecture. In this paper, we present an abstraction for scalable data mining that allows us to explore these trade-offs. Data and computation are distributed to a computing cloud with minimal effort from the user, and multiple models for data management are available depending on the workload and system configuration. We demonstrate the performance and scalability characteristics of our ensembles using a wide variety of datasets and algorithms on a Condor-based pool with Chirp to handle the storage. Christopher Moretti, Karsten Steinhaeuser, Douglas Thain, Nitesh V. Chawla |
ICDM | 4 |
| 2008 | Analyzing PETs on Imbalanced Datasets When Training and Testing Class Distributions Differ
David A. Cieslak, Nitesh V. Chawla |
PAKDD | 2 |
| 2008 | Learning Decision Trees for Unbalanced Data
David A. Cieslak, Nitesh V. Chawla |
ECML/PKDD (1) | 2 |
| 2008 | Automatically countering imbalance and its empirical relationship to cost
Nitesh V. Chawla, David A. Cieslak, Lawrence O. Hall, Ajay Joshi |
Data Min. Knowl. Discov. | 1 |
| 2007 | A Black-Box Approach to Query Cardinality Estimation
Tanu Malik, Randal C. Burns, Nitesh V. Chawla |
CIDR | 3 |
| 2007 | Detecting Fractures in Classifier PerformanceabstractA fundamental tenet assumed by many classification algorithms is the presumption that both training and testing samples are drawn from the same distribution of data - this is the stationary distribution assumption. This entails that the past is strongly indicative of the future. However, in real world applications, many factors may alter the One True Model responsible for generating the data distribution both significantly and subtly. In circumstances violating the stationary distribution assumption, traditional validation schemes such as ten-folds and hold-out become poor performance predictors and classifier rankers. Thus, it becomes critical to discover the fracture points in classifier performance by discovering the divergence between populations. In this paper, we implement a comprehensive evaluation framework to identify bias, enabling selection of a "correct" classifier given the sample bias. To thoroughly evaluate the performance of classifiers within biased distributions, we consider the following three scenarios: missing completely at random (akin to stationary); missing at random; and missing not at random. The latter reflects the canonical sample selection bias problem. David A. Cieslak, Nitesh V. Chawla |
ICDM | 2 |
| 2006 | Evaluation of Summarization Schemes for Learning in StreamsabstractTraditional discretization techniques for machine learning, from examples with continuous feature spaces, are not efficient when the data is in the form of a stream from an unknown, possibly changing, distribution. We present a time-and-memory-efficient discretization technique based on computing ε -approximate exponential frequency quantiles, and prove bounds on the worst-case error introduced in computing information entropy in data streams compared to an offline algorithm that has no efficiency constraints. We compare the empirical performance of the technique, using it for feature selection, with (streaming adaptations of) two popular methods of discretization, equal width binning and equal frequency binning, under a variety of streaming scenarios for real and artificial datasets. Our experiments show that ε -approximate exponential frequency quantiles are remarkably consistent in their performance, in contrast to the simple and efficient equal width binning that perform quite well when the streams are from stationary distributions, and quite poorly otherwise. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Alec Pawling, Nitesh V. Chawla, Amitabh Chaudhary |
PKDD | 2 |
| 2003 | SMOTEBoost: Improving Prediction of the Minority Class in Boosting
Nitesh V. Chawla, Aleksandar Lazarevic, Lawrence O. Hall, Kevin W. Bowyer |
PKDD | 1 |
| 2001 | Creating Ensembles of ClassifiersabstractEnsembles of classifiers offer promise in increasing overall classification accuracy. The availability of extremely large datasets has opened avenues for application of distributed and/or parallel learning to efficiently learn models of them. In this paper, distributed learning is done by training classifiers on disjoint subsets of the data. We examine a random partitioning method to create disjoint subsets and propose a more intelligent way of partitioning into disjoint subsets using clustering. It was observed that the intelligent method of partitioning generally performs better than random partitioning for our datasets. In both methods a significant gain in accuracy may be obtained by applying bagging to each of the disjoint subsets, creating multiple diverse classifiers. The significance of our finding is that a partition strategy for even small/moderate sized datasets when combined with bagging can yield better performance than applying a single learner using the entire dataset. Nitesh V. Chawla, Steven Eschrich, Lawrence O. Hall |
ICDM | 1 |