Slobodan Vucetic

dblp:13/5776 · DBLP profile ↗
← Back
86ranked-venue papers
11as first author
13since 2021 · last 2026
0000-0001-5884-6293ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 31 · 7 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 4 since 2021Computer networks · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Lost in Translation: Understanding Autistic-Neurotypical Communication Style Differences in Job Postings
abstract
Autistic adults often use different communication styles than neurotypical individuals (NTs). While prior research has documented how such gaps disadvantage autistic job seekers, no study has systematically examined when these differences arise in language use and why autistic adults encounter interpretive gaps. This work seeks to datafy and characterize these communication challenges. We built an annotation interface and recruited 20 autistic adults to analyze 10 job postings each that they had selected as cases where they felt “lost in translation.” Participants annotated text spans using six categories informed by speech and language literature: unclear, ambiguous, incomplete, inappropriate, negative, and other. Follow-up interviews showed that lexical difficulties were rarely barriers; rather, challenges stemmed from interpreting implicit social arrangements or unstated expectations. We release the anonymized annotation data as the first-of-its-kind dataset documenting autistic–NT communication style differences. We conclude with implications for designing supports that foster clearer autistic–NT communication.
Huining Feng, Zinat Ara, Andrew Hundt, Slobodan Vucetic, John Joon Young Chung, Sungsoo Ray Hong
CHI4
2025 Towards a Unified Few-Shot Learning Evaluation Framework for RF Fingerprinting
abstract
Radio frequency (RF) fingerprinting is a technique used to identify a wireless device based on its specific and unique hardware characteristics. In recent years, deep learning has been utilized for RF fingerprinting due to its superiority in feature extraction and higher classification accuracy. However, one major challenge of deep learning-based RF fingerprinting is that wireless signals are highly sensitive to environmental conditions, causing the device fingerprints captured in one environment to not transfer well to another. Hence, deep learning models are found to perform well in the same condition but lose their ability to classify devices in the new condition. In this paper, we examine three transfer learning techniques to mitigate the domain shift problem in RF fingerprinting and compare them with two well-defined baselines. The three RF fingerprinting datasets under various scenarios are examined to explore how environmental factors impact RF fingerprinting, such as transmitter locations, transmitter distance, and device configurations. We identify the most challenging scenarios and study how environmental factors lead to model deterioration through t-SNE visualization.
Sai Shi, Vahid Mahzoon, Xuyu Wang, Shiwen Mao, Jie Wu 0001, Slobodan Vucetic
ICCCN6
2025 Deep Learning-Based Pedestrian Simulation with Limited Real-World Training Data: An Evaluation Framework
abstract
Simulating pedestrian movement is important for applications such as disaster management, robotics, and game design. While deep learning models have been extensively used on related problems, their use as pedestrian simulators remains relatively unexplored. This paper aims to encourage more research in this direction in two ways. First, it proposes an evaluation framework that is applicable to both traditional and deep learning based simulators. Second, it proposes and evaluates several ideas related to input representation, choice of neural architecture, exploiting knowledge-based simulators in data poor regimes, and repurposing trajectory prediction models. Our extensive experiments provide several useful insights for future research in pedestrian simulation. The code is available at https://github.com/vmahzoon76/DL-Crowd-Sim.
Vahid Mahzoon, Abigail Liu, Slobodan Vucetic
IJCAI3
2025 Exploring Engagement Opportunities for Autistic Children: Using AAC as a Controller in a Wizard-of-Oz Coloring Game
abstract
Autistic children face significant challenges in vocal communication and social interaction, often leading to social isolation. There is evidence that Augmentative and Alternative Communication (AAC) offers support to mitigate these challenges, enabling them to communicate with non-vocal means through forms of AAC, such as speech-generation devices (SGDs). However, the adoption and use of SGDs are hindered by several factors, including the large amount of practice required to learn to use SGDs and the limited options for highly engaging social learning contexts. Our study introduces the novel approach of using SGDs as game controller for digital and interactive games. With three design goals guiding our work, we conducted a Wizard-of-Oz formative case study with five participants aged 3-5 years, who were learning to use their SGD. We simulated a digital coloring game, integrating the speech-generated output of the participant's SGD to function as the game's controller. From this case study, we observed that all participants engaged with the game using their SGD for at least one turn, and two participants also engaged in emerging joint attention responses with the game and game's facilitator. This paper discusses these findings and contributes directions for future research, with suggestions for the design of future SGD-controlled games and exploration of social connection and collaboration between autistic children who use AAC and their caregivers, siblings, and peers.
Elizabeth Garrison, Stephen MacNeil, Elizabeth Lorah, Christine Holyfield, Slobodan Vucetic
Proc. ACM Hum. Comput. Interact.5
2024 Collaborative Job Seeking for People with Autism: Challenges and Design Opportunities
abstract
Successful job search results from job seekers' well-shaped social communication. While well-known diferences in communication exist between people with autism and neurotypicals, little is known about how people with autism collaborate with their social surroundings to strive in the job market. To better understand the practices and challenges of collaborative job seeking for people with autism, we interviewed 20 participants including applicants with autism, their social surroundings, and career experts. Through the interviews, we identified social challenges that people with autism face during their job seeking; the social support they leverage to be successful; and the technological limitations that hinder their collaboration. We designed four probes that represent major collaborative features found from the interviews-executive planning, communication, stage-wise preparation, and neurodivergent community formation-and discussed their potential usefulness and impact through three focus groups. We provide implications regarding how our findings can enhance collaborative job seeking experiences for people with autism through new designs.
Zinat Ara, Amrita Ganguly, Donna Peppard, Dongjun Chung, Slobodan Vucetic, Vivian Motti 0001, Sungsoo Ray Hong
CHI5
2024 CourtsightTV: An Interactive Visualization Software for Labeling Key Basketball Moments
abstract
Advancements in sensor technology are leading to massive collection of tracking data in sports. There is an increasing interest in analyzing the tracking data to gain competitive advantage. Analyzing and labeling key game moments can provide deep insights into player performance, team dynamics, as well as game strategy. However, the process of manually labeling and analyzing these moments is costly and time-consuming. In this paper, we describe a visual interface for user-friendly and efficient labeling of key moments in basketball games aided by neural networks. We report results of a user study evaluating the labeling interface.
Alexander Russakoff, Kenny Miller, Vahid Mahzoon, Parsa Esmaeilkhani, Christine Cho, Jaffar Alzeidi, Sandro Hauri, Slobodan Vucetic
CIKM8
2023 Group Activity Recognition in Basketball Tracking Data - Neural Embeddings in Team Sports (NETS)
abstract
Like many team sports, basketball involves two groups of players who engage in collaborative and adversarial activities to win a game. Players and teams are executing various complex strategies to gain an advantage over their opponents. Defining, identifying, and analyzing different types of activities is an important task in sports analytics, as it can lead to better strategies and decisions by the players and coaching staff. The objective of this paper is to automatically recognize basketball group activities from tracking data representing locations of players and the ball during a game. We propose a novel deep learning approach for group activity recognition (GAR) in team sports called NETS. To efficiently model the player relations in team sports, we combined a Transformer-based architecture with LSTM embedding, and a team-wise pooling layer to recognize the group activity. Training such a neural network generally requires a large amount of annotated data, which incurs high labeling cost. To alleviate this problem, we pretrain the neural network on a self-supervised trajectory prediction task and fine-tune it using a mix of strong and weak labels. We used a large tracking data set from 632 NBA games to evaluate our approach. The results show that NETS is capable of learning group activities with high accuracy, and that self- and weak-supervised training in NETS have a positive impact on GAR accuracy.
Sandro Hauri, Slobodan Vucetic
ECAI2
2022 MERIT: Minimal SupErvision Through Label Augmentation For Biomedical RelatIon ExTraction
Saman Enayati, Slobodan Vucetic
AMIA2
2022 OpenStance: Real-world Zero-shot Stance Detection
abstract
Prior studies of zero-shot stance detection identify the attitude of texts towards unseen topics occurring in the same document corpus.Such task formulation has three limitations: (i) Single domain/dataset.A system is optimized on a particular dataset from a single domain; therefore, the resulting system cannot work well on other datasets; (ii) the model is evaluated on a limited number of unseen topics; (iii) it is assumed that part of the topics has rich annotations, which might be impossible in real-world applications.These drawbacks will lead to an impractical stance detection system that fails to generalize to open domains and open-form topics.This work defines OpenStance: opendomain zero-shot stance detection, aiming to handle stance detection in an open world with neither domain constraints nor topicspecific annotations.The key challenge of OpenStance lies in the open-domain generalization: learning a system with fully unspecific supervision but capable of generalizing to any dataset.To solve OpenStance, we propose to combine indirect supervision, from textual entailment datasets, and weak supervision, from data generated automatically by pretrained Language Models.Our single system, without any topic-specific supervision, outperforms the supervised method on three popular datasets.To our knowledge, this is the first work that studies stance detection under the open-domain zero-shot setting.All data and code are publicly released.1
Hanzi Xu, Slobodan Vucetic, Wenpeng Yin 0001
CoNLL2
2021 Cannot Predict Comment Volume of a News Article before (a few) Users Read It
Lihong He 0001, Chen Shen 0009, Arjun Mukherjee, Slobodan Vucetic, Eduard C. Dragut
ICWSM4
2021 Sample Efficient Decentralized Stochastic Frank-Wolfe Methods for Continuous DR-Submodular Maximization
abstract
Continuous DR-submodular maximization is an important machine learning problem, which covers numerous popular applications. With the emergence of large-scale distributed data, developing efficient algorithms for the continuous DR-submodular maximization, such as the decentralized Frank-Wolfe method, became an important challenge. However, existing decentralized Frank-Wolfe methods for this kind of problem have the sample complexity of $\mathcal{O}(1/\epsilon^3)$, incurring a large computational overhead. In this paper, we propose two novel sample efficient decentralized Frank-Wolfe methods to address this challenge. Our theoretical results demonstrate that the sample complexity of the two proposed methods is $\mathcal{O}(1/\epsilon^2)$, which is better than $\mathcal{O}(1/\epsilon^3)$ of the existing methods. As far as we know, this is the first published result achieving such a favorable sample complexity. Extensive experimental results confirm the effectiveness of the proposed methods.
Hongchang Gao, Hanzi Xu, Slobodan Vucetic
IJCAI3
2021 Data Science with Human in the Loop
abstract
The aim of this workshop is to stimulate research on human-computer interaction challenges in data science. We invite researchers and practitioners interested in understanding how to optimize the human-computer cooperation and how to minimize human effort along the data science pipeline in a wide range of data science tasks and real-life applications. One over-arching challenge is to raise the level of abstraction of human-computer interaction to more sophisticated interaction models that better reflect a human's conceptual model and understanding. This workshop will bring together the interdisciplinary researchers from academia, research labs and practice to share, exchange, learn, and develop preliminary results, new concepts, ideas, principles, and methodologies on understanding and improving human-computer interaction for cost-effective development of data science models and for knowledge discovery. We expect the workshop to help develop and grow a strong community of researchers who are interested in this topic, and yield future collaborations and scientific exchanges across the relevant areas of data mining, machine learning, data and knowledge management, human-machine interaction, and user interfaces.
Eduard C. Dragut, Yunyao Li 0001, Lucian Popa 0001, Slobodan Vucetic
KDD4
2021 Multi-Modal Trajectory Prediction of NBA Players
abstract
National Basketball Association (NBA) players are highly motivated and skilled experts that solve complex decision making problems at every time point during a game. As a step towards understanding how players make their decisions, we focus on their movement trajectories during games. We propose a method that captures the multi-modal behavior of players, where they might consider multiple trajectories and select the most advantageous one. The method is built on an LSTM-based architecture predicting multiple trajectories and their probabilities, trained by a multi-modal loss function that updates the best trajectories. Experiments on large, fine-grained NBA tracking data show that the proposed method outperforms the state-of-the-art. In addition, the results indicate that the approach generates more realistic trajectories and that it can learn individual playing styles of specific players.
Sandro Hauri, Nemanja Djuric, Vladan Radosavljevic, Slobodan Vucetic
WACV4
2020 Improving Word Embeddings through Iterative Refinement of Word- and Character-level Models
abstract
Embedding of rare and out-of-vocabulary (OOV) words is an important open NLP problem.A popular solution is to train a character-level neural network to reproduce the embeddings from a standard word embedding model.The trained network is then used to assign vectors to any input string, including OOV and rare words.We enhance this approach and introduce an algorithm that iteratively refines and improves both word-and character-level models.We demonstrate that our method outperforms the existing algorithms on 5 word similarity data sets, and that it can be successfully applied to job title normalization, an important problem in the e-recruitment domain that suffers from the OOV problem.
Phong Ha, Shanshan Zhang 0004, Nemanja Djuric, Slobodan Vucetic
COLING4
2020 Growing Adaptive Multi-hyperplane Machines
abstract
Adaptive Multi-hyperplane Machine (AMM) is an online algorithm for learning Multi-hyperplane Machine (MM), a classification model which allows multiple hyperplanes per class. AMM is based on Stochastic Gradient Descent (SGD), with training time comparable to linear Support Vector Machine (SVM) and significantly higher accuracy. On the other hand, empirical results indicate there is a large accuracy gap between AMM and non-linear SVMs. In this paper we show that this performance gap is not due to limited representability of the MM model, as it can represent arbitrary concepts. We set to explain the connection between the AMM and Learning Vector Quantization (LVQ) algorithms, and introduce a novel Growing AMM (GAMM) classifier motivated by Growing LVQ, that imputes duplicate hyperplanes into the MM model during SGD training. We provide theoretical results showing that GAMM has favorable convergence properties, and analyze the generalization bound of the MM models. Experiments indicate that GAMM achieves significantly improved accuracy on non-linear problems, with only slightly slower training compared to AMM. On some tasks GAMM comes close to non-linear SVM, and outperforms other popular classifiers such as Neural Networks and Random Forests.
Nemanja Djuric, Slobodan Vucetic
ICML3
2019 Spatial Aggregation Facilitates Discovery of Spatial Topics
abstract
Spatial aggregation refers to merging of documents created at the same spatial location.We show that by spatial aggregation of a large collection of documents and applying a traditional topic discovery algorithm on the aggregated data we can efficiently discover spatially distinct topics.By looking at topic discovery through matrix factorization lenses we show that spatial aggregation allows low rank approximation of the original document-word matrix, in which spatially distinct topics are preserved and non-spatial topics are aggregated into a single topic.Our experiments on synthetic data confirm this observation.Our experiments on 4.7 million tweets collected during the Sandy Hurricane in 2012 show that spatial and temporal aggregation allows rapid discovery of relevant spatial and temporal topics during that period.Our work indicates that different forms of document aggregation might be effective in rapid discovery of various types of distinct topics from large collections of documents.
Aniruddha Maiti, Slobodan Vucetic
ACL (1)2
2019 Medical Concept Representation Learning from Multi-source Data
abstract
Representing words as low dimensional vectors is very useful in many natural language processing tasks. This idea has been extended to medical domain where medical codes listed in medical claims are represented as vectors to facilitate exploratory analysis and predictive modeling. However, depending on a type of a medical provider, medical claims can use medical codes from different ontologies or from a combination of ontologies, which complicates learning of the representations. To be able to properly utilize such multi-source medical claim data, we propose an approach that represents medical codes from different ontologies in the same vector space. We first modify the Pointwise Mutual Information (PMI) measure of similarity between the codes. We then develop a new negative sampling method for word2vec model that implicitly factorizes the modified PMI matrix. The new approach was evaluated on the code cross-reference problem, which aims at identifying similar codes across different ontologies. In our experiments, we evaluated cross-referencing between ICD-9 and CPT medical code ontologies. Our results indicate that vector representations of codes learned by the proposed approach provide superior cross-referencing when compared to several existing approaches.
Tian Bai 0001, Brian L. Egleston, Richard Bleicher, Slobodan Vucetic
IJCAI4
2019 How to Invest my Time: Lessons from Human-in-the-Loop Entity Extraction
abstract
Recognizing entities that follow or closely resemble a regular expression (regex) pattern is an important task in information extraction. Common approaches for extraction of such entities require humans to either write a regex recognizing an entity or manually label entity mentions in a document corpus. While human effort is critical to build an entity recognition model, surprisingly little is known about how to best invest that effort given a limited time budget. To get an answer, we consider an iterative human-in-the-loop (HIL) framework that allows users to write a regex or manually label entity mentions, followed by training and refining a classifier based on the provided information. We demonstrate on 5 entity recognition tasks that classification accuracy improves over time with either approach. When a user is allowed to choose between regex construction and manual labeling, we discover that (1) if the time budget is low, spending all time for regex construction is often advantageous, (2) if the time budget is high, spending all time for manual labeling seems to be superior, and (3) between those two extremes, writing regexes followed by manual labeling is typically the best approach. Our code and data is available at https://github.com/nymph332088/HILRecognizer.
Shanshan Zhang 0004, Lihong He 0001, Eduard C. Dragut, Slobodan Vucetic
KDD4
2019 Storage on the Edge: Evaluating Cloud Backed Edge Storage in Cyberphysical Systems
abstract
Effective control of emerging cyberphysical systems such as smart transportation, smart health-care, etc. requires edge computing infrastructure that is often organized into three layers, namely edge (IoT) devices, edge controllers (ECs) and the cloud. In large infrastructures, ECs must be deployed densely in the proximity of edge devices and need to satisfy strict constraints on cost, size, cooling, etc. Thus, ECs cannot host large amounts of local storage and instead must make use of cloud storage in the background to provide an impression of large, fast local storage to host the IoT device data needed for online and real-time queries. In this paper, we provide insights into the configuration issues of such an edge storage infrastructure (ESI) based on the evaluation of commercial ESIs on several real-world edge computing workloads. We also show that the current ESI designs are lacking in several respects, and suggest some approaches for enhancing their capabilities to meet the stringent requirements of emerging edge computing applications.
Sanjeev Sondur, Krishna Kant 0001, Slobodan Vucetic, Brandon Byers
MASS3
2019 Improving Medical Code Prediction from Clinical Text via Incorporating Online Knowledge Sources
abstract
Clinical notes contain detailed information about health status of patients for each of their encounters with a health system. Developing effective models to automatically assign medical codes to clinical notes has been a long-standing active research area. Despite a great recent progress in medical informatics fueled by deep learning, it is still a challenge to find the specific piece of evidence in a clinical note which justifies a particular medical code out of all possible codes. Considering the large amount of online disease knowledge sources, which contain detailed information about signs and symptoms of different diseases, their risk factors, and epidemiology, there is an opportunity to exploit such sources. In this paper we consider Wikipedia as an external knowledge source and propose Knowledge Source Integration (KSI), a novel end-to-end code assignment framework, which can integrate external knowledge during training of any baseline deep learning model. The main idea of KSI is to calculate matching scores between a clinical note and disease related Wikipedia documents, and combine the scores with output of the baseline model. To evaluate KSI, we experimented with automatic assignment of ICD-9 diagnosis codes to the emergency department clinical notes from MIMIC-III data set, aided by Wikipedia documents corresponding to the ICD-9 codes. We evaluated several baseline models, ranging from logistic regression to recently proposed deep learning models known to achieve the state-of-the-art accuracy on clinical notes. The results show that KSI consistently improves the baseline models and that it is particularly successful in assignment of rare codes. In addition, by analyzing weights of KSI models, we can gain understanding about which words in Wikipedia documents provide useful information for predictions.
Tian Bai 0001, Slobodan Vucetic
WWW2
2019 A new clustering and nomenclature for beta turns derived from high-resolution protein structures
abstract
Protein loops connect regular secondary structures and contain 4-residue beta turns which represent 63% of the residues in loops. The commonly used classification of beta turns (Type I, I', II, II', VIa1, VIa2, VIb, and VIII) was developed in the 1970s and 1980s from analysis of a small number of proteins of average resolution, and represents only two thirds of beta turns observed in proteins (with a generic class Type IV representing the rest). We present a new clustering of beta-turn conformations from a set of 13,030 turns from 1074 ultra-high resolution protein structures (≤1.2 Å). Our clustering is derived from applying the DBSCAN and k-medoids algorithms to this data set with a metric commonly used in directional statistics applied to the set of dihedral angles from the second and third residues of each turn. We define 18 turn types compared to the 8 classical turn types in common use. We propose a new 2-letter nomenclature for all 18 beta-turn types using Ramachandran region names for the two central residues (e.g., 'A' and 'D' for alpha regions on the left side of the Ramachandran map and 'a' and 'd' for equivalent regions on the right-hand side; classical Type I turns are 'AD' turns and Type I' turns are 'ad'). We identify 11 new types of beta turn, 5 of which are sub-types of classical beta-turn types. Up-to-date statistics, probability densities of conformations, and sequence profiles of beta turns in loops were collected and analyzed. A library of turn types, BetaTurnLib18, and cross-platform software, BetaTurnTool18, which identifies turns in an input protein structure, are freely available and redistributable from dunbrack.fccc.edu/betaturn and github.com/sh-maxim/BetaTurn18. Given the ubiquitous nature of beta turns, this comprehensive study updates understanding of beta turns and should also provide useful tools for protein structure determination, refinement, and prediction programs.
Maxim V. Shapovalov, Slobodan Vucetic, Roland L. Dunbrack Jr.
PLoS Comput. Biol.2
2018 Regular Expression Guided Entity Mention Mining from Noisy Web Data
abstract
Many important entity types in web documents, such as dates, times, email addresses, and course numbers, follow or closely resemble patterns that can be described by Regular Expressions (REs).Due to a vast diversity of web documents and ways in which they are being generated, even seemingly straightforward tasks such as identifying mentions of date in a document become very challenging.It is reasonable to claim that it is impossible to create a RE that is capable of identifying such entities from web documents with perfect precision and recall.Rather than abandoning REs as a go-to approach for entity detection, this paper explores ways to combine the expressive power of REs, ability of deep learning to learn from large data, and human-in-the loop approach into a new integrated framework for entity identification from web data.The framework starts by creating or collecting the existing REs for a particular type of an entity.Those REs are then used over a large document corpus to collect weak labels for the entity mentions and a neural network is trained to predict those RE-generated weak labels.Finally, a human expert is asked to label a small set of documents and the neural network is fine tuned on those documents.The experimental evaluation on several entity identification problems shows that the proposed framework achieves impressive accuracy, while requiring very modest human effort.
Shanshan Zhang 0004, Lihong He 0001, Slobodan Vucetic, Eduard C. Dragut
EMNLP3
2018 Enhancing Disaster Situational Awareness via Automated Summary Dissemination of Social Media Content
abstract
The paper proposes a situational awareness service, named StayTuned that collects information from social media, extracts relevant messages, and broadcasts them to the subscribers through wireless emergency alert system. StayTuned uses automated filtering and summarization of messages and updates subscribers with real-time situational summaries. Extensive experiments were conducted using twitter data collected during the Sandy hurricane to evaluate performance of the automated message extraction.
Shanshan Zhang 0004, Amitangshu Pal, Krishna Kant 0001, Slobodan Vucetic
GLOBECOM4
2018 Interpretable Representation Learning for Healthcare via Capturing Disease Progression through Time
abstract
Various deep learning models have recently been applied to predictive modeling of Electronic Health Records (EHR). In medical claims data, which is a particular type of EHR data, each patient is represented as a sequence of temporally ordered irregularly sampled visits to health providers, where each visit is recorded as an unordered set of medical codes specifying patient's diagnosis and treatment provided during the visit. Based on the observation that different patient conditions have different temporal progression patterns, in this paper we propose a novel interpretable deep learning model, called Timeline. The main novelty of Timeline is that it has a mechanism that learns time decay factors for every medical code. This allows the Timeline to learn that chronic conditions have a longer lasting impact on future visits than acute conditions. Timeline also has an attention mechanism that improves vector embeddings of visits. By analyzing the attention weights and disease progression functions of Timeline, it is possible to interpret the predictions and understand how risks of future visits change over time. We evaluated Timeline on two large-scale real world data sets. The specific task was to predict what is the primary diagnosis category for the next hospital visit given previous visits. Our results show that Timeline has higher accuracy than the state of the art deep learning models based on RNN. In addition, we demonstrate that time decay factors and attentions learned by Timeline are in accord with the medical knowledge and that Timeline can provide a useful insight into its predictions.
Tian Bai 0001, Shanshan Zhang 0004, Brian L. Egleston, Slobodan Vucetic
KDD4
2017 Joint learning of representations of medical concepts and words from EHR data
abstract
There has been an increasing interest in learning low-dimensional vector representations of medical concepts from electronic health records (EHRs). While EHRs contain structured data such as diagnostic codes and laboratory tests, they also contain unstructured clinical notes, which provide more nuanced details on a patient's health status. In this work, we propose a method that jointly learns medical concept and word representations. In particular, we focus on capturing the relationship between medical codes and words by using a novel learning scheme for word2vec model. Our method exploits relationships between different parts of EHRs in the same visit and embeds both codes and words in the same continuous vector space. In the end, we are able to derive clusters which reflect distinct disease and treatment patterns. In our experiments, we qualitatively show how our methods of grouping words for given diagnostic codes compares with a topic modeling approach. We also test how well our representations can be used to predict disease patterns of the next visit. The results show that our approach outperforms several common methods.
Tian Bai 0001, Ashis Kumar Chanda, Brian L. Egleston, Slobodan Vucetic
BIBM4
2016 Joint Learning of Representation and Structure for Sparse Regression on Graphs
abstract
In many applications, including climate science, power systems, and remote sensing, multiple input variables are observed for each output variable and the output variables are dependent. Several methods have been proposed to improve prediction by learning the conditional distribution of the output variables. However, when the relationship between the raw features and the outputs is nonlinear, the existing methods cannot capture both the nonlinearity and the underlying structure well. In this study, we propose a structured model containing hidden variables, which are nonlinear functions of inputs and which are linearly related with the output variables. The parameters modeling the relationships between the input and hidden variables, between the hidden and output variables, as well as among the output variables are learned simultaneously. To demonstrate the effectiveness of our proposed method, we conducted extensive experiments on eight synthetic datasets and three real-world challenging datasets: forecasting wind power, forecasting solar energy, and forecasting precipitation over U.S. The proposed method was more accurate than state-of-the-art structured regression methods.
Chao Han 0003, Shanshan Zhang 0004, Mohamed F. Ghalwash, Slobodan Vucetic, Zoran Obradovic
SDM4
2016 Semi-supervised combination of experts for aerosol optical depth estimation
Nemanja Djuric, Lakesh Kansakar, Slobodan Vucetic
Artif. Intell.3
2015 Gaussian Conditional Random Fields for Aggregation of Operational Aerosol Retrievals
abstract
We present a Gaussian conditional random field model for the aggregation of aerosol optical depth (AOD) retrievals from multiple satellite instruments into a joint retrieval. The model provides aggregated retrievals with higher accuracy and coverage than any of the individual instruments while also providing an estimation of retrieval uncertainty. The proposed model finds an optimal temporally smoothed combination of individual retrievals that minimizes the root-mean-squared error of AOD retrieval. We evaluated the model on five years (2006-2010) of satellite data over North America from five instruments (Aqua and Terra MODIS, MISR, SeaWiFS, and the Ozone Monitoring Instrument), collocated with ground-based Aerosol Robotic Network ground-truth AOD readings, clearly showing that the aggregation of different sources leads to improvements in the accuracy and coverage of AOD retrievals.
Nemanja Djuric, Vladan Radosavljevic, Zoran Obradovic, Slobodan Vucetic
IEEE Geosci. Remote. Sens. Lett.4
2015 Scaling Up Graph-Based Semisupervised Learning via Prototype Vector Machines
abstract
When the amount of labeled data are limited, semisupervised learning can improve the learner's performance by also using the often easily available unlabeled data. In particular, a popular approach requires the learned function to be smooth on the underlying data manifold. By approximating this manifold as a weighted graph, such graph-based techniques can often achieve state-of-the-art performance. However, their high time and space complexities make them less attractive on large data sets. In this paper, we propose to scale up graph-based semisupervised learning using a set of sparse prototypes derived from the data. These prototypes serve as a small set of data representatives, which can be used to approximate the graph-based regularizer and to control model complexity. Consequently, both training and testing become much more efficient. Moreover, when the Gaussian kernel is used to define the graph affinity, a simple and principled method to select the prototypes can be obtained. Experiments on a number of real-world data sets demonstrate encouraging performance and scaling properties of the proposed approach. It also compares favorably with models learned via l1 -regularization at the same level of model sparsity. These results demonstrate the efficacy of the proposed approach in producing highly parsimonious and accurate models for semisupervised learning.
Kai Zhang 0001, Liang Lan, James T. Kwok, Slobodan Vucetic, Bahram Parvin
IEEE Trans. Neural Networks Learn. Syst.4
2014 Non-Linear Label Ranking for Large-Scale Prediction of Long-Term User Interests
abstract
We consider the problem of personalization of online services from the viewpoint of ad targeting, where we seek to find the best ad categories to be shown to each user, resulting in improved user experience and increased advertiser's revenue. We propose to address this problem as a task of ranking the ad categories depending on a user's preference, and introduce a novel label ranking approach capable of efficiently learning non-linear, highly accurate models in large-scale settings. Experiments on real-world advertising data set with more than 3.2 million users show that the proposed algorithm outperforms the existing solutions in terms of both rank loss and top-K retrieval performance, strongly suggesting the benefit of using the proposed model on large-scale ranking problems.
Nemanja Djuric, Mihajlo Grbovic, Vladan Radosavljevic, Narayan L. Bhamidipati, Slobodan Vucetic
AAAI5
2014 Spatial Scan for Disease Mapping on a Mobile Population
abstract
In disease mapping, the spatial scan statistic is used to detect spatial regions where population is exposed to a significantly higher disease risk than expected. In this important application, the current residence is typically used to define the location of individuals from the population. Considering the mobility of humans at various temporal and spatial scales, using only information about the current residence may be an insufficiently informative proxy because it ignores a multitude of exposures that may occur away from home, or which had occurred at previous residences. In this paper, we propose a spatial scan statistic that is appropriate for disease mapping on mobile populations. We formulate a computationally efficient algorithm that uses the proposed statistic to find significant high-risk regions from mobile population's disease status data. The algorithm is applicable on large populations and over dense spatial grids. The experimental results demonstrate that the proposed algorithm is computationally efficient and outperforms the traditional disease clustering approaches at discovering high-risk regions in mobile populations.
Liang Lan, Vuk Malbasa, Slobodan Vucetic
AAAI3
2014 Neural Gaussian Conditional Random Fields
Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
ECML/PKDD (2)2
2014 Frugal Traffic Monitoring with Autonomous Participatory Sensing
abstract
As mobile devices are becoming pervasive, participatory sensing is becoming an attractive way of collecting large quantities of valuable location-based data. An important participatory sensing application is traffic monitoring, where GPS-enabled smartphones can provide invaluable information about traffic conditions. In this paper we propose a strategy for frugal sensing in which the participants send only a fraction of the observed traffic information to reduce costs while achieving high accuracy. The strategy is based on autonomous sensing, in which participants make decisions to send traffic information without guidance from the central server, thus reducing the communication overhead and improving privacy. We propose to use traffic flow theory in deciding whether or not to send an observation to the server. To provide accurate and computationally efficient estimation of the current traffic, we propose to use a budgeted version of the Gaussian Process model on the server side. The model is tuned to be robust to missing observations, which occur quite often in frugal sensing. The experiments on real-life traffic data set indicate that the proposed approach can use up to two orders of magnitude less samples than a baseline approach when estimating traffic speed on a highway network, with only a negligible loss in accuracy.
Vladimir Coric, Nemanja Djuric, Slobodan Vucetic
SDM3
2013 Continuous Conditional Random Fields for Efficient Regression in Large Fully Connected Graphs
abstract
When used for structured regression, powerful Conditional Random Fields (CRFs) are typically restricted to modeling effects of interactions among examples in local neighborhoods. Using more expressive representation would result in dense graphs, making these methods impractical for large-scale applications. To address this issue, we propose an effective CRF model with linear scale-up properties regarding approximate learning and inference for structured regression on large, fully connected graphs. The proposed method is validated on real-world large-scale problems of image de-noising and remote sensing. In conducted experiments, we demonstrated that dense connectivity provides an improvement in prediction accuracy. Inference time of less than ten seconds on graphs with millions of nodes and trillions of edges makes the proposed model an attractive tool for large-scale, structured regression problems.
Kosta Ristovski, Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
AAAI3
2013 Distributed confidence-weighted classification on MapReduce
abstract
Explosive growth in data size, data complexity, and data rates, triggered by emergence of high-throughput technologies such as remote sensing, crowd-sourcing, social networks, or computational advertising, in recent years has led to an increasing availability of data sets of unprecedented scales, with billions of high-dimensional data examples stored on hundreds of terabytes of memory. In order to make use of this large-scale data and extract useful knowledge, researchers in machine learning and data mining communities are faced with numerous challenges, since the classification algorithms designed for standard desktop computers are not capable of addressing these problems due to memory and time constraints. As a result, there exists an evident need for development of novel, more scalable algorithms that can handle large data sets. In this paper we propose such method, named AROW-MR, a linear SVM solver for efficient training of recently proposed confidence-weighted (CW) classifiers. Linear CW models maintain a Gaussian distribution over parameter vectors, thus allowing a user to estimate, in addition to separating hyperplane between two classes, parameter confidence as well. The proposed method employs MapReduce framework to train CW classifier in a distributed way, obtaining significant improvements in both training time and accuracy. This is achieved through training of local CW classifiers on each mapper, followed by optimally combining local classifiers on the reducer to obtain aggregated, more accurate CW linear model. We validated the proposed algorithm on synthetic data, and further showed that AROW-MR algorithm outperforms the baseline classifiers on an industrial, large-scale task of Ad Latency prediction, with nearly one billion examples.
Nemanja Djuric, Mihajlo Grbovic, Slobodan Vucetic
IEEE BigData3
2013 Efficient Visualization of Large-Scale Data Tables through Reordering and Entropy Minimization
abstract
Visualization of data tables with n examples and m columns using heat maps provides a holistic view of the original data. As there are n! ways to order rows and m! ways to order columns, and data tables are typically ordered without regard to visual inspection, heat maps of the original data tables often appear as noisy images. However, if rows and columns of a data table are ordered such that similar rows and similar columns are grouped together, a heat map may provide a deep insight into the underlying data distribution. We propose an information-theoretic approach to produce a well-ordered data table. In particular, we search for ordering that minimizes entropy of residuals of predictive coding applied on the ordered data table. This formalization leads to a novel ordering procedure, EM-ordering, that can be applied separately on rows and columns. For ordering of rows, EM-ordering repeats until convergence the steps of (1) rescaling columns and (2) solving a Traveling Salesman Problem (TSP) where rows are treated as cities. To allow fast ordering of large data tables, we propose an efficient TSP heuristic with modest O(n log(n)) time complexity. When compared to the existing state-of-the-art reordering approaches, we show that the method often provides heat maps of higher visual quality, while being significantly more scalable. Moreover, analysis of real-world traffic and financial data sets using the proposed method, which allowed us to readily gain deeper insights about the data, further confirmed that EM-ordering can be a valuable tool for visual exploration of large-scale data sets.
Nemanja Djuric, Slobodan Vucetic
ICDM2
2013 Semi-Supervised Learning for Integration of Aerosol Predictions from Multiple Satellite Instruments
Nemanja Djuric, Lakesh Kansakar, Slobodan Vucetic
IJCAI3
2013 Multi-Prototype Label Ranking with Novel Pairwise-to-Total-Rank Aggregation
Mihajlo Grbovic, Nemanja Djuric, Slobodan Vucetic
IJCAI3
2013 MS-kNN: protein function prediction by integrating multiple data sources
abstract
BACKGROUND: Protein function determination is a key challenge in the post-genomic era. Experimental determination of protein functions is accurate, but time-consuming and resource-intensive. A cost-effective alternative is to use the known information about sequence, structure, and functional properties of genes and proteins to predict functions using statistical methods. In this paper, we describe the Multi-Source k-Nearest Neighbor (MS-kNN) algorithm for function prediction, which finds k-nearest neighbors of a query protein based on different types of similarity measures and predicts its function by weighted averaging of its neighbors' functions. Specifically, we used 3 data sources to calculate the similarity scores: sequence similarity, protein-protein interactions, and gene expressions. RESULTS: We report the results in the context of 2011 Critical Assessment of Function Annotation (CAFA). Prior to CAFA submission deadline, we evaluated our algorithm on 1,302 human test proteins that were represented in all 3 data sources. Using only the sequence similarity information, MS-kNN had term-based Area Under the Curve (AUC) accuracy of Gene Ontology (GO) molecular function predictions of 0.728 when 7,412 human training proteins were used, and 0.819 when 35,622 training proteins from multiple eukaryotic and prokaryotic organisms were used. By aggregating predictions from all three sources, the AUC was further improved to 0.848. Similar result was observed on prediction of GO biological processes. Testing on 595 proteins that were annotated after the CAFA submission deadline showed that overall MS-kNN accuracy was higher than that of baseline algorithms Gotcha and BLAST, which were based solely on sequence similarity information. Since only 10 of the 595 proteins were represented by all 3 data sources, and 66 by two data sources, the difference between 3-source and one-source MS-kNN was rather small. CONCLUSIONS: Based on our results, we have several useful insights: (1) the k-nearest neighbor algorithm is an efficient and effective model for protein function prediction; (2) it is beneficial to transfer functions across a wide range of organisms; (3) it is helpful to integrate multiple sources of protein information.
Liang Lan, Nemanja Djuric, Yuhong Guo, Slobodan Vucetic
BMC Bioinform.4
2013 BudgetedSVM: a toolbox for scalable SVM approximations
Nemanja Djuric, Liang Lan, Slobodan Vucetic
J. Mach. Learn. Res.3
2013 Supervised clustering of label ranking data using label preference information
Mihajlo Grbovic, Nemanja Djuric, Shengbo Guo, Slobodan Vucetic
Mach. Learn.4
2013 Decentralized Estimation using distortion sensitive learning vector quantization
Mihajlo Grbovic, Slobodan Vucetic
Pattern Recognit. Lett.2
2013 Cold Start Approach for Data-Driven Fault Detection
abstract
A typical assumption in supervised fault detection is that abundant historical data are available prior to model learning, where all types of faults have already been observed at least once. This assumption is likely to be violated in practical settings as new fault types can emerge over time. In this paper we study this often overlooked cold start learning problem in data-driven fault detection, where in the beginning only normal operation data are available and faulty operation data become available as the faults occur. We explored how to leverage strengths of unsupervised and supervised approaches to build a model capable of detecting faults even if none are still observed, and of improving over time, as new fault types are observed. The proposed framework was evaluated on the benchmark Tennessee Eastman Process data. The proposed fusion model performed better on both unseen and seen faults than the stand-alone unsupervised and supervised models.
Mihajlo Grbovic, Weichang Li, Niranjan A. Subrahmanya, Adam K. Usadi, Slobodan Vucetic
IEEE Trans. Ind. Informatics5
2012 Convex Kernelized Sorting
abstract
Kernelized sorting is a method for aligning objects across two domains by considering within-domain similarity, without a need to specify a cross-domain similarity measure. In this paper we present the Convex Kernelized Sorting method where, unlike in the previous approaches, the cross-domain object matching is formulated as a convex optimization problem, leading to simpler optimization and global optimum solution. Our method outputs soft alignments between objects, which can be used to rank the best matches for each object, or to visualize the object matching and verify the correct choice of the kernel. It also allows for computing hard one-to-one alignments by solving the resulting Linear Assignment Problem. Experiments on a number of cross-domain matching tasks show the strength of the proposed method, which consistently achieves higher accuracy than the existing methods.
Nemanja Djuric, Mihajlo Grbovic, Slobodan Vucetic
AAAI3
2012 Sparse Principal Component Analysis with Constraints
abstract
The sparse principal component analysis is a variant of the classical principal component analysis, which finds linear combinations of a small number of features that maximize variance across data. In this paper we propose a methodology for adding two general types of feature grouping constraints into the original sparse PCA optimization procedure.We derive convex relaxations of the considered constraints, ensuring the convexity of the resulting optimization problem. Empirical evaluation on three real-world problems, one in process monitoring sensor networks and two in social networks, serves to illustrate the usefulness of the proposed methodology.
Mihajlo Grbovic, Christopher R. Dance, Slobodan Vucetic
AAAI3
2012 A connectivity-based popularity prediction approach for social networks
abstract
In social media websites, such as Twitter and Digg, certain content will attract much more visitors than others. Predicting which content will become popular is of interest to website owners and market analysts. In this paper, we present a novel technique to predict popularity using the connection features of individuals and their community. Our approach is based on the hypothesis that connection plays a dominant role in spreading content on social media. The resulting predictor is more efficient than approaches which estimate popularity by complex graph properties, and more accurate than approaches that use simple visit counts. We evaluated the proposed approach empirically on several real-life data sets. Results indicate that, compared with the conventional methods, our approach is both accurate and computationally efficient.
Huangmao Quan, Ana Milicic, Slobodan Vucetic, Jie Wu 0001
ICC3
2012 Supervised Clustering of Label Ranking Data
abstract
In this paper we study supervised clustering in the context of label ranking data. Segmentation of such complex data has many potential real-world applications. For example, in target marketing, the goal is to cluster customers in the feature space by taking into consideration the assigned, potentially incomplete product preferences, such that the preferences of instances within a cluster are more similar than the preferences of customers in the other clusters. We establish several heuristic baselines for this application that make use of well-known algorithms such as K-means, and propose a principled algorithm specifically tailored for this type of clustering. It is based on the Plackett-Luce (PL) probabilistic ranking model. Each cluster is represented as a union of Voronoi cells defined by a set of prototypes and is assigned a set of PL label scores that determine the cluster-specific label ranking. The unknown cluster PL parameters and prototype positions are determined using a supervised learning technique. Cluster membership and ranking for a new instance is determined by membership of its nearest prototype. The proposed algorithms were empirically evaluated on synthetic and reallife label ranking data. The PL-based method was superior to the heuristically-based supervised clustering approaches. The proposed PL-based algorithm was also evaluated on the task of label ranking prediction. The results showed that it is highly competitive to the state of the art label ranking algorithms, and that it is particularly accurate on data with partial rankings.
Mihajlo Grbovic, Nemanja Djuric, Slobodan Vucetic
SDM3
2012 Breaking the curse of kernelization: budgeted stochastic gradient descent for large-scale SVM training
Koby Crammer, Slobodan Vucetic
J. Mach. Learn. Res.3
2012 Uncertainty Analysis of Neural-Network-Based Aerosol Retrieval
abstract
Neural networks have the ability to represent and learn complex regression functions and are very suitable for retrieval of geophysical parameters from remotely sensed data. Neural networks trained to minimize the mean square error are able to estimate the conditional expectation of target variables. In many remote sensing applications, it is also critical to provide estimates of prediction uncertainty. In this paper, we evaluate an approach that, in addition to training a neural network for retrievals, also trains a neural-network-based estimator of retrieval uncertainty. The uncertainty estimator is built under the assumption that uncertainty is a function of input variables. The methodology was evaluated on aerosol-optical-depth retrieval. The data set consists of 38 238 collocated Moderate Resolution Imaging Spectrometer (MODIS) satellite instrument and Aerosol Robotic Network ground-based instrument measurements collected over the entire Earth during two years (in 2005-2006). The results indicate that a neural network ensemble is more accurate than the operational MODIS retrieval algorithm called Collection 5 and that the retrieval uncertainty of the ensemble can be estimated with satisfactory accuracy.
Kosta Ristovski, Slobodan Vucetic, Zoran Obradovic
IEEE Trans. Geosci. Remote. Sens.2
2012 Mixture Model for Multiple Instance Regression and Applications in Remote Sensing
abstract
The multiple instance regression (MIR) problem arises when a data set is a collection of bags, where each bag contains multiple instances sharing the identical real-valued label. The goal is to train a regression model that can accurately predict label of an unlabeled bag. Many remote sensing applications can be studied within this setting. We propose a novel probabilistic framework for MIR that represents bag labels with a mixture model. It is based on an assumption that each bag contains a prime instance which is responsible for the bag label. An expectation-maximization algorithm is proposed to maximize the likelihood of the mixture model. The mixture model MIR framework is quite flexible, and several existing MIR algorithms can be described as its special cases. The proposed algorithms were evaluated on synthetic data and remote sensing data for aerosol retrieval and crop yield prediction. The results show that the proposed MIR algorithms achieve higher accuracy than the previous state of the art.
Liang Lan, Slobodan Vucetic
IEEE Trans. Geosci. Remote. Sens.3
2011 Spatially regularized logistic regression for disease mapping on large moving populations
abstract
Spatial analysis of disease risk, or disease mapping, typically relies on information about the residence and health status of individuals from population under study. However, residence information has its limitations because people are exposed to numerous disease risks as they spend time outside of their residences. Thanks to the wide-spread use of mobile phones and GPS-enabled devices, it is becoming possible to obtain a detailed record about the movement of human populations. Availability of movement information opens up an opportunity to improve the accuracy of disease mapping. Starting with an assumption that an individual's disease risk is a weighted average of risks at the locations which were visited, we show that disease mapping can be accomplished by spatially regularized logistic regression. Due to the inherent sparsity of movement data, the proposed approach can be applied to large populations and over large spatial grids. In our experiments, we were able to map disease for a simulated population with 1.6 million people and a spatial grid with 65 thousand locations in several minutes. The results indicate that movement information can improve the accuracy of disease mapping as compared to residential data only. We also studied a privacy-preserving scenario in which only the aggregate statistics are available about the movement of the overall population, while detailed movement information is available only for individuals with disease. The results indicate that the accuracy of disease mapping remains satisfactory when learning from movement data sanitized in this way.
Vuk Malbasa, Slobodan Vucetic
KDD2
2011 Trading representability for scalability: adaptive multi-hyperplane machine for nonlinear classification
abstract
Support Vector Machines (SVMs) are among the most popular and successful classification algorithms. Kernel SVMs often reach state-of-the-art accuracies, but suffer from the curse of kernelization due to linear model growth with data size on noisy data. Linear SVMs have the ability to efficiently learn from truly large data, but they are applicable to a limited number of domains due to low representational power. To fill the representability and scalability gap between linear and nonlinear SVMs, we propose the Adaptive Multi-hyperplane Machine (AMM) algorithm that accomplishes fast training and prediction and has capability to solve nonlinear classification problems. AMM model consists of a set of hyperplanes (weights), each assigned to one of the multiple classes, and predicts based on the associated class of the weight that provides the largest prediction. The number of weights is automatically determined through an iterative algorithm based on the stochastic gradient descent algorithm which is guaranteed to converge to a local optimum. Since the generalization bound decreases with the number of weights, a weight pruning mechanism is proposed and analyzed. The experiments on several large data sets show that AMM is nearly as fast during training and prediction as the state-of-the-art linear SVM solver and that it can be orders of magnitude faster than kernel SVM. In accuracy, AMM is somewhere between linear and kernel SVMs. For example, on an OCR task with 8 million highly dimensional training examples, AMM trained in 300 seconds on a single-core processor had 0.54% error rate, which was significantly lower than 2.03% error rate of a linear SVM trained in the same time and comparable to 0.43% error rate of a kernel SVM trained in 2 days on 512 processors. The results indicate that AMM could be an attractive option when solving large-scale classification problems. The software is available at www.dabi.temple.edu/~vucetic/AMM.html.
Nemanja Djuric, Koby Crammer, Slobodan Vucetic
KDD4
2011 Tracking Concept Change with Incremental Boosting by Minimization of the Evolving Exponential Loss
Mihajlo Grbovic, Slobodan Vucetic
ECML/PKDD (1)2
2010 Continuous Conditional Random Fields for Regression in Remote Sensing
abstract
Conditional random fields (CRF) are widely used for predicting output variables that have some internal structure. Most of the CRF research has been done on structured classification where the outputs are discrete. In this study we propose a CRF probabilistic model for structured regression that uses multiple non-structured predictors as its features. We construct features as squared prediction errors and show that this results in a Gaussian predictor. Learning becomes a convex optimization problem leading to a global solution for a set of parameters. Inference can be conveniently conducted through matrix computation. Experimental results on the remote sensing problem of estimating Aerosol Optical Depth (AOD) provide strong evidence that the proposed CRF model successfully exploits the inherent spatio-temporal properties of AOD data. The experiments revealed that CRF are more accurate than the baseline neural network and domain-based predictors.
Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
ECAI2
2010 Multi-Class Pegasos on a Budget
Koby Crammer, Slobodan Vucetic
ICML3
2010 A Data-Mining Technique for Aerosol Retrieval Across Multiple Accuracy Measures
abstract
A typical approach in supervised learning is to select an accuracy measure and train a predictor that maximizes it. This can be insufficient in remote-sensing applications where predictor performance is often evaluated over multiple domain-specific accuracy measures. Here, we test the hypothesis that predictors can be trained to maximize performance over multiple accuracy measures. To do this, we evaluate several metalearning algorithms on the problem of aerosol optical depth (AOD) retrieval. The multiple accuracy measures included mean squared error, correlation, relative squared error, and fraction of satisfactory predictions. The proposed metalearning algorithms have a two-layer architecture, where the first layer consists of multiple neural networks, each trained using a different accuracy measure, and the second layer aggregates decisions of the first layer predictors. To evaluate AOD predictors, we used nearly 70 000 collocated data points whose attributes were radiances, solar and view angles, and terrain elevation from MODerate resolution Imaging Spectrometer (MODIS) instrument satellite observations and whose target AOD variable was obtained from the ground-based AEROsol robotic NETwork (AERONET) instruments. The data were collected at 221 AERONET locations over the globe in the period between 2005 and 2007. AOD prediction accuracies of neural networks were compared to the recently developed operational MODISC005retrieval algorithm and to several other data-mining methods. Results showed that neural networks are better at reproducing the test data than the operational retrieval algorithm and that predictors obtained by metalearning are robust over multiple accuracy measures.
Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
IEEE Geosci. Remote. Sens. Lett.2
2009 A Multi-task Feature Selection Filter for Microarray Classification
abstract
A major challenge in microarray classification and biomarker discovery is dealing with small-sample high-dimensional data where the number of genes used as features is typically orders of magnitude larger than the number of labeled microarrays. One way to address this challenge is by leveraging information from the publicly accessible repositories of microarray data. Following this idea, a multi-task feature selection filter is proposed that borrows strength from the auxiliary microarray classification data sets. The filter uses Kruskal-Wallis test on auxiliary data sets and ranks genes based on their aggregated p-values. Expressions of the top-ranked genes are used as features to build a classifier on the target data set. The proposed approach was evaluated on 9 microarray data sets related to 9 different types of cancers. Comparison of the classification accuracies reveals that the multi-task feature selection is superior to single-task feature selection. Furthermore, the results strongly suggest that multi-task algorithms could improve microarray classification by exploiting auxiliary data during feature selection and learning.
Liang Lan, Slobodan Vucetic
BIBM2
2009 Decentralized Estimation Using Learning Vector Quantization
abstract
A decentralized estimation system consists of n distributed data sources S1... Snand a fusion center. The data sources produce multivariate random vectors X1... Xnthat are transmitted to the fusion center in the form of messages Z1... Zn, Zi= alpha1(X1). Due to communication constraints, Ziis a discrete variable with cardinality Mirepresented as an integer from a set {1...Mi}. At the fusion center, the goal is to estimate the conditional expectation of unobserved variable Y, E(Y|x1... xn), by fusion function h(z1... zn). The challenge is to find quantization functions alpha1... alphanand fusion function h such that the estimation error is minimized under given communication constraints.
Mihajlo Grbovic, Slobodan Vucetic
DCC2
2009 Compressed Kernel Perceptrons
abstract
Kernel machines are a popular class of machine learning algorithms that achieve state of the art accuracies on many real-life classification problems. Kernel perceptrons are among the most popular online kernel machines that are known to achieve high-quality classification despite their simplicity. They are represented by a set of B prototype examples, called support vectors, and their associated weights. To obtain a classification, a new example is compared to the support vectors. Both space to store a prediction model and time to provide a single classification scale as O(B). A problem with kernel perceptrons is that on noisy data the number of support vectors tends to grow without bounds with the number of training examples. To reduce the strain at computational resources, budget kernel perceptrons have been developed by upper bounding the number of support vectors. In this work, we propose a new budget algorithm that upper bounds the number of bits needed to store kernel perceptron. Setting the bitlength constraint could facilitate development of hardware and software implementations of kernel perceptrons on resource-limited devices such as microcontrollers. The proposed compressed kernel perceptron algorithm decides on the optimal tradeoff between number of support vectors and their bit precision. The algorithm was evaluated on several benchmark data sets and the results indicate that it can train highly accurate classifiers even when the available memory budget is below 1Kbit. This promising result points to a possibility of implementing powerful learning algorithms even on the most resource-constrained computational devices.
Slobodan Vucetic, Vladimir Coric
DCC1
2009 Active Selection of Sensor Sites in Remote Sensing Applications
abstract
In a data-mining approach, a model for estimation of aerosol optical depth (AOD) from satellite observations is learned using collocated satellite and ground-based observations. For accurate learning of such a spatio-temporal model, it is important to collect ground-based data from a large number of sites. The objective of this project is to determine appropriate locations for the next set of ground-based data collection sites to maximize accuracy of AOD estimation. Ideally, a new site should capture the most significant unseen aerosol patterns and should be the least correlated with the previously observed patterns. We propose achieving this aim by selecting the locations on which the existing prediction model is the most uncertain. Several criteria were considered for site selection, including uncertainty, spatial diversity, similarity in temporal pattern, and their combination. Extensive experiments on globally distributed data over 90 AERONET sites from the years 2005 and 2006 provide strong evidence that sites selected using the proposed algorithms improve the overall AOD prediction accuracy at a faster rate than those selected randomly or based on spatial diversity among sites.
Debasish Das, Zoran Obradovic, Slobodan Vucetic
ICDM3
2009 Regression Learning Vector Quantization
abstract
Learning vector quantization (LVQ) is a popular class of nearest prototype classifiers for multiclass classification. Learning algorithms from this family are widely used because of their intuitively clear learning process and ease of implementation. In this paper we propose an extension of the LVQ algorithm to regression. Just like the LVQ algorithm, the proposed modification uses a supervised learning procedure to learn the best prototype positions, but unlike LVQ algorithm for classification, it also learns the best prototype target values. This results in the effective partition of the feature space, similar to the one the K-means algorithm would make. Experimental results on benchmark datasets showed that the proposed regression LVQ algorithm performs better than the nearest prototype competitors that choose prototypes randomly or through K-means clustering, classification LVQ on quantized target values, and similarly to the memory-based Parzen window and nearest neighbor algorithms.
Mihajlo Grbovic, Slobodan Vucetic
ICDM2
2009 Fast Online Training of Ramp Loss Support Vector Machines
abstract
A fast online algorithm OnlineSVMRfor training Ramp-Loss Support Vector Machines (SVMRs) is proposed. It finds the optimal SVMRfor t + 1 training examples using SVMR built on t previous examples. The algorithm retains the Karush-Kuhn-Tucker conditions on all previously observed examples. This is achieved by an SMO-style incremental learning and decremental unlearning under the Concave-Convex Procedure framework. Further speedup of training time could be achieved by dropping the requirement of optimality. A variant, called OnlineASVMR, is a greedy approach that approximately optimizes the SVMRobjective function and is suitable for online active learning. The proposed algorithms were comprehensively evaluated on 9 large benchmark data sets. The results demonstrate that OnlineSVMR(1) has the similar computational cost as its offline counterpart; (2) outperforms IDSVM, its competing online algorithm that uses hinge-loss, in terms of accuracy, model sparsity and training time. The experiments on online active learning show that for a fixed number of label queries OnlineASVMR(1) achieves consistently better accuracy than QueryAll and competitive accuracy to Greedy approach; (2) outperforms the active learning version of IDSVM.
Slobodan Vucetic
ICDM2
2009 Learning Vector Quantization with adaptive prototype addition and removal
abstract
Learning Vector Quantization (LVQ) is a popular class of nearest prototype classifiers for multiclass classification. Learning algorithms from this family are widely used because of their intuitively clear learning process and ease of implementation. They run efficiently and in many cases provide state of the art performance. In this paper we propose a modification of the LVQ algorithm that addresses problems of determining appropriate number of prototypes, sensitivity to initialization, and sensitivity to noise in data. The proposed algorithm allows adaptive addition of prototypes at potentially beneficial locations and removal of harmful or less useful prototypes. The prototype addition and removal steps can be easily implemented on top of many existing LVQ algorithms. Experimental results on synthetic and benchmark datasets showed that the proposed modifications can significantly improve LVQ classification accuracy while at the same time determining the appropriate number of prototypes and avoiding the problems of initialization.
Mihajlo Grbovic, Slobodan Vucetic
IJCNN2
2009 Tighter Perceptron with improved dual use of cached data for model representation and validation
abstract
Kernel perceptrons are represented by a subset of training points, called the support vectors, and their associated weights. To address the issue of unlimited growth in model size during training, budget kernel perceptrons maintain the fixed number of support vectors and thus achieve the constant update time and space complexity. In this paper, a new kernel perceptron algorithm for online learning on a budget is proposed. Following the idea of tighter perceptron, upon exceeding the budget, the algorithm removes the support vector with the minimal impact on classification accuracy. To optimize memory use, instead on maintaining a separate validation data set for accuracy estimation, the proposed algorithm only uses the support vectors for both model representation and validation. This is achieved by estimating posterior class probability of each support vector and using this information in validation. The experimental results on 11 benchmark data sets indicate that the proposed algorithm is significantly more accurate than the competing budget kernel perceptrons and that it has comparable accuracy to the resource unbounded perceptrons, including the original kernel perceptron and the tighter perceptron that uses whole training data set for validation.
Slobodan Vucetic
IJCNN2
2009 Twin Vector Machines for Online Learning on a Budget
abstract
This paper proposes Twin Vector Machine (TVM), a constant space and sublinear time Support Vector Machine (SVM) algorithm for online learning. TVM achieves its favorable scaling by maintaining only a fixed number of examples, called the twin vectors, and their associated information in memory during training. In addition, TVM guarantees that Kuhn-Tucker conditions are satisfied on all twin vectors at any time. To maximize the accuracy of TVM, twin vectors are adjusted during the training phase to approximate the data distribution near the decision boundary. Given a new training example, TVM is updated in three steps. First, the new example is added as a new twin vector if it is near the decision boundary. If this happens, two twin vectors are selected and merged into a single twin vector to maintain the budget. Finally, TVM is updated by incremental and decremental learning to account for the change. Several methods for twin vector merging were proposed and experimentally evaluated. TVMs were thoroughly tested on 12 large data sets. In most cases, the accuracy of low-budget TVMs was comparable to the state of the art resource-unconstrained SVMs. Additionally, the TVM accuracy was substantially larger than that of SVM trained on a random sample of the same size. Even larger difference in accuracy was observed when comparing to Forgetron, a popular kernel perceptron algorithm on a budget. The results illustrate that highly accurate online SVMs could be trained from large data streams using devices with severely limited memory budgets.
Slobodan Vucetic
SDM2
2008 Reducing Need for Collocated Ground and Satellite based Observations in Statistical Aerosol Optical Depth Estimation
abstract
One of the biggest challenges of current climate research is to characterize and quantify the effect of aerosols on the global and local weather. This requires an accurate prediction of aerosol optical density (AOD) which is defined as the amount of loss a beam of light incurs when it passes through the atmosphere. In this paper a neural network-based data-driven prediction model is considered which uses collocated satellite (MODIS) observation and ground-based (AERONET) AOD retrievals as predictors and target respectively. This paper studies an active learning-based data collection method which will facilitate the learning of a sufficiently accurate AOD prediction model using a minimal set of labeled training data by querying the labels of only the most informative data points.
Debasish Das, Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
IGARSS (2)3
2008 Spatio-Temporal Partitioning for Improving Aerosol Prediction Accuracy
abstract
In supervised learning, on data collected over space and time, different relationships can be found over different spatio-temporal regions. In such situations, an appropriate spatio-temporal data partitioning followed by building specialized predictors could often achieve higher overall prediction accuracy than when learning a single predictor on all the data. In practice, partitions are typically decided based on prior knowledge. As an alternative to domain-based partitioning, we propose a method that automatically discovers a spatio-temporal partitioning through the competition of regression models. The method is evaluated on a challenging problem using satellite observations to predict Aerosol Optical Depth (AOD), which represents the amount of depletion that a beam of radiation undergoes as it passes through the atmosphere. Our experiments used more than 20,000 labeled data points collected during 3 years from more than 100 sites worldwide. Our partitioning-based approach was compared to the recently developed operational AOD prediction algorithm, called C5, which uses domain knowledge for spatio-temporal partitioning of the Earth and implements a region-specific deterministic predictor that utilizes forward simulations from the postulated physical models. Data partitioning used in C5 divides the world into three spatio-temporal regions that differ based on the location and the time of the year as decided by domain experts. The results showed that a neural network predictor trained on all the data has accuracy comparable to C5. When specialized neural network predictors were learned on C5-based partitions, the overall prediction accuracy was not improved. On the other hand, our competition-based spatio-temporal data partitioning approach resulted in large accuracy improvements. The most accurate results were obtained when (1) the data from each of the sites were split into two temporal subsets, one for winter-spring months and another for summer-fall months; and (2) two neural network predictors were competing for each of the identified spatio-temporal subsets.
Vladan Radosavljevic, Slobodan Vucetic, Zoran Obradovic
SDM2
2008 A Data-Mining Approach for the Validation of Aerosol Retrievals
abstract
Operational algorithms for retrieval of aerosols from satellite observations are typically created manually based on the domain knowledge. Validation studies, where the retrievals are compared to the available ground-truth data, are periodically performed with the goal of understanding how to further improve the quality of the retrieval algorithms. This letter describes a data-mining approach aimed to facilitate this highly labor-intensive process. It is based on training a neural network for retrieval and comparing its performance with that of the operational algorithm. The situations, where a neural network is more accurate, point to the weaknesses of the operational algorithm that could be corrected. Use of decision trees is proposed to provide easily interpretable descriptions of such situations. The approach was applied on 3646 collocated Moderate Resolution Imaging Spectroradiometer and AERONET observations over the continental U.S. related to the retrieval of aerosol optical thickness. The experiments showed that the approach is feasible and that it can be a valuable tool for the domain scientists working on the development of retrieval algorithms.
Slobodan Vucetic, Wen Mi, Zhanquing Li, Zoran Obradovic
IEEE Geosci. Remote. Sens. Lett.1
2007 A Reservoir Sampling Algorithm with Adaptive Estimation of Conditional Expectation
abstract
Resource-constrained data mining introduces many constraints when learning from large datasets. It is often not practical or possible to keep the entire data set in main memory and often the data could be observed in a single run in the order in which they are presented. Traditional reservoir-based approaches perform well in this situation. One drawback of these approaches is that the examples not included in the final reservoir are often ignored. To remedy this situation we propose a modification to the baseline reservoir algorithm. Instead of keeping the actual target values of reservoir examples, an estimate of their conditional expectation is kept and updated online as new data are observed from the stream. The estimate is obtained by averaging target values of the similar examples. The proposed algorithm uses a paired t-test to determine the similarity threshold. Thorough evaluation on generated two dimensional data shows that the proposed algorithm is producing reservoirs with considerably reduced target noise. This property allows training of significantly improved prediction models as compared with the baseline reservoir-based approach.
Vuk Malbasa, Slobodan Vucetic
IJCNN2
2006 A Fast Algorithm for Lossless Compression of Data Tables by Reordering
abstract
Summary form only given. An algorithm for lossless compression of tables with numeric attributes based on row ordering is proposed. Extensive experiments were performed on randomly generated and scientific multidimensional tables with numerical attributes. The results showed that ordering is useful for compression of moderately large to large tables with intrinsic dimensionality below 20 and with attributes represented with low to moderate precision. The benefits of the iterative ordering procedure are the largest on data tables with correlated attributes and heterogeneous attribute types
Slobodan Vucetic
DCC1
2006 Data Mining Support for the Improvement of MODIS Aerosol Retrievals
abstract
Abstract—This paper describes data mining approach for improving the accuracy of aerosol retrieval algorithms. The approach was applied on 1,722 collocated MODIS and AERONET observations over the western part of the continental U.S. Neural networks were trained to predict AERONET Aerosol Optical Thickness (AOT) using attributes derived from observations made by MODIS instrument onboard TERRA satellite. The results showed that neural networks provide more accurate retrievals than the operational MODIS algorithm. Study of differences between neural networks and the MODIS algorithm revealed useful information that can help domain scientists improve quality of the MODIS algorithm.
Zoran Obradovic, Zhanquing Li, Slobodan Vucetic
IGARSS4
2006 Substring selection for biomedical document classification
abstract
MOTIVATION: Attribute selection is a critical step in development of document classification systems. As a standard practice, words are stemmed and the most informative ones are used as attributes in classification. Owing to high complexity of biomedical terminology, general-purpose stemming algorithms are often conservative and could also remove informative stems. This can lead to accuracy reduction, especially when the number of labeled documents is small. To address this issue, we propose an algorithm that omits stemming and, instead, uses the most discriminative substrings as attributes. RESULTS: The approach was tested on five annotated sets of abstracts from iProLINK that report on the experimental evidence about five types of protein post-translational modifications. The experiments showed that Naive Bayes and support vector machine classifiers perform consistently better [with area under the ROC curve (AUC) accuracy in range 0.92-0.97] when using the proposed attribute selection than when using attributes obtained by the Porter stemmer algorithm (AUC in 0.86-0.93 range). The proposed approach is particularly useful when labeled datasets are small.
Zoran Obradovic, Zhang-Zhi Hu, Cathy H. Wu, Slobodan Vucetic
Bioinform.5
2006 Length-dependent prediction of protein intrinsic disorder
abstract
BACKGROUND: Due to the functional importance of intrinsically disordered proteins or protein regions, prediction of intrinsic protein disorder from amino acid sequence has become an area of active research as witnessed in the 6th experiment on Critical Assessment of Techniques for Protein Structure Prediction (CASP6). Since the initial work by Romero et al. (Identifying disordered regions in proteins from amino acid sequences, IEEE Int. Conf. Neural Netw., 1997), our group has developed several predictors optimized for long disordered regions (>30 residues) with prediction accuracy exceeding 85%. However, these predictors are less successful on short disordered regions (< or =30 residues). A probable cause is a length-dependent amino acid compositions and sequence properties of disordered regions. RESULTS: We proposed two new predictor models, VSL2-M1 and VSL2-M2, to address this length-dependency problem in prediction of intrinsic protein disorder. These two predictors are similar to the original VSL1 predictor used in the CASP6 experiment. In both models, two specialized predictors were first built and optimized for short (< or = 30 residues) and long disordered regions (>30 residues), respectively. A meta predictor was then trained to integrate the specialized predictors into the final predictor model. As the 10-fold cross-validation results showed, the VSL2 predictors achieved well-balanced prediction accuracies of 81% on both short and long disordered regions. Comparisons over the VSL2 training dataset via 10-fold cross-validation and a blind-test set of unrelated recent PDB chains indicated that VSL2 predictors were significantly more accurate than several existing predictors of intrinsic protein disorder. CONCLUSION: The VSL2 predictors are applicable to disordered regions of any length and can accurately identify the short disordered regions that are often misclassified by our previous disorder predictors. The success of the VSL2 predictors further confirmed the previously observed differences in amino acid compositions and sequence properties between short and long disordered regions, and justified our approaches for modelling short and long disordered regions separately. The VSL2 predictors are freely accessible for non-commercial use at http://www.ist.temple.edu/disprot/predictorVSL2.php.
Kang Peng, Predrag Radivojac, Slobodan Vucetic, A. Keith Dunker, Zoran Obradovic
BMC Bioinform.3
2006 A statistical complement to deterministic algorithms for the retrieval of aerosol optical thickness from radiance data
Slobodan Vucetic, Amy Braverman, Zoran Obradovic
Eng. Appl. Artif. Intell.2
2005 Accuracy-Optimized Quantization for High-Dimensional Data Fusion
abstract
Summary form only given. Decentralized estimation is an essential problem for a number of data fusion applications. The accuracy can be defined in terms of the mean square quantization error, MSQE. In this work, a computationally efficient and robust algorithm was developed for high-dimensional and high-rate decentralized estimation scenarios. Experiments were performed on a 2-source 21-dimensional problem. The proposed algorithm is compared with standard vector quantization (VQ), due to lack of alternative high-rate and high-dimensional decentralized estimation algorithms. The results showed that the proposed algorithm was consistently more accurate than standard VQ and that the difference increased with increase in number of codewords and size of the data set.
Slobodan Vucetic
DCC1
2005 Correcting Sampling Bias in Structural Genomics through Iterative Selection of Underrepresented Targets
abstract
In this study we proposed an iterative procedure for correcting sampling bias in labeled datasets for supervised learning applications. Given a much larger and unbiased unlabeled dataset, our approach relies on training contrast classifiers to iteratively select unlabeled examples most highly underrepresented in the labeled dataset. Once labeled, these examples could greatly reduce the sampling bias present in the labeled dataset. Unlike active learning methods, the actual labeling is not necessary in order to determine the most appropriate sampling schedule. The proposed procedure was applied on an important bioinformatics problem of prioritizing protein targets for structural genomics projects. We show that the procedure is capable of identifying protein targets that are underrepresented in current protein structure database, the Protein Data Bank (PDB). We argue that these proteins should be given higher priorities for experimental structural characterization to achieve faster sampling bias reduction in current PDB and make it more representative of the protein space.
Kang Peng, Slobodan Vucetic, Zoran Obradovic
SDM2
2005 DisProt: a database of protein disorder
abstract
UNLABELLED: The Database of Protein Disorder (DisProt) is a curated database that provides structure and function information about proteins that lack a fixed three-dimensional (3D) structure under putatively native conditions, either in their entirety or in part. Starting from the central premise that intrinsic disorder is an important structural class of protein and in order to meet the increasing interest thereof, DisProt is aimed at becoming a central repository of disorder-related information. For each disordered protein, the database includes the name of the protein, various aliases, accession codes, amino acid sequence, location of the disordered region(s), and methods used for structural (disorder) characterization. If applicable, most entries also list the biological function(s) of each disordered region, how each region of disorder is used for function, as well as provide links to PubMed abstracts and major protein databases. AVAILABILITY: www.disprot.org
Slobodan Vucetic, Zoran Obradovic, Vladimir Vacic, Predrag Radivojac, Kang Peng, Lilia M. Iakoucheva, Marc S. Cortese, J. David Lawson, Celeste J. Brown, Jason G. Sikes, Crystal D. Newton, A. Keith Dunker
Bioinform.1
2005 Collaborative Filtering Using a Regression-Based Approach
Slobodan Vucetic, Zoran Obradovic
Knowl. Inf. Syst.1
2004 Towards Efficient Learning of Neural Network Ensembles from Arbitrarily Large Datasets
Kang Peng, Zoran Obradovic, Slobodan Vucetic
ECAI3
2004 Feature Selection Filters Based on the Permutation Test
Predrag Radivojac, Zoran Obradovic, A. Keith Dunker, Slobodan Vucetic
ECML4
2003 Exploiting Unlabeled Data for Improving Accuracy of Predictive Data Mining
abstract
Predictive data mining typically relies on labeled data without exploiting a much larger amount of available unlabeled data. We show that using unlabeled data can be beneficial in a range of important prediction problems and therefore should be an integral part of the learning process. Given an unlabeled dataset representative of the underlying distribution and a K-class labeled sample that might be biased, our approach is to learn K contrast classifiers each trained to discriminate a certain class of labeled data from the unlabeled population. We illustrate that contrast classifiers can be useful in one-class classification, outlier detection, density estimation, and learning from biased data. The advantages of the proposed approach are demonstrated by an extensive evaluation on synthetic data followed by real-life bioinformatics applications for (1) ranking PubMed articles by their relevance to protein disorder and (2) cost-effective enlargement of a disordered protein database.
Kang Peng, Slobodan Vucetic, Hongbo M. Xie, Zoran Obradovic
ICDM2
2003 Detection of Underrepresented Biological Sequences using Class-Conditional Distribution Models
abstract
A labeled sequence data set related to a certain biological property is often biased and, therefore, does not completely capture its diversity in nature. To reduce this sampling bias problem a data mining procedure is proposed for detecting underrepresented relevant sequences. The procedure is aimed at helping domain experts achieve a cost-effective qualitative enlargement of knowledge through an in-depth study of a small number of statistically underrepresented and functionally interesting sequences. Our procedure consists of: (i) learning a class-conditional distribution model on each class of labeled data; (ii) applying the models to select statistically underrepresented unlabeled sequences; and (iii) automatically evaluating their interestingness. An application of the proposed approach is illustrated on an important problem of increasing the data set of confirmed disordered proteins. The obtained results demonstrate the promise of the proposed approach for an efficient reduction of sampling bias in biological databases.
Slobodan Vucetic, Dragoljub Pokrajac, Hongbo M. Xie, Zoran Obradovic
SDM1
2001 Classification on Data with Biased Class Distribution
Slobodan Vucetic, Zoran Obradovic
ECML1
2000 Discovering Homogeneous Regions in Spatial Data through Competition
Slobodan Vucetic, Zoran Obradovic
ICML1
2000 Performance Controlled Data Reduction for Knowledge Discovery in Distributed Databases
Slobodan Vucetic, Zoran Obradovic
PAKDD1
1999 A data partitioning scheme for spatial regression
abstract
Precision agriculture data consisting of crop yield and topographic features are examined with the objective of explaining yield variability as a function of topographic attributes in order to extrapolate this knowledge to unseen agricultural sites. It is demonstrated that random data partitioning into training, validation and test subsets is not appropriate when dealing with agricultural problems characterized with strong spatial data correlation. A simple spatial data partitioning scheme that leads to significantly faster neural network training and slightly better generalization is proposed. Also, integration of predictors formed from spatially partitioned data led to improved generalization over a bagging integration procedure in experiments. The margin between the best spatial model and a trivial predictor for our precision agriculture problem was small indicating that topographic features alone could explain only a small amount of the yield variability.
Slobodan Vucetic, Tim Fiez, Zoran Obradovic
IJCNN1