Russell Greiner

dblp:g/RussellGreiner · also Russ Greiner · DBLP profile ↗
← Back
132ranked-venue papers
27as first author
12since 2021 · last 2026
0000-0001-8327-934XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 100 · 25 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-authorSoftware engineering, systems software and programming languages · 5Theory of computation · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Temporally consistent survival prediction for non-uniform longitudinal data
abstract
OBJECTIVE: Traditional survival prediction models use a patient's covariates at a single time point to estimate the time until a specific event occurs, such as death or hospital readmission. However, in many longitudinal datasets, patient covariates are recorded at multiple time points, typically with varying intervals. Our objective is to learn a survival prediction model by training on longitudinal datasets with non-uniform time intervals between covariate measurements, both within and across patient trajectories. METHODS: We propose a new algorithm, Temporally Consistent Multi-Task Logistic Regression (TC-MTLR), which incorporates concepts from distributional reinforcement learning to model survival outcomes. Unlike existing dynamic survival prediction algorithms, TC-MTLR is designed to leverage the non-uniformity of longitudinal measurements. We evaluate this method against two standard and two dynamic survival prediction algorithms across three short and three long longitudinal datasets, including two related to healthcare. RESULTS: On short datasets, TC-MTLR achieves top performance in Concordance Index (C-Index) and Uncensored Mean Average Error (MAE-Uncensored) while displaying mixed results according to Integrated Brier Score (IBS) and Pseudo-Observable MAE (MAE-PO). However, on long datasets, TC-MTLR achieves similar C-Index performance as the other survival predictions methods while outperforming them according to MAE-PO and achieving top performance according to MAE-Uncensored and IBS. CONCLUSION: TC-MTLR effectively utilizes the non-uniform temporal structure of longitudinal datasets, offering a competitive and often superior alternative to existing survival prediction models.
Harrison Fah, Russell Greiner, Roger A. Dixon
J. Biomed. Informatics2
2026 Learning dynamic binary treatment policies under treatment selection bias: A Conservative Q-Learning approach with representation balancing
abstract
OBJECTIVE: Learn safe, robust dynamic treatment regimes (DTRs) from observational trajectories that exhibit treatment selection bias, using an offline reinforcement learning (RL) approach. METHODS: We propose CQL-RB, which augments Conservative Q-Learning (CQL) with a representation-balancing penalty based on an integral probability metric (IPM) (instantiated as either a maximum mean discrepancy (MMD) or an energy-distance penalty). The penalty aligns latent patient representations across treatment groups to reduce action-conditioned distribution shift while preserving CQL's conservative policy estimation. We evaluate CQL-RB on two clinically realistic simulators: EpiCare (eight environments) and AhnChemo from DTR-Bench, both modeling longitudinal healthcare decisions with binary actions at each stage. To emulate selection bias, we implement clinician-like behavior policies that assign treatment as a function of patient covariates. Baselines include BOWL, ACWL, T-RL, RL-NN, and standard CQL. Outcomes are expected return and adverse-event counts from simulator rollouts; model selection uses weighted importance sampling off-policy evaluation on held-out data. Ablations vary both the IPM weight β and the choice of IPM metric. RESULTS: Across all eight EpiCare environments and the challenging AhnChemo task, CQL-RB with either MMD or energy-distance penalties consistently achieves higher returns than competing methods while yielding lower (or comparable) adverse-event rates. Removing the balancing term degrades both return and safety, confirming its contribution. Performance is robust for moderate penalty weights (e.g., β∈{1,10,100}), with degradation only at overly large values (e.g., β≥1000 for MMD or β=10000 for energy distance). CONCLUSION: Representation balancing materially strengthens conservative offline RL for DTR learning under treatment selection bias. By aligning patient representations without altering CQL's safety mechanics, CQL-RB delivers policies that are both effective (higher returns) and safer (fewer adverse events) in realistic healthcare simulations. These findings underscore the importance of addressing treatment selection bias when learning robust and safe dynamic treatment policies.
Animesh Kumar Paul, Russell Greiner
J. Biomed. Informatics2
2025 Extendable and Iterative Structure Learning Strategy for Bayesian Networks
abstract
Learning the structure of Bayesian networks is a fundamental yet computationally intensive task, especially as the number of variables grows. Traditional algorithms require retraining from scratch when new variables are introduced, making them impractical for dynamic or large-scale applications. In this paper, we propose an extendable structure learning strategy that efficiently incorporates a new variable $Y$ into an existing Bayesian network graph $\mathcal{G}$ over variables $\mathcal{X}$, resulting in an updated P-map graph $\bar{\mathcal{G}}$ on $\bar{\mathcal{X}} = \mathcal{X} \cup \{Y\}$. By leveraging the information encoded in $\mathcal{G}$, our method significantly reduces computational overhead compared to learning $\bar{\mathcal{G}}$ from scratch. Empirical evaluations demonstrate runtime reductions of up to 1300x without compromising accuracy. Building on this approach, we introduce a novel iterative paradigm for structure learning over $\mathcal{X}$. Starting with a small subset $\mathcal{U} \subset \mathcal{X}$, we iteratively add the remaining variables using our extendable algorithms to construct a P-map graph over the full set. This method offers runtime advantages comparable to common algorithms while maintaining similar accuracy. Our contributions provide a scalable solution for Bayesian network structure learning, enabling efficient model updates in real-time and high-dimensional settings.
Hamid Kalantari, Russell Greiner, Pouria Ramazi
ICLR2
2025 Early detection of disease outbreaks and non-outbreaks using incidence data: A framework using feature-based time series classification and machine learning
abstract
Forecasting the occurrence and absence of novel disease outbreaks is essential for disease management, yet existing methods are often context-specific, require a long preparation time, and non-outbreak prediction remains understudied. To address this gap, we propose a novel framework using a feature-based time series classification (TSC) method to forecast outbreaks and non-outbreaks. We tested our methods on synthetic data from a Susceptible-Infected-Recovered (SIR) model for slowly changing, noisy disease dynamics. Outbreak sequences give a transcritical bifurcation within a specified future time window, whereas non-outbreak (null bifurcation) sequences do not. We identified incipient differences, reflected in 22 statistical features and 5 early warning signal indicators, in time series of infectives leading to future outbreaks and non-outbreaks. Classifier performance, given by the area under the receiver-operating curve (AUC), ranged from 0 . 99 for large expanding windows of training data to 0 . 7 for small rolling windows. The framework is further evaluated on four empirical datasets: COVID-19 incidence data from Singapore, 18 other countries, and Edmonton, Canada, as well as SARS data from Hong Kong, with two classifiers exhibiting consistently high accuracy. Our results highlight detectable statistical features distinguishing outbreak and non-outbreak sequences well before potential occurrence, in both synthetic and real-world datasets presented in this study.
Amit K. Chakraborty, Russell Greiner, Mark A. Lewis, Hao Wang 0027
PLoS Comput. Biol.3
2024 Conformalized Survival Distributions: A Generic Post-Process to Increase Calibration
abstract
Discrimination and calibration represent two important properties of survival analysis, with the former assessing the model’s ability to accurately rank subjects and the latter evaluating the alignment of predicted outcomes with actual events. With their distinct nature, it is hard for survival models to simultaneously optimize both of them especially as many previous results found improving calibration tends to diminish discrimination performance. This paper introduces a novel approach utilizing conformal regression that can improve a model’s calibration without degrading discrimination. We provide theoretical guarantees for the above claim, and rigorously validate the efficiency of our approach across 11 real-world datasets, showcasing its practical applicability and robustness in diverse scenarios.
Shiang Qi, Yakun Yu, Russell Greiner
ICML3
2024 MassSpecGym: A benchmark for the discovery and identification of molecules
abstract
The discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure from its mass spectrum is exceptionally challenging, even when performed by human experts. As a result, the vast majority of acquired MS/MS spectra remain uninterpreted, thereby limiting our understanding of the underlying (bio)chemical processes. Despite decades of progress in machine learning applications for predicting molecular structures from MS/MS spectra, the development of new methods is severely hindered by the lack of standard datasets and evaluation protocols. To address this problem, we propose MassSpecGym -- the first comprehensive benchmark for the discovery and identification of molecules from MS/MS data. Our benchmark comprises the largest publicly available collection of high-quality MS/MS spectra and defines three MS/MS annotation challenges: \textit{de novo} molecular structure generation, molecule retrieval, and spectrum simulation. It includes new evaluation metrics and a generalization-demanding data split, therefore standardizing the MS/MS annotation tasks and rendering the problem accessible to the broad machine learning community. MassSpecGym is publicly available at \url{https://github.com/pluskal-lab/MassSpecGym}.
Roman Bushuiev, Anton Bushuiev, Niek F. de Jonge, Adamo Young, Fleming Kretschmer, Raman Samusevich, Janne Heirman, Fei Wang 0062, Luke Zhang, Kai Dührkop, Marcus Ludwig, Nils A. Haupt, Apurva Kalia, Corinna Brungs, Robin Schmid, Russell Greiner, Bo Wang 0044, David S. Wishart, Liping Liu 0001, Juho Rousu, Wout Bittremieux, Hannes L. Röst, Tytus D. Mak, Soha Hassoun, Florian Huber 0001, Justin J. J. van der Hooft, Michael A. Stravs, Sebastian Böcker, Josef Sivic, Tomás Pluskal
NeurIPS16
2024 Toward Conditional Distribution Calibration in Survival Prediction
abstract
Survival prediction often involves estimating the time-to-event distribution from censored datasets. Previous approaches have focused on enhancing discrimination and marginal calibration. In this paper, we highlight the significance of *conditional calibration* for real-world applications – especially its role in individual decision-making. We propose a method based on conformal prediction that uses the model’s predicted individual survival probability at that instance’s observed time. This method effectively improves the model’s marginal and conditional calibration, without compromising discrimination. We provide asymptotic theoretical guarantees for both marginal and conditional calibration and test it extensively across 15 diverse real-world datasets, demonstrating the method’s practical effectiveness and versatility in various settings.
Shiang Qi, Yakun Yu, Russell Greiner
NeurIPS3
2024 The ACROBAT 2022 challenge: Automatic registration of breast cancer tissue
abstract
The alignment of tissue between histopathological whole-slide-images (WSI) is crucial for research and clinical applications. Advances in computing, deep learning, and availability of large WSI datasets have revolutionised WSI analysis. Therefore, the current state-of-the-art in WSI registration is unclear. To address this, we conducted the ACROBAT challenge, based on the largest WSI registration dataset to date, including 4,212 WSIs from 1,152 breast cancer patients. The challenge objective was to align WSIs of tissue that was stained with routine diagnostic immunohistochemistry to its H&E-stained counterpart. We compare the performance of eight WSI registration algorithms, including an investigation of the impact of different WSI properties and clinical covariates. We find that conceptually distinct WSI registration methods can lead to highly accurate registration performances and identify covariates that impact performances across methods. These results provide a comparison of the performance of current WSI registration methods and guide researchers in selecting and developing methods.
Philippe Weitz, Masi Valkonen, Leslie Solorzano, Circe Carr, Kimmo Kartasalo, Constance Boissin, Sonja Koivukoski, Aino Kuusela, Dusan Rasic, Yanbo Feng, Sandra Kristiane Sinius Pouplier, Kajsa Ledesma Eriksson, Stephanie Robertson, Christian Marzahl, Chandler Gatenbee, Alexander R. A. Anderson, Marek Wodzinski, Artur Jurgas, Niccolò Marini, Manfredo Atzori, Henning Müller, Daniel Budelmann, Nick Weiss, Stefan Heldmann, Johannes Lotz 0002, Jelmer M. Wolterink, Bruno De Santi, Abhijeet Patil, Amit Sethi, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Mahtab Farrokh, Neeraj Kumar 0002, Russell Greiner, Leena Latonen, Anne-Vibeke Laenkholm, Johan Hartman, Pekka Ruusuvuori, Mattias Rantalainen
Medical Image Anal.36
2023 Exploring Language-Agnostic Speech Representations Using Domain Knowledge for Detecting Alzheimer's Dementia
abstract
We explore ways to use speech data to screen for indications of Alzheimer’s dementia (AD). In particular, we describe our approach to the ICASSP 2023 Signal Processing Grand Challenge, which involves extrapolating from models learned from English speech samples, to Greek speech samples, to determine which subjects have AD. By using acoustic and linguistic features, inspired by clinical research on AD, our top-performing classification model achieves 69% accuracy in distinguishing AD patients from healthy controls, and our regression model attains an RMSE of 4.8 for inferring cognitive testing scores. These outcomes underscore the potential of our explainable model for detecting cognitive decline in AD patients via speech, and its applicability in clinical settings.
Zehra Shah, Shiang Qi, Fei Wang 0062, Mahtab Farrokh, Mashrura Tasnim, Eleni Stroulia, Russell Greiner, Manos Plitsis, Athanasios Katsamanis
ICASSP7
2023 An Effective Meaningful Way to Evaluate Survival Models
abstract
One straightforward metric to evaluate a survival prediction model is based on the Mean Absolute Error (MAE) – the average of the absolute difference between the time predicted by the model and the true event time, over all subjects. Unfortunately, this is challenging because, in practice, the test set includes (right) censored individuals, meaning we do not know when a censored individual actually experienced the event. In this paper, we explore various metrics to estimate MAE for survival datasets that include (many) censored individuals. Moreover, we introduce a novel and effective approach for generating realistic semi-synthetic survival datasets to facilitate the evaluation of metrics. Our findings, based on the analysis of the semi-synthetic datasets, reveal that our proposed metric (MAE using pseudo-observations) is able to rank models accurately based on their performance, and often closely matches the true MAE – in particular, is better than several alternative methods.
Shiang Qi, Neeraj Kumar 0002, Mahtab Farrokh, Weijie Sun 0004, Li-Hao Kuan, Rajesh Ranganath, Ricardo Henao, Russell Greiner
ICML8
2023 Copula-based deep survival models for dependent censoring
abstract
A survival dataset describes a set of instances (e.g. patients) and provides, for each, either the time until an event (e.g. death), or the censoring time (e.g. when lost to follow-up - which is a lower bound on the time until the event). We consider the challenge of survival prediction: learning, from such data, a predictive model that can produce an individual survival distribution for a novel instance. Many contemporary methods of survival prediction implicitly assume that the event and censoring distributions are independent conditional on the instance’s covariates - a strong assumption that is difficult to verify (as we observe only one outcome for each instance) and which can induce significant bias when it does not hold. This paper presents a parametric model of survival that extends modern non-linear survival analysis by relaxing the assumption of conditional independence. On synthetic and semi-synthetic data, our approach significantly improves estimates of survival distributions compared to the standard that assumes conditional independence in the data.
Ali Hossein Gharari Foomani, Michael Cooper, Russell Greiner, Rahul G. Krishnan
UAI3
2021 Sample efficient learning of image-based diagnostic classifiers via probabilistic labels
abstract
Deep learning approaches often require huge datasets to achieve good generalization. This complicates its use in tasks like image-based medical diagnosis, where the small training datasets are usually insufficient to learn appropriate data representations. For such sensitive tasks it is also important to provide the confidence in the predictions. Here, we propose a way to learn and use probabilistic labels to train accurate and calibrated deep networks from relatively small datasets. We observe gains of up to 22% in the accuracy of models trained with these labels, as compared with traditional approaches, in three classification tasks: diagnosis of hip dysplasia, fatty liver, and glaucoma. The outputs of models trained with probabilistic labels are calibrated, allowing the interpretation of its predictions as proper probabilities. We anticipate this approach will apply to other tasks where few training instances are available and expert knowledge can be encoded as probabilities.
Roberto Vega, Pouneh Gorji, Zichen Zhang 0001, Xuebin Qin, Abhilash Rakkunedeth Hareendranathan, Jeevesh Kapur, Jacob L. Jaremko, Russell Greiner
AISTATS8
2020 Learning Disentangled Representations for CounterFactual Regression
Negar Hassanpour, Russell Greiner
ICLR2
2020 Domain Aggregation Networks for Multi-Source Domain Adaptation
abstract
In many real-world applications, we want to exploit multiple source datasets to build a model for a different but related target dataset. Despite the recent empirical success, most existing research has used ad-hoc methods to combine multiple sources, leading to a gap between theory and practice. In this paper, we develop a finite-sample generalization bound based on domain discrepancy and accordingly propose a theoretically justified optimization procedure. Our algorithm, Domain AggRegation Network (DARN), can automatically and dynamically balance between including more data to increase effective sample size and excluding irrelevant data to avoid negative effects during training. We find that DARN can significantly outperform the state-of-the-art alternatives on multiple real-world tasks, including digit/object recognition and sentiment analysis.
Junfeng Wen, Russell Greiner, Dale Schuurmans
ICML2
2020 Shared Space Transfer Learning for analyzing multi-site fMRI data
abstract
Multi-voxel pattern analysis (MVPA) learns predictive models from task-based functional magnetic resonance imaging (fMRI) data, for distinguishing when subjects are performing different cognitive tasks — e.g., watching movies or making decisions. MVPA works best with a well-designed feature set and an adequate sample size. However, most fMRI datasets are noisy, high-dimensional, expensive to collect, and with small sample sizes. Further, training a robust, generalized predictive model that can analyze homogeneous cognitive tasks provided by multi-site fMRI datasets has additional challenges. This paper proposes the Shared Space Transfer Learning (SSTL) as a novel transfer learning (TL) approach that can functionally align homogeneous multi-site fMRI datasets, and so improve the prediction performance in every site. SSTL first extracts a set of common features for all subjects in each site. It then uses TL to map these site-specific features to a site-independent shared space in order to improve the performance of the MVPA. SSTL uses a scalable optimization procedure that works effectively for high-dimensional fMRI datasets. The optimization procedure extracts the common features for each site by using a single-iteration algorithm and maps these site-specific common features to the site-independent shared space. We evaluate the effectiveness of the proposed method for transferring between various cognitive tasks. Our comprehensive experiments validate that SSTL achieves superior performance to other state-of-the-art analysis techniques.
Muhammad Yousefnezhad, Alessandro Selvitella, Daoqiang Zhang, Andrew J. Greenshaw, Russell Greiner
NeurIPS5
2020 Effective Ways to Build and Evaluate Individual Survival Distributions
abstract
An accurate model of a patient’s individual survival distribution can help determine the appropriate treatment for terminal patients. Unfortunately, risk scores (for example from Cox Proportional Hazard models) do not provide survival probabilities, single-time probability models (for instance the Gail model, predicting 5 year probability) only provide for a single time point, and standard Kaplan-Meier survival curves provide only population averages for a large class of patients, meaning they are not specific to individual patients. This motivates an alternative class of tools that can learn a model that provides an individual survival distribution for each subject, which gives survival probabilities across all times, such as extensions to the Cox model, Accelerated Failure Time, an extension to Random Survival Forests, and Multi-Task Logistic Regression. This paper first motivates such 'individual survival distribution' (ISD) models, and explains how they differ from standard models. It then discusses ways to evaluate such models — namely Concordance, 1-Calibration, Integrated Brier score, and versions of L1-loss — then motivates and defines a novel approach, 'D-Calibration', which determines whether a model's probability estimates are meaningful. We also discuss how these measures differ, and use them to evaluate several ISD prediction tools over a range of survival data sets. We also provide a code base for all of these survival models and evaluation measures, at https://github.com/haiderstats/ISDEvaluation.
Humza Haider, Bret Hoehn, Sarah Davis, Russell Greiner
J. Mach. Learn. Res.4
2019 Ischemic Stroke Lesion Prediction in CT Perfusion Scans Using Multiple Parallel U-Nets Following by a Pixel-Level Classifier
abstract
It is critical to know what brain regions are affected by an ischemic stroke, as this enables doctors to make more effective decisions about stroke patient therapy. These regions are often identified by segmenting computed tomography perfusion (CTP) images. Previously, this task has been done manually by an expert. However, manual segmentation is an extremely tedious and time-consuming process, that is not suitable for ischemic stroke lesion segmentation as it is highly time sensitive. In addition, these approaches require an expert to do the segmentation task, who may not be available and are prone to errors. Several automatic medical image analysis methods have been proposed for ischemic stroke lesion segmentation. These approaches, typically, use hand-crafted features that are predefined to represent the input data. However, because of the irregular and physiologically shapes, ischemic stroke lesions cannot be properly predicted, in an automatic way, using simple predefined features. In this work, we propose an automatic prediction algorithm that learns an effective model for segmenting the ischemic stroke lesion. This learned model first uses four 2D U-Nets to, separately, extract valuable information about the location of the stroke lesion from four CTP maps (CBV, CBF, MTT, Tmax). The model then combines the probability maps extracted by the U-Nets, to decide whether the pixels are either lesion or healthy tissues. This approach uses information about each pixel, as well as its neighborhood, to learn the stroke lesion, despite their varying shapes. The segmentation performance is evaluated using dice similarity coefficient (DSC), volume similarity (VS), and Recall. We have used this new algorithm on ISLES 2018 challenge dataset and found that our approach achieved results that are better than state-of-the-art approaches.
Mohsen Soltanpour, Russell Greiner, Pierre Boulanger, Brian Buck
BIBE2
2019 CounterFactual Regression with Importance Sampling Weights
abstract
Perhaps the most pressing concern of a patient diagnosed with cancer is her life expectancy under various treatment options. For a binary-treatment case, this translates into estimating the difference between the outcomes (e.g., survival time) of the two available treatment options – i.e., her Individual Treatment Effect (ITE). This is especially challenging to estimate from observational data, as that data has selection bias: the treatment assigned to a patient depends on that patient's attributes. In this work, we borrow ideas from domain adaptation to address the distributional shift between the source (outcome of the administered treatment, appearing in the observed training data) and target (outcome of the alternative treatment) that exists due to selection bias. We propose a context-aware importance sampling re-weighing scheme, built on top of a representation learning module, for estimating ITEs. Empirical results on two publicly available benchmarks demonstrate that the proposed method significantly outperforms state-of-the-art.
Negar Hassanpour, Russell Greiner
IJCAI2
2019 Simultaneous Prediction Intervals for Patient-Specific Survival Curves
abstract
Accurate models of patient survival probabilities provide important information to clinicians prescribing care for life-threatening and terminal ailments. A recently developed class of models -- known as individual survival distributions (ISDs) -- produces patient-specific survival functions that offer greater descriptive power of patient outcomes than was previously possible. Unfortunately, at the time of writing, ISD models almost universally lack uncertainty quantification. In this paper we demonstrate that an existing method for estimating simultaneous prediction intervals from samples can easily be adapted for patient-specific survival curve analysis and yields accurate results. Furthermore, we introduce both a modification to the existing method and a novel method for estimating simultaneous prediction intervals and show that they offer competitive performance. It is worth emphasizing that these methods are not limited to survival analysis and can be applied in any context in which sampling the distribution of interest is tractable. Code is available at https://github.com/ssokota/spie.
Samuel Sokota, Ryan D'Orazio, Khurram Javed, Humza Haider, Russell Greiner
IJCAI5
2019 Learning Macroscopic Brain Connectomes via Group-Sparse Factorization
abstract
Mapping structural brain connectomes for living human brains typically requires expert analysis and rule-based models on diffusion-weighted magnetic resonance imaging. A data-driven approach, however, could overcome limitations in such rule-based approaches and improve precision mappings for individuals. In this work, we explore a framework that facilitates applying learning algorithms to automatically extract brain connectomes. Using a tensor encoding, we design an objective with a group-regularizer that prefers biologically plausible fascicle structure. We show that the objective is convex and has unique solutions, ensuring identifiable connectomes for an individual. We develop an efficient optimization strategy for this extremely high-dimensional sparse problem, by reducing the number of parameters using a greedy algorithm designed specifically for the problem. We show that this greedy algorithm significantly improves on a standard greedy algorithm, called Orthogonal Matching Pursuit. We conclude with an analysis of the solutions found by our method, showing we can accurately reconstruct the diffusion information while maintaining contiguous fascicles with smooth direction changes.
Farzane Aminmansour, Andrew Patterson, Lei Le, Yisu Peng, Daniel Mitchell, Franco Pestilli, Cesar F. Caiafa, Russell Greiner, Martha White
NeurIPS8
2019 Low-Dimensional Perturb-and-MAP Approach for Learning Restricted Boltzmann Machines
abstract
This paper introduces a new approach to maximum likelihood learning of the parameters of a restricted Boltzmann machine (RBM). The proposed method is based on the Perturb-and-MAP (PM) paradigm that enables sampling from the Gibbs distribution. PM is a two step process: (i) perturb the model using Gumbel perturbations, then (ii) find the maximum a posteriori (MAP) assignment of the perturbed model. We show that under certain conditions the resulting MAP configuration of the perturbed model is an unbiased sample from the original distribution. However, this approach requires an exponential number of perturbations, which is computationally intractable. Here, we apply an approximate approach based on the first order (low-dimensional) PM to calculate the gradient of the log-likelihood in binary RBM. Our approach relies on optimizing the energy function with respect to observable and hidden variables using a greedy procedure. First, for each variable we determine whether flipping this value will decrease the energy, and then we utilize the new local maximum to approximate the gradient. Moreover, we show that in some cases our approach works better than the standard coordinate-descent procedure for finding the MAP assignment and compare it with the Contrastive Divergence algorithm. We investigate the quality of our approach empirically, first on toy problems, then on various image datasets and a text dataset.
Jakub M. Tomczak, Szymon Zareba, Siamak Ravanbakhsh, Russell Greiner
Neural Process. Lett.4
2018 Analyzing the effects of test driven development in GitHub
abstract
Testing is an integral part of the software development lifecycle, approached with varying degrees of rigor by different process models. Agile process models recommend Test Driven Development (TDD) as a key practice for reducing costs and improving code quality. The objective of this work is to perform a cost-benefit analysis of this practice. Previous work by Fucci et al. [2, 3] engaged in laboratory studies of developers actively engaged in test-driven development practices. Fucci et al. found little difference between test-first behaviour of TDD and test-later behaviour. To that end, we opted to conduct a study about TDD behaviours in the "wild" rather than in the laboratory. Thus we have conducted a comparative analysis of GitHub repositories that adopts TDD to a lesser or greater extent, in order to determine how TDD affects software development productivity and software quality. We classified GitHub repositories archived in 2015 in terms of how rigorously they practiced TDD, thus creating a TDD spectrum. We then matched and compared various subsets of these repositories on this TDD spectrum with control sets of equal size. The control sets were samples from all GitHub repositories that matched certain characteristics, and that contained at least one test file. We compared how the TDD sets differed from the control sets on the following characteristics: number of test files, average commit velocity, number of bug-referencing commits, number of issues recorded, usage of continuous integration, number of pull requests, and distribution of commits per author. We found that Java TDD projects were relatively rare. In addition, there were very few significant differences in any of the metrics we used to compare TDD-like and non-TDD projects; therefore, our results do not reveal any observable benefits from using TDD.
Neil Borle, Meysam Feghhi, Eleni Stroulia, Russell Greiner, Abram Hindle
ICSE4
2018 Analyzing the effects of test driven development in GitHub
Neil Borle, Meysam Feghhi, Eleni Stroulia, Russell Greiner, Abram Hindle
Empir. Softw. Eng.4
2017 Deep Green: Modelling Time-Series of Software Energy Consumption
abstract
Inefficient mobile software kills battery life. Yet, developers lack the tools necessary to detect and solve energy bugs in software. In addition, developers are usually tasked with the creation of software features and triaging existing bugs. This means that most developers do not have the time or resources to research, build, or employ energy debugging tools. We present a new method for predicting software energy consumption to help debug software energy issues. Our approach enables developers to align traces of software behavior with traces of software energy consumption. This allows developers to match run-time energy hot spots to the corresponding execution. We accomplish this by applying recent neural network models to predict time series of energy consumption given a software's behavior. We compare our time series models to prior state-of-the-art models that only predict total software energy consumption. We found that machine learning based time series based models, and LSTM based time series based models, can often be more accurate at predicting instantaneous power use and total energy consumption.
Stephen Romansky, Neil Borle, Shaiful Alam Chowdhury, Abram Hindle, Russell Greiner
ICSME5
2017 Detecting duplicate bug reports with software engineering domain knowledge
abstract
Bug deduplication, ie, recognizing bug reports that refer to the same problem, is a challenging task in the software‐engineering life cycle. Researchers have proposed several methods primarily relying on information‐retrieval techniques. Our work motivated by the intuition that domain knowledge can provide the relevant context to enhance effectiveness, attempts to improve the use of information retrieval by augmenting with software‐engineering knowledge. In our previous work, we proposed the software‐literature‐context method for using software‐engineering literature as a source of contextual information to detect duplicates. If bug reports relate to similar subjects, they have a better chance of being duplicates. Our method, being largely automated, has a potential to substantially decrease the level of manual effort involved in conventional techniques with a minor trade‐off in accuracy. In this study, we extend our work by demonstrating that domain‐specific features can be applied across projects than project‐specific features demonstrated previously while still maintaining performance. We also introduce a hierarchy‐of‐context to capture the software‐engineering knowledge in the realms of contextual space to produce performance gains. We also highlight the importance of domain‐specific contextual features through cross‐domain contexts: adding context improved accuracy; Kappa scores improved by at least 3.8% to 10.8% per project.
Karan Aggarwal, Finbarr Timbers, Tanner Rutgers, Abram Hindle, Eleni Stroulia, Russell Greiner
J. Softw. Evol. Process.6
2016 Stochastic Neural Networks with Monotonic Activation Functions
abstract
We propose a Laplace approximation that creates a stochastic unit from any smooth monotonic activation function, using only Gaussian noise. This paper investigates the application of this stochastic approximation in training a family of Restricted Boltzmann Machines (RBM) that are closely linked to Bregman divergences. This family, that we call exponential family RBM (Exp-RBM), is a subset of the exponential family Harmoniums that expresses family members through a choice of smooth monotonic non-linearity for each neuron. Using contrastive divergence along with our Gaussian approximation, we show that Exp-RBM can learn useful representations using novel stochastic units.
Siamak Ravanbakhsh, Barnabás Póczos, Jeff G. Schneider, Dale Schuurmans, Russell Greiner
AISTATS5
2016 Boolean Matrix Factorization and Noisy Completion via Message Passing
abstract
Boolean matrix factorization and Boolean matrix completion from noisy observations are desirable unsupervised data-analysis methods due to their interpretability, but hard to perform due to their NP-hardness. We treat these problems as maximum a posteriori inference problems in a graphical model and present a message passing approach that scales linearly with the number of observations and factors. Our empirical study demonstrates that message passing is able to recover low-rank Boolean matrices, in the boundaries of theoretically possible recovery and compares favorably with state-of-the-art in real-world applications, such collaborative filtering with large-scale Boolean data.
Siamak Ravanbakhsh, Barnabás Póczos, Russell Greiner
ICML3
2015 Correcting Covariate Shift with the Frank-Wolfe Algorithm
Junfeng Wen, Russell Greiner, Dale Schuurmans
IJCAI2
2015 Detecting duplicate bug reports with software engineering domain knowledge
abstract
In previous work by Alipour et al., a methodology was proposed for detecting duplicate bug reports by comparing the textual content of bug reports to subject-specific contextual material, namely lists of software-engineering terms, such as non-functional requirements and architecture keywords. When a bug report contains a word in these word-list contexts, the bug report is considered to be associated with that context and this information tends to improve bug-deduplication methods. In this paper, we propose a method to partially automate the extraction of contextual word lists from software-engineering literature. Evaluating this software-literature context method on real-world bug reports produces useful results that indicate this semi-automated method has the potential to substantially decrease the manual effort used in contextual bug deduplication while suffering only a minor loss in accuracy.
Karan Aggarwal, Tanner Rutgers, Finbarr Timbers, Abram Hindle, Russell Greiner, Eleni Stroulia
SANER5
2015 Perturbed message passing for constraint satisfaction problems
Siamak Ravanbakhsh, Russell Greiner
J. Mach. Learn. Res.2
2014 Budgeted transcript discovery: A framework for joint exploration and validation studies
abstract
This paper presents the budgeted transcript discovery problem (BTD): deciding how to spend a given research budget collecting data, using a combination of microarrays and PCRs, to discover which transcripts are differentially expressed with respect to a given phenotype. We present algorithms that address this task by sequentially analyzing the data collected so far, to decide which data would be most informative to collect next. We provide empirical studies that demonstrate their effectiveness.
Sheehan Khan, Russell Greiner
BIBM2
2014 A robust convergence index filter for breast cancer cell segmentation
abstract
COnvergence INdex (COIN) filter, a successful tool for cell localization, evaluates the degree of convergence of the gradient vectors within the neighborhood (region of support) toward a pixel of interest. All previous efforts were to increase the adaptability of the region of support to make the COIN filter robust and accurate. However, improving the quality of the image gradient map was ignored, which results in poor performance of the members of the COIN family in noisy settings. We propose a new Robust Convergence Index (RCI) filter that tailors the COIN filter in a noisy environment by (a) spreading the gradient vectors within non-homogeneous object regions by convolving an Aggregated Edge Probability Map (AEPM) with an edge preserving gradient vector kernel, and (b) increasing the convergence of the gradient vectors through the integration of the sine and cosine distribution as well as the magnitude of the gradient vectors. AEPM is computed through the consensus of the responses of a number of edge detectors over a wide range of scales, which lessens the effects of clutter by enforcing higher weights to the actual edges, and a non-parametric Kernel Density Estimation (KDE) is used to compute the edge probability map. Experimental results demonstrate that it obtains state-of-the-art performance.
Baidya Nath Saha, Amritpal Saini, Nilanjan Ray, Russell Greiner, Judith Hugh, Mauro Tambasco
ICIP4
2014 Min-Max Problems on Factor Graphs
abstract
We study the min-max problem in factor graphs, which seeks the assignment that minimizes the maximum value over all factors. We reduce this problem to both min-sum and sum-product inference, and focus on the later. This approach reduces the min-max inference problem to a sequence of constraint satisfaction problems (CSPs) which allows us to sample from a uniform distribution over the set of solutions. We demonstrate how this scheme provides a message passing solution to several NP-hard combinatorial problems, such as min-max clustering (a.k.a. K-clustering), the asymmetric K-center problem, K-packing and the bottleneck traveling salesman problem. Furthermore we theoretically relate the min-max reductions to several NP hard decision problems, such as clique cover, set cover, maximum clique and Hamiltonian cycle, therefore also providing message passing solutions for these problems. Experimental results suggest that message passing often provides near optimal min-max solutions for moderate size instances.
Siamak Ravanbakhsh, Christopher Srinivasa, Brendan J. Frey, Russell Greiner
ICML4
2014 Robust Learning under Uncertain Test Distributions: Relating Covariate Shift to Model Misspecification
abstract
Many learning situations involve learning the conditional distribution p(y|x) when the training instances are drawn from the training distribution p_tr(x), even though it will later be used to predict for instances drawn from a different test distribution p_te(x). Most current approaches focus on learning how to reweigh the training examples, to make them resemble the test distribution. However, reweighing does not always help, because (we show that) the test error also depends on the correctness of the underlying model class. This paper analyses this situation by viewing the problem of learning under changing distributions as a game between a learner and an adversary. We characterize when such reweighing is needed, and also provide an algorithm, robust covariate shift adjustment (RCSA), that provides relevant weights. Our empirical studies, on UCI datasets and a real-world cancer prognostic prediction dataset, show that our analysis applies, and that our RCSA works effectively.
Junfeng Wen, Chun-Nam Yu, Russell Greiner
ICML3
2014 Budgeted Learning for Developing Personalized Treatment
abstract
There is increased interest in using patient-specific information to personalize treatment. Personalized treatment decision rules can be learned using data from standard clinical trials, but such trials are very costly to run. This paper explores the use of budgeted learning techniques to design more efficient clinical trials, by effectively determining which type of patients to recruit, at each time, throughout the duration of the trial. We propose a Bayesian bandit model and discuss the computational challenges and issues pertaining to this approach. We compare our budgeted learning algorithm, which approximately minimizes the Bayes risk, using both simulated data and data modeled after a clinical trial for treating depressed individuals, with other plausible algorithms. We show that our budgeted learning algorithm demonstrated excellent performance across a wide variety of situations.
Russell Greiner, Susan A. Murphy
ICMLA2
2014 Augmentative Message Passing for Traveling Salesman Problem and Graph Partitioning
Siamak Ravanbakhsh, Reihaneh Rabbany, Russell Greiner
NIPS3
2013 Online Learning with Costly Features and Labels
abstract
This paper introduces the online probing" problem: In each round, the learner is able to purchase the values of a subset of feature values. After the learner uses this information to come up with a prediction for the given round, he then has the option of paying for seeing the loss that he is evaluated against. Either way, the learner pays for the imperfections of his predictions and whatever he chooses to observe, including the cost of observing the loss function for the given round and the cost of the observed features. We consider two variations of this problem, depending on whether the learner can observe the label for free or not. We provide algorithms and upper and lower bounds on the regret for both variants. We show that a positive cost for observing the label significantly increases the regret of the problem."
Navid Zolghadr, Gábor Bartók, Russell Greiner, András György 0001, Csaba Szepesvári
NIPS3
2013 Breast cancer prediction using genome wide single nucleotide polymorphism data
abstract
BACKGROUND: This paper introduces and applies a genome wide predictive study to learn a model that predicts whether a new subject will develop breast cancer or not, based on her SNP profile. RESULTS: We first genotyped 696 female subjects (348 breast cancer cases and 348 apparently healthy controls), predominantly of Caucasian origin from Alberta, Canada using Affymetrix Human SNP 6.0 arrays. Then, we applied EIGENSTRAT population stratification correction method to remove 73 subjects not belonging to the Caucasian population. Then, we filtered any SNP that had any missing calls, whose genotype frequency was deviated from Hardy-Weinberg equilibrium, or whose minor allele frequency was less than 5%. Finally, we applied a combination of MeanDiff feature selection method and KNN learning method to this filtered dataset to produce a breast cancer prediction model. LOOCV accuracy of this classifier is 59.55%. Random permutation tests show that this result is significantly better than the baseline accuracy of 51.52%. Sensitivity analysis shows that the classifier is fairly robust to the number of MeanDiff-selected SNPs. External validation on the CGEMS breast cancer dataset, the only other publicly available breast cancer dataset, shows that this combination of MeanDiff and KNN leads to a LOOCV accuracy of 60.25%, which is significantly better than its baseline of 50.06%. We then considered a dozen different combinations of feature selection and learning method, but found that none of these combinations produces a better predictive model than our model. We also considered various biological feature selection methods like selecting SNPs reported in recent genome wide association studies to be associated with breast cancer, selecting SNPs in genes associated with KEGG cancer pathways, or selecting SNPs associated with breast cancer in the F-SNP database to produce predictive models, but again found that none of these models achieved accuracy better than baseline. CONCLUSIONS: We anticipate producing more accurate breast cancer prediction models by recruiting more study subjects, providing more accurate labelling of phenotypes (to accommodate the heterogeneity of breast cancer), measuring other genomic alterations such as point mutations and copy number variations, and incorporating non-genetic information about subjects such as environmental and lifestyle factors.
Mohsen Hajiloo, Babak Damavandi, Metanat HooshSadat, Farzad Sangi, John R. Mackey, Carol E. Cass, Russell Greiner, Sambasivarao Damaraju
BMC Bioinform.7
2013 ETHNOPRED: a novel machine learning method for accurate continental and sub-continental ancestry identification and population stratification correction
abstract
BACKGROUND: Population stratification is a systematic difference in allele frequencies between subpopulations. This can lead to spurious association findings in the case-control genome wide association studies (GWASs) used to identify single nucleotide polymorphisms (SNPs) associated with disease-linked phenotypes. Methods such as self-declared ancestry, ancestry informative markers, genomic control, structured association, and principal component analysis are used to assess and correct population stratification but each has limitations. We provide an alternative technique to address population stratification. RESULTS: We propose a novel machine learning method, ETHNOPRED, which uses the genotype and ethnicity data from the HapMap project to learn ensembles of disjoint decision trees, capable of accurately predicting an individual's continental and sub-continental ancestry. To predict an individual's continental ancestry, ETHNOPRED produced an ensemble of 3 decision trees involving a total of 10 SNPs, with 10-fold cross validation accuracy of 100% using HapMap II dataset. We extended this model to involve 29 disjoint decision trees over 149 SNPs, and showed that this ensemble has an accuracy of ≥ 99.9%, even if some of those 149 SNP values were missing. On an independent dataset, predominantly of Caucasian origin, our continental classifier showed 96.8% accuracy and improved genomic control's λ from 1.22 to 1.11. We next used the HapMap III dataset to learn classifiers to distinguish European subpopulations (North-Western vs. Southern), East Asian subpopulations (Chinese vs. Japanese), African subpopulations (Eastern vs. Western), North American subpopulations (European vs. Chinese vs. African vs. Mexican vs. Indian), and Kenyan subpopulations (Luhya vs. Maasai). In these cases, ETHNOPRED produced ensembles of 3, 39, 21, 11, and 25 disjoint decision trees, respectively involving 31, 502, 526, 242 and 271 SNPs, with 10-fold cross validation accuracy of 86.5% ± 2.4%, 95.6% ± 3.9%, 95.6% ± 2.1%, 98.3% ± 2.0%, and 95.9% ± 1.5%. However, ETHNOPRED was unable to produce a classifier that can accurately distinguish Chinese in Beijing vs. Chinese in Denver. CONCLUSIONS: ETHNOPRED is a novel technique for producing classifiers that can identify an individual's continental and sub-continental heritage, based on a small number of SNPs. We show that its learned classifiers are simple, cost-efficient, accurate, transparent, flexible, fast, applicable to large scale GWASs, and robust to missing values.
Mohsen Hajiloo, Yadav Sapkota, John R. Mackey, Paula Robson, Russell Greiner, Sambasivarao Damaraju
BMC Bioinform.5
2013 Exploiting Syntactic, Semantic, and Lexical Regularities in Language Modeling via Directed Markov Random Fields
abstract
We present a directed Markov random field (MRF) model that combinesn‐gram models, probabilistic context‐free grammars (PCFGs), and probabilistic latent semantic analysis (PLSA) for the purpose of statistical language modeling. Even though the composite directed MRF model potentially has an exponential number of loops and becomes a context‐sensitive grammar, we are nevertheless able to estimate its parameters in cubic time using an efficient modified Expectation‐Maximization (EM) method,the generalized inside–outside algorithm, which extends the inside–outside algorithm to incorporate the effects of then‐gram and PLSA language models. We generalize various smoothing techniques to alleviate the sparseness ofn‐gram counts in cases where there are hidden variables. We also derive an analogous algorithm to find the most likely parse of a sentence and to calculate the probability of initial subsequence of a sentence, all generated by the composite language model. Our experimental results on theWall Street Journalcorpus show that we obtain significant reductions in perplexity compared to the state‐of‐the‐art baseline trigram model with Good–Turing and Kneser–Ney smoothing techniques.
Shaomin Wang, Li Cheng 0001, Russell Greiner, Dale Schuurmans
Comput. Intell.4
2012 A Generalized Loop Correction Method for Approximate Inference in Graphical Models
Siamak Ravanbakhsh, Chun-Nam Yu, Russell Greiner
ICML3
2012 Learning to predict ice accretion on electric power lines
Ashkan Zarnani, Petr Musilek, Xiaodi Ke, Russell Greiner
Eng. Appl. Artif. Intell.6
2012 An experimental methodology for response surface optimization methods
Daniel J. Lizotte, Russell Greiner, Dale Schuurmans
J. Glob. Optim.2
2011 Learning Patient-Specific Cancer Survival Distributions as a Sequence of Dependent Regressors
abstract
An accurate model of patient survival time can help in the treatment and care of cancer patients. The common practice of providing survival time estimates based only on population averages for the site and stage of cancer ignores many important individual differences among patients. In this paper, we propose a local regression method for learning patient-specific survival time distribution based on patient attributes such as blood tests and clinical assessments. When tested on a cohort of more than 2000 cancer patients, our method gives survival time predictions that are much more accurate than popular survival analysis models such as the Cox and Aalen regression models. Our results also show that using patient-specific attributes can reduce the prediction error on survival time by as much as 20% when compared to using cancer site and stage only.
Chun-Nam Yu, Russell Greiner, Hsiu-Chin Lin, Vickie Baracos
NIPS2
2011 Using Classifier-Based Nominal Imputation to Improve Machine Learning
Xiaoyuan Su, Russell Greiner, Taghi M. Khoshgoftaar, Amri Napolitano
PAKDD (1)2
2010 A Cross-Entropy Method that Optimizes Partially Decomposable Problems: A New Way to Interpret NMR Spectra
abstract
Some real-world problems are partially decomposable, in that they can be decomposed into a set of coupled sub- problems, that are each relatively easy to solve. However, when these sub-problem share some common variables, it is not sufficient to simply solve each sub-problem in isolation. We develop a technology for such problems, and use it to address the challenge of finding the concentrations of the chemicals that appear in a complex mixture, based on its one-dimensional 1H Nuclear Magnetic Resonance (NMR) spectrum. As each chemical involves clusters of spatially localized peaks, this requires finding the shifts for the clusters and the concentrations of the chemicals, that collectively pro- duce the best match to the observed NMR spectrum. Here, each sub-problem requires finding the chemical concentrations and cluster shifts that can appear within a limited spectrum range; these are coupled as these limited regions can share many chemicals, and so must agree on the concentrations and cluster shifts of the common chemicals. This task motivates CEED: a novel extension to the Cross-Entropy stochastic optimization method constructed to address such partially decomposable problems. Our experimental results in the NMR task show that our CEED system is superior to other well-known optimization methods, and indeed produces the best-known results in this important, real-world application.
Siamak Ravanbakhsh, Barnabás Póczos, Russell Greiner
AAAI3
2010 Budgeted Distribution Learning of Belief Net Parameters
Liuyang Li, Barnabás Póczos, Csaba Szepesvári, Russell Greiner
ICML4
2010 Mind change optimal learning of Bayes net structure from dependency and independency data
Oliver Schulte, Wei Luo 0001, Russell Greiner
Inf. Comput.3
2009 Segmentation of Lung Tumours in Positron Emission Tomography Scans: A Machine Learning Approach
Aliaksei Kerhet, Cormac Small, Harvey Quon, Terence Riauka, Russell Greiner, Alexander McEwan, Wilson Roa
AIME5
2009 A new hybrid method for Bayesian network learning With dependency constraints
abstract
A Bayes net has qualitative and quantitative aspects: The qualitative aspect is its graphical structure that corresponds to correlations among the variables in the Bayes net. The quantitative aspects are the net parameters. This paper develops a hybrid criterion for learning Bayes net structures that is based on both aspects. We combine model selection criteria measuring data fit with correlation information from statistical tests: Given a sample d, search for a structure G that maximizes score(G, d), over the set of structures G that satisfy the dependencies detected in d. We rely on the statistical test only to accept conditional dependencies, not conditional independencies. We show how to adapt local search algorithms to accommodate the observed dependencies. Simulation studies with GES search and the BDeu/BIC scores provide evidence that the additional dependency information leads to Bayes nets that better fit the target model in distribution and structure.
Oliver Schulte, Gustavo Frigo, Russell Greiner, Wei Luo 0001, Hassan Khosravi
CIDM3
2009 Learning to segment from a few well-selected training images
abstract
We address the task of actively learning a segmentation system: given a large number of unsegmented images, and access to an oracle that can segment a given image, decide which images to provide, to quickly produce a segmenter (here, a discriminative random field) that is accurate over this distribution of images. We extend the standard models for active learner to define a system for this task that first selects the image whose expected label will reduce the uncertainty of the other unlabeled images the most, and then after greedily selects, from the pool of unsegmented images, the most informative image. The results of our experiments, over two real-world datasets (segmenting brain tumors within magnetic resonance images; and segmenting the sky in real images) show that training on very few informative images (here, as few as 2) can produce a segmenter that is as good as training on the entire dataset.
Alireza Farhangfar, Russell Greiner, Csaba Szepesvári
ICML2
2009 Learning when to stop thinking and do something!
abstract
An anytime algorithm is capable of returning a response to the given task at essentially any time; typically the quality of the response improves as the time increases. Here, we consider the challenge of learning when we should terminate such algorithms on each of a sequence of iid tasks, to optimize the expected average reward per unit time. We provide a system for addressing this challenge, which combines the global optimizer Cross-Entropy method with local gradient ascent. This paper theoretically investigates how far the estimated gradient is from the true gradient, then empirically demonstrates that this system is effective by applying it to a toy problem, as well as on a real-world face detection task.
Barnabás Póczos, Yasin Abbasi-Yadkori, Csaba Szepesvári, Russell Greiner, Nathan R. Sturtevant
ICML4
2009 Improved Mean and Variance Approximations for Belief Net Responses via Network Doubling
Peter Hooper, Yasin Abbasi-Yadkori, Russell Greiner, Bret Hoehn
UAI3
2009 Predicting homologous signaling pathways using machine learning
abstract
MOTIVATION: In general, each cell signaling pathway involves many proteins, each with one or more specific roles. As they are essential components of cell activity, it is important to understand how these proteins work-and in particular, to determine which of the species' proteins participate in each role. Experimentally determining this mapping of proteins to roles is difficult and time consuming. Fortunately, many pathways are similar across species, so we may be able to use known pathway information of one species to understand the corresponding pathway of another. RESULTS: We present an automatic approach, Predict Signaling Pathway (PSP), which uses the signaling pathways in well-studied species to predict the roles of proteins in less-studied species. We use a machine learning approach to create a predictor that achieves a generalization F-measure of 78.2% when applied to 11 different pathways across 14 different species. We also show our approach is very effective in predicting the pathways that have not yet been experimentally studied completely. AVAILABILITY: The list of predicted proteins for all pathways over all considered species is available at http://www.cs.ualberta.ca/~bioinfo/signaling.
Babak Bostan, Russell Greiner, Duane Szafron, Paul Lu
Bioinform.2
2008 Constrained Classification on Structured Data
Matthew R. G. Brown, Russell Greiner, Albert Murtha
AAAI3
2008 Supervised image segmentation via ground truth decomposition
abstract
This paper proposes a data driven image segmentation algorithm, based on decomposing the target output (ground truth). Classical pixel labeling methods utilize machine learning algorithms that induce a mapping from pixel features to individual pixel labels. In contrast we propose to first extract features from both images and labels. Subsequently we induce a mapping from pixel features to label features and synthesize the final output by combining the newly derived label components. We demonstrate the effectiveness of the proposed approach by applying log-Gabor filters to both input and ground truth images of mineral ore. Subsequently we train perceptrons and regression trees to produce individual output components that are combined in frequency space to create the final segmentation. Experimental results show significant improvements over contextual pixel labeling and over ensemble methods.
Ilya Levner, Russell Greiner, Hong Zhang 0013
ICIP2
2008 Does Wikipedia Information Help Netflix Predictions?
abstract
We explore several ways to estimate movie similarity from the free encyclopedia Wikipedia with the goal of im-proving our predictions for the Netflix Prize. Our system first uses the content and hyperlink structure of Wikipedia articles to identify similarities between movies. We then predict a user’s unknown ratings by using these similarities in conjunction with the user’s known ratings to initialize matrix factorization and k-Nearest Neighbours algorithms. We blend these results with existing ratings-based predic-tors. Finally, we discuss our empirical results, which sug-gest that external Wikipedia data does not significantly im-prove the overall prediction accuracy. 1
John D. Lees-Miller, Fraser Anderson, Bret Hoehn, Russell Greiner
ICMLA4
2008 Using Imputation Techniques to Help Learn Accurate Classifiers
abstract
It is difficult to learn good classifiers when training data is missing attribute values. Conventional techniques for dealing with such omissions, such as mean imputation, generally do not significantly improve the performance of the resulting classifier. We proposed imputation-helped classifiers, which use accurate imputation techniques, such as Bayesian multiple imputation (BMI), predictive mean matching (PMM), and Expectation Maximization (EM), as preprocessors for conventional machine learning algorithms. Our empirical results show that EM-helped and BMI-helped classifiers work effectively when the data is "missing completely at random", generally improving predictive performance over most of the original machine learned classifiers we investigated.
Xiaoyuan Su, Taghi M. Khoshgoftaar, Russell Greiner
ICTAI (1)3
2008 Segmenting Brain Tumors Using Pseudo-Conditional Random Fields
Albert Murtha, Matthew R. G. Brown, Russell Greiner
MICCAI (1)5
2008 Speeding Up Planning in Markov Decision Processes via Automatically Constructed Abstraction
Alejandro Isaza, Csaba Szepesvári, Vadim Bulitko, Russell Greiner
UAI4
2008 Imputed Neighborhood Based Collaborative Filtering
abstract
Collaborative filtering (CF) is one of the most effective types of recommender systems. As data sparsity remains a significant challenge for CF, we consider basing predictions on imputed data, and find this often improves performance on very sparse rating data. In this paper, we propose two imputed neighborhood based collaborative filtering (INCF) algorithms: imputed nearest neighborhood CF (INN-CF) and imputed densest neighborhood CF (IDN-CF), each of which first imputes the user rating data using an imputation technique, before using a traditional Pearson correlation-based CF algorithm on the resulting imputed data of the most similar neighbors or the densest neighbors to make CF predictions for a specific user. We compared an extension of Bayesian multiple imputation (eBMI) and the mean imputation (MEI) in these INCF algorithms, with the commonly-used neighborhood based CF, Pearson correlation-based CF, as well as a densest neighborhood based CF. Our empirical results show that IDN-CF using eBMI significantly outperforms its rivals and takes less time to make its best predictions.
Xiaoyuan Su, Taghi M. Khoshgoftaar, Russell Greiner
Web Intelligence3
2008 Quantifying the uncertainty of a belief net response: Bayesian error-bars for belief net inference
Tim Van Allen, Ajit Singh, Russell Greiner, Peter Hooper
Artif. Intell.3
2008 Improving subcellular localization prediction using text classification and the gene ontology
abstract
MOTIVATION: Each protein performs its functions within some specific locations in a cell. This subcellular location is important for understanding protein function and for facilitating its purification. There are now many computational techniques for predicting location based on sequence analysis and database information from homologs. A few recent techniques use text from biological abstracts: our goal is to improve the prediction accuracy of such text-based techniques. We identify three techniques for improving text-based prediction: a rule for ambiguous abstract removal, a mechanism for using synonyms from the Gene Ontology (GO) and a mechanism for using the GO hierarchy to generalize terms. We show that these three techniques can significantly improve the accuracy of protein subcellular location predictors that use text extracted from PubMed abstracts whose references are recorded in Swiss-Prot.
Alona Fyshe, Yifeng Liu 0001, Duane Szafron, Russell Greiner, Paul Lu
Bioinform.4
2008 Clustering high dimensional data: A graph-based relaxed optimization approach
Osmar R. Zaïane, Ho-Hyun Park, Jiayuan Huang, Russell Greiner
Inf. Sci.5
2007 Mind Change Optimal Learning of Bayes Net Structure
Oliver Schulte, Wei Luo 0001, Russell Greiner
COLT3
2007 Optimistic Active-Learning Using Mutual Information
Yuhong Guo, Russell Greiner
IJCAI2
2007 Hybrid Collaborative Filtering Algorithms Using a Mixture of Experts
abstract
Collaborative filtering (CF) is one of the most successful approaches for recommendation. In this paper, we propose two hybrid CF algorithms, sequential mixture CF and joint mixture CF, each combining advice from multiple experts for effective recommendation. These proposed hybrid CF models work particularly well in the common situation when data are very sparse. By combining multiple experts to form a mixture CF, our systems are able to cope with sparse data to obtain satisfactory performance. Empirical studies show that our algorithms outperform their peers, such as memory-based, pure model-based, pure content-based CF algorithms, and the content- boosted CF (a representative hybrid CF algorithm), especially when the underlying data are very sparse.
Xiaoyuan Su, Russell Greiner, Taghi M. Khoshgoftaar, Xingquan Zhu 0001
Web Intelligence2
2006 Visual Explanation of Evidence with Additive Classifiers
Brett Poulin, Roman Eisner, Duane Szafron, Paul Lu, Russell Greiner, David S. Wishart, Alona Fyshe, Brandon Pearcy, John Anvik
AAAI5
2006 Semi-Supervised Conditional Random Fields for Improved Sequence Segmentation and Labeling
abstract
We present a new semi-supervised training procedure for conditional random fields (CRFs) that can be used to train sequence segmentors and labelers from a combination of labeled and unlabeled training data. Our approach is based on extending the minimum entropy regularization framework to the structured prediction case, yielding a training objective that combines unlabeled conditional entropy with labeled conditional likelihood. Although the training objective is no longer concave, it can still be used to improve an initial model (e.g. obtained from supervised training) by iterative ascent. We apply our new training algorithm to the problem of identifying gene and protein mentions in biological texts, and show that incorporating unlabeled data improves the performance of the supervised CRF in this case.
Feng Jiao, Russell Greiner, Dale Schuurmans
ACL4
2006 Learning to Detect Objects of Many Classes Using Binary Classifiers
Ramana Isukapalli, Ahmed M. Elgammal, Russell Greiner
ECCV (1)3
2006 Using query-specific variance estimates to combine Bayesian classifiers
abstract
Many of today's best classification results are obtained by combining the responses of a set of base classifiers to produce an answer for the query. This paper explores a novel "query specific" combination rule: After learning a set of simple belief network classifiers, we produce an answer to each query by combining their individual responses, using weights based inversely on their respective variances around their responses. These variances are based on the uncertainty of the network parameters, which in turn depend on the training datasample. In essence, this variance quantifies the base classifier's confidence of its response to this query. Our experimental results show that these "mixture-using-variance belief net classifiers" MUVS work effectively, especially when the base classifiers are learned using balanced bootstrap samples and when their results are combined using James-Stein shrinkage. We also found that our variance-based combination rule performed better than both bagging and AdaBoost, even on the set of base classifiers produced by AdaBoost itself. Finally, this framework is extremely efficient, as both the learning and the classification components require only straight-line code.
Russell Greiner
ICML2
2006 Automatic construction of personalized customer interfaces
abstract
Interface personalization can improve a user's performance and subjective impression of interface quality and responsiveness. Personalization is difficult to implement as it requires an accurate model of a user's intentions and a formal model of how an interface meets a user's need. We present a novel model for tractable inference of consumer intentions in the context of grocery shopping. The model makes unique use of a priori temporal relations to simplify inference. We then present a simple interface generation framework that was inspired by viewing user interface interaction as a channel coding problem. The resulting model defines a simplified but clear notion of a user's utility for an interface. We demonstrate the effectiveness of the research prototype on some simple data, and explain how the model can be augmented with richer user modeling to create a deployable application.
Robert Price, Russell Greiner, Gerald Häubl, Alden Flatt
IUI2
2006 Learning to Model Spatial Dependency: Semi-Supervised Discriminative Random Fields
abstract
We present a novel, semi-supervised approach to training discriminative random fields (DRFs) that efficiently exploits labeled and unlabeled training data to achieve improved accuracy in a variety of image processing tasks. We formulate DRF training as a form of MAP estimation that combines conditional loglikelihood on labeled data, given a data-dependent prior, with a conditional entropy regularizer defined on unlabeled data. Although the training objective is no longer concave, we develop an efficient local optimization procedure that produces classifiers that are more accurate than ones based on standard supervised DRF training. We then apply our semi-supervised approach to train DRFs to segment both synthetic and real data sets, and demonstrate significant improvements over supervised DRFs in each case.
Feng Jiao, Dale Schuurmans, Russell Greiner
NIPS5
2006 Information Marginalization on Subgraphs
Jiayuan Huang, Tingshao Zhu, Russell Greiner, Dengyong Zhou, Dale Schuurmans
PKDD3
2006 Efficient Spatial Classification Using Decoupled Conditional Random Fields
Russell Greiner, Osmar R. Zaïane
PKDD2
2006 Finding optimal satisficing strategies for and-or trees
Russell Greiner, Ryan B. Hayward, Magdalena Jankowska, Michael Molloy 0001
Artif. Intell.1
2005 Discriminative Model Selection for Belief Net Structures
Yuhong Guo, Russell Greiner
AAAI2
2005 The Proteome Analyst Suite of Automated Function Prediction Tools
Brett Poulin, Duane Szafron, Paul Lu, Russell Greiner, David S. Wishart, Roman Eisner, Alona Fyshe, Brandon Pearcy, Luca Pireddu
AAAI4
2005 Goal-Directed Site-Independent Recommendations from Passive Observations
Tingshao Zhu, Russell Greiner, Gerald Häubl, Kevin Jewell, Robert Price
AAAI2
2005 Improving Protein Function Prediction Using the Hierarchical Structure of the Gene Ontology
Roman Eisner, Brett Poulin, Duane Szafron, Paul Lu, Russell Greiner
CIBCB5
2005 Learning and Classifying Under Hard Budgets
Aloak Kapoor, Russell Greiner
ECML2
2005 Exploiting syntactic, semantic and lexical regularities in language modeling via directed Markov random fields
abstract
We present a directed Markov random field (MRF) model that combines n-gram models, probabilistic context free grammars (PCFGs) and probabilistic latent semantic analysis (PLSA) for the purpose of statistical language modeling. Even though the composite directed MRF model potentially has an exponential number of loops and becomes a context sensitive grammar, we are nevertheless able to estimate its parameters in cubic time using an efficient modified EM method, the generalized inside-outside algorithm, which extends the inside-outside algorithm to incorporate the effects of the n-gram and PLSA language models. We generalize various smoothing techniques to alleviate the sparseness of n-gram counts in cases where there are hidden variables. We also derive an analogous algorithm to calculate the probability of initial subsequence of a sentence, generated by the composite language model. Our experimental results on the Wall Street Journal corpus show that we obtain significant reductions in perplexity compared to the state-of-the-art baseline trigram model with Good-Turing and Kneser-Ney smoothings.
Shaomin Wang, Russell Greiner, Dale Schuurmans, Li Cheng 0001
ICML3
2005 Segmenting brain tumors using alignment-based features
abstract
Detecting and segmenting brain tumors in magnetic resonance images (MRI) is an important but time-consuming task performed by medical experts. Automating this process is a challenging task due to the often high degree of intensity and textural similarity between normal areas and tumor areas. Several recent projects have explored ways to use an aligned spatial 'template' image to incorporate spatial anatomic information about the brain, but it is not obvious what types of aligned information should be used. This work quantitatively evaluates the performance of 4 different types of alignment-based (AB) features encoding spatial anatomic information for use in supervised pixel classification. This is the first work to (1) compare several types of AB features, (2) explore ways to combine different types of AB features, and (3) explore combining AB features with textural features in a learning framework. We considered situations where existing methods perform poorly, and found that combining textural and AB features allows a substantial performance increase, achieving segmentations that very closely resemble expert annotations.
Mark Schmidt 0001, Ilya Levner, Russell Greiner, Albert Murtha, Aalo Bistritz
ICMLA3
2005 Learning Coordination Classifiers
Yuhong Guo, Russell Greiner, Dale Schuurmans
IJCAI2
2005 Using Learned Browsing Behavior Models to Recommend Relevant Web Pages
Tingshao Zhu, Russell Greiner, Gerald Häubl, Kevin Jewell, Robert Price
IJCAI2
2005 Support Vector Random Fields for Spatial Classification
Russell Greiner, Mark Schmidt 0001
PKDD2
2005 Structural Extension to Logistic Regression: Discriminative Parameter Learning of Belief Net Classifiers
Russell Greiner, Xiaoyuan Su
Mach. Learn.1
2004 The Budgeted Multi-armed Bandit Problem
Omid Madani, Daniel J. Lizotte, Russell Greiner
COLT3
2004 Batch Reinforcement Learning with State Importance
Lihong Li 0001, Vadim Bulitko, Russell Greiner
ECML3
2004 Active Model Selection
Omid Madani, Daniel J. Lizotte, Russell Greiner
UAI3
2004 Predicting subcellular localization of proteins using machine-learned classifiers
abstract
MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.
Zhiyong Lu, Duane Szafron, Russell Greiner, Paul Lu, David S. Wishart, Brett Poulin, John Anvik, Roman Eisner
Bioinform.3
2003 Discriminative Parameter Learning of General Bayesian Network Classifiers
abstract
Greiner and Zhou (1988) presented ELR, a discriminative parameter-learning algorithm that maximizes conditional likelihood (CL) for a fixed Bayesian belief network (BN) structure, and demonstrated that it often produces classifiers that are more accurate than the ones produced using the generative approach (OFE), which finds maximal likelihood parameters. This is especially true when learning parameters for incorrect structures, such as naive Bayes (NB). In searching for algorithms to learn better BN classifiers, this paper uses ELR to learn parameters of more nearly correct BN structures - e.g., of a general Bayesian network (GBN) learned from a structure-learning algorithm by Greiner and Zhou (2002). While OFE typically produces more accurate classifiers with GBN (vs. NB), we show that ELR does not, when the training data is not sufficient for the GBN structure learner to produce a good model. Our empirical studies also suggest that the better the BN structure is, the less advantages ELR has over OFE, for classification purposes. ELR learning on NB (i.e., with little structural knowledge) still performs about the same as OFE on GBN in classification accuracy, over a large number of standard benchmark datasets.
Xiaoyuan Su, Russell Greiner, Petr Musilek, Corrine Cheng
ICTAI3
2003 Lookahead Pathologies for Single Agent Search
Vadim Bulitko, Lihong Li 0001, Russell Greiner, Ilya Levner
IJCAI3
2003 Use of Off-line Dynamic Programming for Efficient Image Interpretation
Ramana Isukapalli, Russell Greiner
IJCAI2
2003 Budgeted Learning of Naive-Bayes Classifiers
Daniel J. Lizotte, Omid Madani, Russell Greiner
UAI3
2002 Learning Bayesian networks from data: An information-theory based approach
Russell Greiner, Jonathan Kelly, David A. Bell, Weiru Liu
Artif. Intell.2
2002 Learning cost-sensitive active classifiers
Russell Greiner, Adam J. Grove, Dan Roth 0001
Artif. Intell.1
2001 Efficient Car Recognition Policies
abstract
This paper addresses the challenges of producing recognition systems that consider both of these objectives. In general, a "(recognition) policy" specifies when to apply which "imaging operators", which can range from low-level edge-detectors and region-growers through high-level token-combination-rules and expectation-driven object-detectors. Given the costs of these operators and the distribution of possible images, we can determine both the expected cost and expected accuracy of any such policy. Our task is to find a maximally effective policy - typically one with sufficient accuracy, whose cost is minimal. We compare various ways to produce such policies in general, and show that policies that select the operators that maximize information gain per unit cost work effectively.
Ramana Isukapalli, Russell Greiner
ICRA2
2001 Efficient Interpretation Policies
Ramana Isukapalli, Russell Greiner
IJCAI2
2001 Bayesian Error-Bars for Belief Net Inference
Tim Van Allen, Russell Greiner, Peter Hooper
UAI2
2000 Model Selection Criteria for Learning Belief Nets: An Empirical Comparison
Tim Van Allen, Russell Greiner
ICML2
1999 Comparing Bayesian Network Classifiers
Russell Greiner
UAI2
1997 Why Experimentation can be better than "Perfect Guidance"
Tobias Scheffer, Russell Greiner, Christian J. Darken
ICML2
1997 Learning Bayesian Nets that Perform Well
Russell Greiner, Adam J. Grove, Dale Schuurmans
UAI1
1997 Knowing what doesn't Matter: Exploiting the Omission of Irrelevant Data
Russell Greiner, Adam J. Grove, Alexander Kogan
Artif. Intell.1
1997 The Relevance of Relevance (Editorial)
Devika Subramanian, Russell Greiner, Judea Pearl
Artif. Intell.2
1996 Exploiting the Omission of Irrelevant Data
Russell Greiner, Adam J. Grove, Alexander Kogan
ICML1
1996 Learning Active Classifiers
Russell Greiner, Adam J. Grove, Dan Roth 0001
ICML1
1996 PALO: A Probabilistic Hill-Climbing Algorithm
Russell Greiner
Artif. Intell.1
1996 Probably Approximately Optimal Satisficing Strategies
Russell Greiner, Pekka Orponen
Artif. Intell.1
1996 Learning to select useful landmarks
abstract
To navigate effectively, an autonomous agent must be able to quickly and accurately determine its current location. Given an initial estimate of its position (perhaps based on dead-reckoning) and an image taken of a known environment, our agent first attempts to locate a set of landmarks (real-world objects at known locations), then uses their angular separation to obtain an improved estimate of its current position. Unfortunately, some landmarks may not be visible, or worse, may be confused with other landmarks, resulting in both time wasted in searching for the undetected landmarks, and in further errors in the agent's estimate of its position. To address these problems, we propose a method that uses previous experiences to learn a selection function that, given the set of landmarks that might be visible, returns the subset that can be used to reliably provide an accurate registration of the agent's position. We use statistical techniques to prove that the learned selection function is, with high probability, effectively at a local optimum in the space of such functions. This paper also presents empirical evidence, using real-world data, that demonstrate the effectiveness of our approach.
Russell Greiner, Ramana Isukapalli
IEEE Trans. Syst. Man Cybern. Part B1
1995 Sequential PAC Learning
abstract
We consider the use of "on-line" stopping rules to reduce the number of training examples needed to pac-learn. Rather than collect a large training sample that can be proved sufficient to eliminate all bad hypotheses a priori, the idea is instead to observe training examples one-at-a-time and decide "on-line" whether to stop and return a hypothesis, or continue training. The primary benefit of this approach is that we can detect when a hypothesizer has actually "converged," and halt training before the standard fixed-sample-size bounds. This paper presents a series of such sequential learning procedures for: distribution-free pac-learning, "mistake-bounded to pac" conversion, and distribution-specific pac-learning, respectively. We analyze the worst case expected training sample size of these procedures, and show that this is often smaller than existing fixed sample size bounds --- while providing the exact same worst case pac-guarantees. We also provide lower bounds that show these r...
Dale Schuurmans, Russell Greiner
COLT2
1995 The Challenge of Revising an Impure Theory
Russell Greiner
ICML1
1995 The Complexity of Theory Revision
Russell Greiner
IJCAI1
1995 Practical PAC Learning
Dale Schuurmans, Russell Greiner
IJCAI2
1994 Learning to Select Useful Landmarks
Russell Greiner, Ramana Isukapalli
AAAI1
1993 D. B. Lenat and R. V. Guha, Building Large Knowledge-Based Systems: Representation and Inference in the Cyc Project
Charles Elkan, Russell Greiner
Artif. Intell.2
1992 A Statistical Approach to Solving the EBL Utility Problem
Russell Greiner, Igor Jurisica
AAAI1
1992 Learning Useful Horn Approximations
Russell Greiner, Dale Schuurmans
KR1
1992 Learning Efficient Query Processing Strategies
abstract
A query processor qp uses the rules in a rule base to reduce a given query to a series of attempted retrievals from a database of facts. The qp's expected cost is the average time it requires to find an answer, averaged over its anticipated set of queries. This cost depends on the qp's strategy, which specifies the order in which it considers the possible rules and retrievals. This paper provides two related learning algorithms, pib and pao, for improving the qp's strategy, i.e., for producing new strategies with lower expected costs. Each algorithm first monitors the qp's operations over a set of queries, observing how often each path of rules leads to a sufficient set of successful retrievals, and then uses these statistics to suggest a new strategy. pib hill-climbs to strategies that are, with high probability, successively better; and pao produces a new strategy that probably is approximately optimal. We describe how to implement both learning systems unobtrusively, discuss thei...
Russell Greiner
PODS1
1991 Measuring and Improving the Effectiveness of Representations
Russell Greiner, Charles Elkan
IJCAI1
1991 Probably Approximately Optimal Derivation Strategies
Russell Greiner, Pekka Orponen
KR1
1991 Finding Optimal Derivation Strategies in Redundant Knowledge Bases
Russell Greiner
Artif. Intell.1
1989 Towards a Formal Analysis of EBL
Russell Greiner
ML1
1989 Incorporating Redundant Learned Rules: A Preliminary Formal Analysis of EBL
Russell Greiner, J. Likuski
IJCAI1
1989 A Correction to the Algorithm in Reiter's Theory of Diagnosis
Russell Greiner, Barbara A. Smith, Ralph W. Wilkerson
Artif. Intell.1
1988 Signal abstractions in the machine analysis of radar signals for ice profiling
abstract
Describes the design and implementation of an automated system for interpreting impulse radar signals for ice thickness profiling. The authors have adopted an integrated approach which includes numeric computation in the form of deconvolution filtering with rule-based classification of signal features at multiple levels. Noise reduction and deconvolution techniques are used to enhance the radar signals for better resolution of overlapping events. Motivated by human perceptual (visual) knowledge, a hierarchy of data structures is constructed as representations of signal characteristics at various levels of abstraction. Classification rules, based on the protocols collected from an expert, physical constraints on the helicopter motion and the nature of the radar signals are used to produce the current signal interpretation. A prototype system has been implemented on the Symbolics Lisp Machine and tested on real data.>
Evangelos E. Milios, Russell Greiner, James R. Rossiter
ICASSP3
1988 Learning by Understanding Analogies
Russell Greiner
Artif. Intell.1
1988 Against the unjustified use of probabilities
Russell Greiner
Comput. Intell.1
1988 A Review of Machine Learning at AAAI-87
Russell Greiner
Mach. Learn.1
1983 What's New? A Semantic Definition of Novelty
Russell Greiner, Michael R. Genesereth
IJCAI1
1980 A Representation Language Language
Russell Greiner, Douglas B. Lenat
AAAI1