Vasant G. Honavar

dblp:h/VasantHonavar · DBLP profile ↗
← Back
153ranked-venue papers
5as first author
25since 2021 · last 2025
0000-0001-5399-3489ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 86 · 4 first-author · 15 since 2021Databases, data management, data science and information retrieval · 38 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 12Security and privacy · 4 · 2 since 2021Theory of computation · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Checking Consistency of CP-Theory Preferences in Polynomial Time
abstract
We investigate the problem of checking the consistency of qualitative preferences expressed in CP-theory. This problem is PSPACE-Complete even when the preferences are locally consistent or the preference variables have binary domain. We present a new sufficient condition for consistency of preferences and show that the condition can be checked in polynomial time in settings of practical relevance (locally consistent or binary domain preference variables). We further show how the resulting sufficient condition can be used to efficiently identify a subset of outcomes that are non-dominated with respect to a set of qualitative preferences.
Erik Rauer, Samik Basu 0001, Vasant G. Honavar
AAAI3
2025 Reinforcement Learning for Large Language Models via Group Preference Reward Shaping
abstract
Huaisheng Zhu, Siyuan Xu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou, Shuyue Hu, Vasant G. Honavar. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Huaisheng Zhu, Hangfan Zhang, Teng Xiao, Zhimeng Guo, Shijie Zhou 0008, Shuyue Hu, Vasant G. Honavar
EMNLP8
2025 SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters
abstract
Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasing both the complexity and time required for fine-tuning large language models. In this paper, we propose a simple yet effective hyperparameter-free preference optimization algorithm for alignment. We observe that promising performance can be achieved simply by optimizing inverse perplexity, which is calculated as the inverse of the exponentiated average log-likelihood of the chosen and rejected responses in the preference dataset. The resulting simple learning objective, SimPER, is easy to implement and eliminates the need for expensive hyperparameter tuning and a reference model, making it both computationally and memory efficient. Extensive experiments on widely used real-world benchmarks, including MT-Bench, AlpacaEval 2, and 10 key benchmarks of the Open LLM Leaderboard with 5 base models, demonstrate that SimPER consistently and significantly outperforms existing approaches—even without any hyperparameters or a reference model. For example, despite its simplicity, SimPER outperforms state-of-the-art methods by up to 5.7 points on AlpacaEval 2 and achieves the highest average ranking across 10 benchmarks on the Open LLM Leaderboard. The source code for SimPER is publicly available at: https://github.com/tengxiao1/SimPER.
Teng Xiao, Yige Yuan, Zhengyu Chen 0001, Mingxiao Li 0004, Shangsong Liang, Zhaochun Ren, Vasant G. Honavar
ICLR7
2025 On a Connection Between Imitation Learning and RLHF
abstract
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcement learning from human feedback RLHF and imitation learning (IL), revealing that RLHF implicitly performs imitation learning on the preference data distribution. Building on this connection, we propose DIL, a principled framework that directly optimizes the imitation learning objective. DIL provides a unified imitation learning perspective on alignment, encompassing existing alignment algorithms as special cases while naturally introducing new variants. By bridging IL and RLHF, DIL offers new insights into alignment with RLHF. Extensive experiments demonstrate that DIL outperforms existing methods on various challenging benchmarks.
Teng Xiao, Yige Yuan, Mingxiao Li 0004, Zhengyu Chen 0001, Vasant G. Honavar
ICLR5
2025 DSPO: Direct Score Preference Optimization for Diffusion Model Alignment
abstract
Diffusion-based Text-to-Image (T2I) models have achieved impressive success in generating high-quality images from textual prompts. While large language models (LLMs) effectively leverage Direct Preference Optimization (DPO) for fine-tuning on human preference data without the need for reward models, diffusion models have not been extensively explored in this area. Current preference learning methods applied to T2I diffusion models immediately adapt existing techniques from LLMs. However, this direct adaptation introduces an estimated loss specific to T2I diffusion models. This estimation can potentially lead to suboptimal performance through our empirical results. In this work, we propose Direct Score Preference Optimization (DSPO), a novel algorithm that aligns the pretraining and fine-tuning objectives of diffusion models by leveraging score matching, the same objective used during pretraining. It introduces a new perspective on preference learning for diffusion models. Specifically, DSPO distills the score function of human-preferred image distributions into pretrained diffusion models, fine-tuning the model to generate outputs that align with human preferences. We theoretically show that DSPO shares the same optimization direction as reinforcement learning algorithms in diffusion models under certain conditions. Our experimental results demonstrate that DSPO outperforms preference learning baselines for T2I diffusion models in human preference evaluation tasks and enhances both visual appeal and prompt alignment of generated images.
Huaisheng Zhu, Teng Xiao, Vasant G. Honavar
ICLR3
2025 Simple Distillation for One-Step Diffusion Models
abstract
Diffusion models have established themselves as leading techniques for image generation. However, their reliance on an iterative denoising process results in slow sampling speeds, which limits their applicability to interactive and creative applications. An approach to overcoming this limitation involves distilling multistep diffusion models into efficient one-step generators. However, existing distillation methods typically suffer performance degradation or require complex iterative training procedures which increase their complexity and computational cost. In this paper, we propose Contrastive Energy Distillation (CED), a simple yet effective approach to distill multistep diffusion models into effective one-step generators. Our key innovation is the introduction of an unnormalized joint energy-based model (EBM) that represents the generator and an auxiliary score model. CED optimizes a Noise Contrastive Estimation (NCE) objective to efficiently transfers knowledge from a multistep teacher diffusion model without additional modules or iterative training complexity. We further show that CED implicitly optimizes the KL divergence between the distributions modeled by the multistep diffusion model and the one-step generator. We present results of experiments which demonstrate that CED achieves competitive performance with the representative baselines for distilling multistep diffusion models while maintaining excellent memory efficiency.
Huaisheng Zhu, Teng Xiao, Shijie Zhou 0008, Zhimeng Guo, Hangfan Zhang, Vasant G. Honavar
NeurIPS7
2025 Hyperdimensional Representation Learning for Node Classification and Link Prediction
abstract
We introduce Hyperdimensional Graph Learner (HDGL), a novel method for node classification and link prediction in graphs. HDGL maps node features into a very high-dimensional space (hyperdimensional or HD space for short) using the injectivity property of node representations in a family of Graph Neural Networks (GNNs) and then uses HD operators such as bundling and binding to aggregate information from the local neighborhood of each node yielding latent node representations that can support both node classification and link prediction tasks. HDGL, unlike GNNs that rely on computationally expensive iterative optimization and hyperparameter tuning, requires only a single pass through the data set. We report results of experiments using widely used benchmark datasets which demonstrate that, on the node classification task, HDGL achieves accuracy that is competitive with that of the state-of-the-art GNN methods at substantially reduced computational cost; and on the link prediction task, HDGL matches the performance of DeepWalk and related methods, although it falls short of computationally demanding state-of-the-art GNNs.
Abhishek Dalvi, Vasant G. Honavar
WSDM2
2025 DGX: Uncovering General Behavior of Deep Graph Models With Model-Level Explanation
abstract
Deep graph learning models have recently been developed to learn from various graphs that are prevalent in describing and modeling complex systems, including those in bioinformatics. However, a versatile explanation method for uncovering the general graph patterns that guide deep graph models in making predictions remains elusive. In this paper, we propose DGX, a novel deep graph model explainer that generates explanatory graphs to explain trained, opaque-box deep graph models. Its effectiveness is demonstrated by producing multiple graphs that collectively encode the structural knowledge captured by the graph neural network on both synthetic and real graph data. Importantly, DGX can produce diverse explanations by generating a set of distinguishable graphs and can provide customized explanations based on prior knowledge or constraints specified by users. We apply DGX to explain a mutagenicity prediction model by exploring the underlying groups of mutagenic compounds, and we explain the model on brain functional networks by revealing the structural patterns that enable the model to differentiate autism spectrum disorder from healthy controls. These findings offer an effective, diverse, and customized approach to explaining the underlying mechanisms and enhancing the understanding of models learned from real graph data, particularly in fields such as biomedicine and bioinformatics.
Jinlong Hu 0002, Shoubin Dong, Bin Liao 0005, Vasant G. Honavar
IEEE Trans. Comput. Biol. Bioinform.6
2024 Inducing Clusters Deep Kernel Gaussian Process for Longitudinal Data
abstract
We consider the problem of predictive modeling from irregularly and sparsely sampled longitudinal data with unknown, complex correlation structures and abrupt discontinuities. To address these challenges, we introduce a novel inducing clusters longitudinal deep kernel Gaussian Process (ICDKGP). ICDKGP approximates the data generating process by a zero-mean GP with a longitudinal deep kernel that models the unknown complex correlation structure in the data and a deterministic non-zero mean function to model the abrupt discontinuities. To improve the scalability and interpretability of ICDKGP, we introduce inducing clusters corresponding to centers of clusters in the training data. We formulate the training of ICDKGP as a constrained optimization problem and derive its evidence lower bound. We introduce a novel relaxation of the resulting problem which under rather mild assumptions yields a solution with error bounded relative to the original problem. We describe the results of extensive experiments demonstrating that ICDKGP substantially outperforms the state-of-the-art longitudinal methods on data with both smoothly and non-smoothly varying outcomes.
Weijieying Ren, Hanifi Sahar, Vasant G. Honavar
AAAI4
2024 GeomCLIP: Contrastive Geometry-Text Pre-training for Molecules
abstract
Pretraining molecular representations is crucial for drug and material discovery. Recent methods focus on learning representations from geometric structures, effectively capturing 3D position information. Yet, they overlook the rich information in biomedical texts, which detail molecules’ properties and substructures. With this in mind, we set up a data collection effort for 200K pairs of ground-state geometric structures and biomedical texts, resulting in a PubChem3D dataset. Based on this dataset, we propose the GeomCLIP framework to enhance geometric pretraining and understanding by biomedical texts. During pre-training, we design two types of tasks, i.e., multimodal representation alignment and unimodal denoising pretraining, to align the 3D geometric encoder with textual information and, at the same time, preserve its original representation power. Experimental results show the effectiveness of GeomCLIP in various tasks such as molecule property prediction, zero-shot text-molecule retrieval, and 3D molecule captioning. Our code and collected dataset are available at https://github.com/xiaocui3737/GeomCLIP.
Teng Xiao, Chao Cui, Huaisheng Zhu, Vasant G. Honavar
BIBM4
2024 How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
abstract
This paper introduces a novel generalized selfimitation learning (GSIL) framework, which effectively and efficiently aligns large language models with offline demonstration data.We develop GSIL by deriving a surrogate objective of imitation learning with density ratio estimates, facilitating the use of self-generated data and optimizing the imitation learning objective with simple classification losses.GSIL eliminates the need for complex adversarial training in standard imitation learning, achieving lightweight and efficient fine-tuning for large language models.In addition, GSIL encompasses a family of offline losses parameterized by a general class of convex functions for density ratio estimation and enables a unified view for alignment with demonstration data.Extensive experiments show that GSIL consistently and significantly outperforms baselines in many challenging benchmarks, such as coding (HuamnEval), mathematical reasoning (GSM8K) and instruction-following benchmark (MT-Bench).Code is public available at https://github.com/tengxiao1/GSIL.
Teng Xiao, Mingxiao Li 0004, Yige Yuan, Huaisheng Zhu, Chao Cui, Vasant G. Honavar
EMNLP6
2024 TabLog: Test-Time Adaptation for Tabular Data Using Logic Rules
abstract
We consider the problem of test-time adaptation of predictive models trained on tabular data. Effective solution of this problem requires adaptation of predictive models trained on the source domain to a target domain, using only unlabeled target domain data, without access to source domain data. Existing test-time adaptation methods for tabular data have difficulty coping with the heterogeneous features and their complex dependencies inherent in tabular data. To overcome these limitations, we consider test-time adaptation in the setting wherein the logical structure of the rules is assumed to remain invariant despite distribution shift between source and target domains whereas the numerical parameters associated with the rules and the weights assigned to them can vary to accommodate distribution shift. TabLog discretizes numerical features, models dependencies between heterogeneous features, introduces a novel contrastive loss for coping with distribution shift, and presents an end-to-end framework for efficient training and test-time adaptation by taking advantage of a logical neural network representation of a rule ensemble. We present results of experiments using several benchmark data sets that demonstrate TabLog is competitive with or improves upon the state-of-the-art methods for test-time adaptation of predictive models trained on tabular data. Our code is available at https://github.com/WeijieyingRen/TabLog.
Weijieying Ren, Xiaoting Li 0001, Huiyuan Chen, Vineeth Rakesh, Zhuoyi Wang, Mahashweta Das, Vasant G. Honavar
ICML7
2024 Efficient Contrastive Learning for Fast and Accurate Inference on Graphs
abstract
Graph contrastive learning has made remarkable advances in settings where there is a scarcity of task-specific labels. Despite these advances, the significant computational overhead for representation inference incurred by existing methods that rely on intensive message passing makes them unsuitable for latency-constrained applications. In this paper, we present GraphECL, a simple and efficient contrastive learning method for fast inference on graphs. GraphECL does away with the need for expensive message passing during inference. Specifically, it introduces a novel coupling of the MLP and GNN models, where the former learns to computationally efficiently mimic the computations performed by the latter. We provide a theoretical analysis showing why MLP can capture essential structural information in neighbors well enough to match the performance of GNN in downstream tasks. The extensive experiments on widely used real-world benchmarks that show that GraphECL achieves superior performance and inference efficiency compared to state-of-the-art graph constrastive learning (GCL) methods on homophilous and heterophilous graphs.
Teng Xiao, Huaisheng Zhu, Zhiwei Zhang 0028, Zhimeng Guo, Charu C. Aggarwal, Suhang Wang, Vasant G. Honavar
ICML7
2024 Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
abstract
We study the problem of aligning large language models (LLMs) with human preference data. Contrastive preference optimization has shown promising results in aligning LLMs with available preference data by optimizing the implicit reward associated with the policy. However, the contrastive objective focuses mainly on the relative values of implicit rewards associated with two responses while ignoring their actual values, resulting in suboptimal alignment with human preferences. To address this limitation, we propose calibrated direct preference optimization (Cal-DPO), a simple yet effective algorithm. We show that substantial improvement in alignment with the given preferences can be achieved simply by calibrating the implicit reward to ensure that the learned implicit rewards are comparable in scale to the ground-truth rewards. We demonstrate the theoretical advantages of Cal-DPO over existing approaches. The results of our experiments on a variety of standard benchmarks show that Cal-DPO remarkably improves off-the-shelf methods.
Teng Xiao, Yige Yuan, Huaisheng Zhu, Mingxiao Li 0004, Vasant G. Honavar
NeurIPS5
2024 EsaCL: An Efficient Continual Learning Algorithm
abstract
A key challenge in the continual learning setting is to efficiently learn a sequence of tasks without forgetting how to perform previously learned tasks. Many existing approaches to this problem work by either retraining the model on previous tasks or by expanding the model to accommodate new tasks. However, these approaches typically suffer from increased storage and computational requirements, a problem that is worsened in the case of sparse models due to need for expensive re-training after sparsification. To address this challenge, we propose a new method for efficient continual learning of sparse models (EsaCL) that can automatically prune redundant parameters without adversely impacting the model's predictive power, and circumvent the need of retraining. We conduct a theoretical analysis of loss landscapes with parameter pruning, and design a directional pruning (SDP) strategy that is informed by the sharpness of the loss function with respect to the model parameters. SDP ensures model with minimal loss of predictive accuracy, accelerating the learning of sparse models at each stage. To accelerate model update, we introduce an intelligent data selection (IDS) strategy that can identify critical instances for estimating loss landscape, yielding substantially improved data efficiency. The results of our experiments show that EsaCL achieves performance that is competitive with the state-of-the-art methods.
Weijieying Ren, Vasant G. Honavar
SDM2
2023 Representing and Reasoning with Multi-Stakeholder Qualitative Preference Queries
abstract
Many decision-making scenarios, e.g., public policy, healthcare, business, and disaster response, require accommodating the preferences of multiple stakeholders. We offer the first formal treatment of reasoning with multi-stakeholder qualitative preferences in a setting where stakeholders express their preferences in a qualitative preference language, e.g., CP-net, CI-net, TCP-net, CP-Theory. We introduce a query language for expressing queries against such preferences over sets of outcomes that satisfy specified criteria, e.g., ψ1PAψ2 (read loosely as the set of outcomes satisfying ψ1 that are preferred over outcomes satisfying ψ2 by a set of stakeholders A). Motivated by practical application scenarios, we introduce and analyze several alternative semantics for such queries, and examine their interrelationships. We provide a provably correct algorithm for answering multi-stakeholder qualitative preference queries using model checking in alternation-free μ-calculus. We present experimental results that demonstrate the feasibility of our approach.
Samik Basu 0001, Vasant G. Honavar, Ganesh Ram Santhanam, Jia Tao 0001
ECAI2
2023 License Forecasting and Scheduling for HPC
abstract
This work focuses on forecasting future license usage for high-performance computing environments and using such predictions to improve the effectiveness of job scheduling. Specifically, we propose a model that carries out both short-term and long-term license usage forecasting and a method of using forecasts to improve job scheduling. Our long-term forecasting model achieves a Mean Absolute Percentage Error (MAPE) as low as 0.26 for a 12-month forecast of daily peak license usage. Our job scheduling experimental results also indicate that wasted work from jobs with insufficient licenses can be reduced by up to 92% without increasing the average license-using job completion times, during periods of high license usage, with our proposed license-aware scheduler.
Ahmed Burak Gulhan, Gulsum Gudukbay Akbulut, Amit Amritkar, Jack Sampson, Vasant G. Honavar, Adam Focht, Chuck Pavloski, Mahmut T. Kandemir
MASCOTS5
2023 Forecasting User Interests Through Topic Tag Predictions in Online Health Communities
abstract
The increasing reliance on online communities for healthcare information by patients and caregivers has led to the increase in the spread of misinformation, or subjective, anecdotal and inaccurate or non-specific recommendations, which, if acted on, could cause serious harm to the patients. Hence, there is an urgent need to connect users with accurate and tailored health information in a timely manner to prevent such harm. This article proposes an innovative approach to suggesting reliable information to participants in online communities as they move through different stages in their disease or treatment. We hypothesize that patients with similar histories of disease progression or course of treatment would have similar information needs at comparable stages. Specifically, we pose the problem of predicting topic tags or keywords that describe the future information needs of users based on their profiles, traces of their online interactions within the community (past posts, replies) and the profiles and traces of online interactions of other users with similar profiles and similar traces of past interaction with the target users. The result is a variant of the collaborative information filtering or recommendation system tailored to the needs of users of online health communities. We report results of our experiments on two unique datasets from two different social media platforms which demonstrates the superiority of the proposed approach over the state of the art baselines with respect to accurate and timely prediction of topic tags (and hence information sources of interest).
Amogh Subbakrishna Adishesha, Lily Jakielaszek, Fariha Azhar, Peixuan Zhang, Vasant G. Honavar, Fenglong Ma, Chandra Belani, Prasenjit Mitra 0001, Sharon X. Huang
IEEE J. Biomed. Health Informatics5
2022 Detecting and Interpreting Changes in Scanning Behavior in Large Network Telescopes
abstract
Network telescopes or “Darknets” received unsolicited Internet-wide traffic, thus providing a unique window into macroscopic Internet activities associated with malware propagation, denial of service attacks, network reconnaissance, misconfigurations and network outages. Analysis of the resulting data can provide actionable insights to security analysts that can be used to prevent or mitigate cyber-threats. Large network telescopes, however, observe millions of nefarious scanning activities on a daily basis which makes the transformation of the captured information into meaningful threat intelligence challenging. To address this challenge, we present a novel framework for characterizing the structure and temporal evolution of scanning behaviors observed in network telescopes. The proposed framework includes four components. It (i) extracts a rich, high-dimensional representation ofscanning profilescomposed of features distilled from network telescope data; (ii) learns, in an unsupervised fashion, information-preservingsuccinct representationsof these scanning behaviors usingdeep representation learningthat is amenable to clustering; (iii) performsclusteringof the scanner profiles in the resulting latent representation space on daily Darknet data, and (iv)detects temporal changesin scanning behavior using techniques fromoptimal mass transport. We robustly evaluate the proposed system using both synthetic data and real-world Darknet data. We demonstrate its ability to detect real-world, high-impact cybersecurity incidents such as the onset of the Mirai botnet in late 2016 and several interesting cluster formations in early 2022 (e.g., heavy scanners, evolved Mirai variants, Darknet “backscatter” activities, etc.). Comparisons with state-of-the-art methods showcase that the integration of the proposed features with the deep representation learning scheme leads to better classification performance of Darknet scanners.
Michael G. Kallitsis, Rupesh Prajapati, Vasant G. Honavar, Dinghao Wu, John Yen
IEEE Trans. Inf. Forensics Secur.3
2021 Longitudinal Deep Kernel Gaussian Process Regression
abstract
Gaussian processes offer an attractive framework for predictive modeling from longitudinal data, \ie irregularly sampled, sparse observations from a set of individuals over time. However, such methods have two key shortcomings: (i) They rely on ad hoc heuristics or expensive trial and error to choose the effective kernels, and (ii) They fail to handle multilevel correlation structure in the data. We introduce Longitudinal deep kernel Gaussian process regression (L-DKGPR) to overcome these limitations by fully automating the discovery of complex multilevel correlation structure from longitudinal data. Specifically, L-DKGPR eliminates the need for ad hoc heuristics or trial and error using a novel adaptation of deep kernel learning that combines the expressive power of deep neural networks with the flexibility of non-parametric kernel methods. L-DKGPR effectively learns the multilevel correlation with a novel additive kernel that simultaneously accommodates both time-varying and the time-invariant effects. We derive an efficient algorithm to train L-DKGPR using latent space inducing points and variational inference. Results of extensive experiments on several benchmark data sets demonstrate that L-DKGPR significantly outperforms the state-of-the-art longitudinal data analysis (LDA) methods.
Yanting Wu, Dongkuan Xu, Vasant G. Honavar
AAAI4
2021 Shedding light into the darknet: scanning characterization and detection of temporal changes
abstract
Network telescopes provide a unique window into Internet-wide malicious activities associated with malware propagation, denial of service attacks, network reconnaissance, and others. Analyses of this telescope data can highlight ongoing malicious events in the Internet which can be used to prevent or mitigate cyber-threats in real-time. However, large telescopes observe millions of events on a daily basis which renders the task of transforming this knowledge to meaningful insights challenging. In order to address this, we present a novel framework for characterizing Internet's background radiation and for tracking its temporal evolution. The proposed framework: (i) Extracts a high dimensional representation of telescope scanners composed of features distilled from telescope data and learns an information-preserving low-dimensional representation of these events that is amenable to clustering; (ii) Performs clustering of resulting representation space to characterize the scanners and (iii) Utilizes the clustering outcomes as "signatures" to detect temporal changes in the network telescope.
Rupesh Prajapati, Vasant G. Honavar, Dinghao Wu, John Yen, Michael G. Kallitsis
CoNEXT2
2021 FARE: Enabling Fine-grained Attack Categorization under Low-quality Labeled Data
Wenbo Guo 0002, Tongbo Luo, Vasant G. Honavar, Gang Wang 0011, Xinyu Xing 0001
NDSS4
2021 Functional Autoencoders for Functional Data Representation Learning
Tsung-Yu Hsieh, Suhang Wang, Vasant G. Honavar
SDM4
2021 Explainable Multivariate Time Series Classification: A Deep Neural Network Which Learns to Attend to Important Variables As Well As Time Intervals
abstract
Many real-world applications, e.g., healthcare, present multi-variate time series prediction problems. In such settings, in addition to the predictive accuracy of the models, model transparency and explainability are paramount. We consider the problem of building explainable classifiers from multi-variate time series data. A key criterion to understand such predictive models involves elucidating and quantifying the contribution of time varying input variables to the classification. Hence, we introduce a novel, modular, convolution-based feature extraction and attention mechanism that simultaneously identifies the variables as well as time intervals which determine the classifier output. We present results of extensive experiments with several benchmark data sets that show that the proposed method outperforms the state-of-the-art baseline methods on multi-variate time series classification task. The results of our case studies demonstrate that the variables and time intervals identified by the proposed method make sense relative to available domain knowledge.
Tsung-Yu Hsieh, Suhang Wang, Vasant G. Honavar
WSDM4
2021 SrVARM: State Regularized Vector Autoregressive Model for Joint Learning of Hidden State Transitions and State-Dependent Inter-Variable Dependencies from Multi-variate Time Series
abstract
Many applications, e.g., healthcare, education, call for effective methods methods for constructing predictive models from high dimensional time series data where the relationship between variables can be complex and vary over time. In such settings, the underlying system undergoes a sequence of unobserved transitions among a finite set of hidden states. Furthermore, the relationships between the observed variables and their temporal dynamics may depend on the hidden state of the system. To further complicate matters, the hidden state sequences underlying the observed data from different individuals may not be aligned relative to a common frame of reference. Against this background, we consider the novel problem of jointly learning the state-dependent inter-variable relationships as well as the pattern of transitions between hidden states from multi-variate time series data. To solve this problem, we introduce the State-Regularized Vector Autoregressive Model (SrVARM) which combines a state-regularized recurrent neural network to learn the dynamics of transitions between discrete hidden states with an augmented autoregressive model which models the inter-variable dependencies in each state using a state-dependent directed acyclic graph (DAG). We propose an efficient algorithm for training SrVARM by leveraging a recently introduced reformulation of the combinatorial problem of optimizing the DAG structure with respect to a scoring function into a continuous optimization problem. We report results of extensive experiments with simulated data as well as a real-world benchmark that show that SrVARM outperforms state-of-the-art baselines in recovering the unobserved state transitions and discovering the state-dependent relationships among variables.
Tsung-Yu Hsieh, Xianfeng Tang, Suhang Wang, Vasant G. Honavar
WWW5
2020 Algorithmic Bias in Recidivism Prediction: A Causal Perspective (Student Abstract)
abstract
ProPublica's analysis of recidivism predictions produced by Correctional Offender Management Profiling for Alternative Sanctions (COMPAS) software tool for the task, has shown that the predictions were racially biased against African American defendants. We analyze the COMPAS data using a causal reformulation of the underlying algorithmic fairness problem. Specifically, we assess whether COMPAS exhibits racial bias against African American defendants using FACT, a recently introduced causality grounded measure of algorithmic fairness. We use the Neyman-Rubin potential outcomes framework for causal inference from observational data to estimate FACT from COMPAS data. Our analysis offers strong evidence that COMPAS exhibits racial bias against African American defendants. We further show that the FACT estimates from COMPAS data are robust in the presence of unmeasured confounding.
Aria Khademi, Vasant G. Honavar
AAAI2
2020 LMLFM: Longitudinal Multi-Level Factorization Machine
abstract
We consider the problem of learning predictive models from longitudinal data, consisting of irregularly repeated, sparse observations from a set of individuals over time. Such data often exhibit longitudinal correlation (LC) (correlations among observations for each individual over time), cluster correlation (CC) (correlations among individuals that have similar characteristics), or both. These correlations are often accounted for using mixed effects models that include fixed effects and random effects, where the fixed effects capture the regression parameters that are shared by all individuals, whereas random effects capture those parameters that vary across individuals. However, the current state-of-the-art methods are unable to select the most predictive fixed effects and random effects from a large number of variables, while accounting for complex correlation structure in the data and non-linear interactions among the variables. We propose Longitudinal Multi-Level Factorization Machine (LMLFM), to the best of our knowledge, the first model to address these challenges in learning predictive models from longitudinal data. We establish the convergence properties, and analyze the computational complexity, of LMLFM. We present results of experiments with both simulated and real-world longitudinal data which show that LMLFM outperforms the state-of-the-art methods in terms of predictive accuracy, variable selection ability, and scalability to data with large number of variables. The code and supplemental material is available at https://github.com/junjieliang672/LMLFM.
Dongkuan Xu, Vasant G. Honavar
AAAI4
2020 Adversarial Attacks on Graph Neural Networks via Node Injections: A Hierarchical Reinforcement Learning Approach
abstract
Graph Neural Networks (GNN) offer the powerful approach to node classification in complex networks across many domains including social media, E-commerce, and FinTech. However, recent studies show that GNNs are vulnerable to attacks aimed at adversely impacting their node classification performance. Existing studies of adversarial attacks on GNN focus primarily on manipulating the connectivity between existing nodes, a task that requires greater effort on the part of the attacker in real-world applications. In contrast, it is much more expedient on the part of the attacker to inject adversarial nodes, e.g., fake profiles with forged links, into existing graphs so as to reduce the performance of the GNN in classifying existing nodes.
Suhang Wang, Xianfeng Tang, Tsung-Yu Hsieh, Vasant G. Honavar
WWW5
2020 iScore: a novel graph kernel-based function for scoring protein-protein docking models
abstract
MOTIVATION: Protein complexes play critical roles in many aspects of biological functions. Three-dimensional (3D) structures of protein complexes are critical for gaining insights into structural bases of interactions and their roles in the biomolecular pathways that orchestrate key cellular processes. Because of the expense and effort associated with experimental determinations of 3D protein complex structures, computational docking has evolved as a valuable tool to predict 3D structures of biomolecular complexes. Despite recent progress, reliably distinguishing near-native docking conformations from a large number of candidate conformations, the so-called scoring problem, remains a major challenge. RESULTS: Here we present iScore, a novel approach to scoring docked conformations that combines HADDOCK energy terms with a score obtained using a graph representation of the protein-protein interfaces and a measure of evolutionary conservation. It achieves a scoring performance competitive with, or superior to, that of state-of-the-art scoring functions on two independent datasets: (i) Docking software-specific models and (ii) the CAPRI score set generated by a wide variety of docking approaches (i.e. docking software-non-specific). iScore ranks among the top scoring approaches on the CAPRI score set (13 targets) when compared with the 37 scoring groups in CAPRI. The results demonstrate the utility of combining evolutionary, topological and energetic information for scoring docked conformations. This work represents the first successful demonstration of graph kernels to protein interfaces for effective discrimination of near-native and non-native conformations of protein complexes. AVAILABILITY AND IMPLEMENTATION: The iScore code is freely available from Github: https://github.com/DeepRank/iScore (DOI: 10.5281/zenodo.2630567). And the docking models used are available from SBGrid: https://data.sbgrid.org/dataset/684). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Cunliang Geng, Yong Jung, Nicolas Renaud, Vasant G. Honavar, Alexandre M. J. J. Bonvin, Li C. Xue
Bioinform.4
2019 Minimum Intervention Cover of a Causal Graph
Saravanan Kandasamy 0002, Arnab Bhattacharyya 0001, Vasant G. Honavar
AAAI3
2019 MEGAN: A Generative Adversarial Network for Multi-View Network Embedding
abstract
Data from many real-world applications can be naturally represented by multi-view networks where the different views encode different types of relationships (e.g., friendship, shared interests in music, etc.) between real-world individuals or entities. There is an urgent need for methods to obtain low-dimensional, information preserving and typically nonlinear embeddings of such multi-view networks. However, most of the work on multi-view learning focuses on data that lack a network structure, and most of the work on network embeddings has focused primarily on single-view networks. Against this background, we consider the multi-view network representation learning problem, i.e., the problem of constructing low-dimensional information preserving embeddings of multi-view networks. Specifically, we investigate a novel Generative Adversarial Network (GAN) framework for Multi-View Network Embedding, namely MEGAN, aimed at preserving the information from the individual network views, while accounting for connectivity across (and hence complementarity of and correlations between) different views. The results of our experiments on two real-world multi-view data sets show that the embeddings obtained using MEGAN outperform the state-of-the-art methods on node classification, link prediction and visualization tasks.
Suhang Wang, Tsung-Yu Hsieh, Xianfeng Tang, Vasant G. Honavar
IJCAI5
2019 Towards Robust Relational Causal Discovery
Sanghack Lee, Vasant G. Honavar
UAI2
2019 Improving Image Captioning by Leveraging Knowledge Graphs
abstract
We explore the use of a knowledge graphs, that capture general or commonsense knowledge, to augment the information extracted from images by the state-of-the-art methods for image captioning. We compare the performance of image captioning systems that as measured by CIDEr-D, a performance measure that is explicitly designed for evaluating image captioning systems, on several benchmark data sets such as MS COCO. The results of our experiments show that the variants of the state-of-the-art methods for image captioning that make use of the information extracted from knowledge graphs can substantially outperform those that rely solely on the information extracted from images.
Yimin Zhou 0009, Vasant G. Honavar
WACV3
2019 Fairness in Algorithmic Decision Making: An Excursion Through the Lens of Causality
abstract
As virtually all aspects of our lives are increasingly impacted by algorithmic decision making systems, it is incumbent upon us as a society to ensure such systems do not become instruments of unfair discrimination on the basis of gender, race, ethnicity, religion, etc. We consider the problem of determining whether the decisions made by such systems are discriminatory, through the lens of causal models. We introduce two definitions of group fairness grounded in causality: fair on average causal effect (FACE), and fair on average causal effect on the treated (FACT). We use the Rubin-Neyman potential outcomes framework for the analysis of cause-effect relationships to robustly estimate FACE and FACT. We demonstrate the effectiveness of our proposed approach on synthetic data. Our analyses of two real-world data sets, the Adult income data set from the UCI repository (with gender as the protected attribute), and the NYC Stop and Frisk data set (with race as the protected attribute), show that the evidence of discrimination obtained by FACE and FACT, or lack thereof, is often in agreement with the findings from other studies. We further show that FACT, being somewhat more nuanced compared to FACE, can yield findings of discrimination that differ from those obtained using FACE.
Aria Khademi, Sanghack Lee, David Foley, Vasant G. Honavar
WWW4
2018 Top-N-Rank: A Scalable List-wise Ranking Method for Recommender Systems
abstract
We propose Top-N-Rank, a novel family of list-wise Learning-to-Rank models for reliably recommending the N top-ranked items. The proposed models optimize a variant of the widely used cumulative discounted gain (DCG) objective function which differs from DCG in two important aspects: (i) It limits the evaluation of DCG only on the top N items in the ranked lists, thereby eliminating the impact of low-ranked items on the learned ranking function; and (ii) it incorporates weights that allow the model to leverage multiple types of implicit feedback with differing levels of reliability or trustworthiness. Because the resulting objective function is non-smooth and hence challenging to optimize, we consider two smooth approximations of the objective function, using the traditional sigmoid function and the rectified linear unit (ReLU). We propose a family of learning-to-rank algorithms (Top-N-Rank) that work with any smooth objective function. Then, a more efficient variant, Top-N-Rank.ReLU, is introduced, which effectively exploits the properties of ReLU function to reduce the computational complexity of Top-N-Rank from quadratic to linear in the average number of items rated by users. The results of our experiments using two widely used benchmarks, namely, the MovieLens data set and the Amazon Video Games data set demonstrate that: (i) The "top-N truncation" of the objective function substantially improves the ranking quality of the top N recommendations; (ii) using the ReLU for smoothing the objective function yields significant improvement in both ranking quality as well as runtime as compared to using the sigmoid; and (iii) Top-N-Rank.ReLU substantially outperforms the well-performing list-wise ranking methods in terms of ranking quality.
Jinlong Hu 0002, Shoubin Dong, Vasant G. Honavar
IEEE BigData4
2018 Compositional Stochastic Average Gradient for Machine Learning and Related Applications
Tsung-Yu Hsieh, Yasser El-Manzalawy, Vasant G. Honavar
IDEAL (1)4
2018 A user similarity-based Top-N recommendation approach for mobile in-application advertising
Jinlong Hu 0002, Yuezhen Kuang, Vasant G. Honavar
Expert Syst. Appl.4
2017 Sleep/wake state prediction and sleep parameter estimation using unsupervised classification via clustering
abstract
Sleep quality impacts virtually all aspects of life, including health, mood, emotions, cognition, memory, behavior, and performance. Actigraphy offers a lower-cost alternative to conventional polysomnography (PSG), the gold standard for measuring sleep quality. Effective use of actigraphy for assessing sleep quality requires reliable methods for detecting sleep/wake states from actigraphy measurements. Machine learning offers a promising approach to building sleep/wake state detectors from actigraphy data. However, current machine learning approaches rely on expert labeled training data that can be expensive and laborious to acquire. In this work, we introduce a novel approach for integrating unsupervised learning algorithms and domain knowledge heuristics, based on statistical properties of clustered sleep and wake epochs, to develop reliable sleep/wake state prediction models using unlabeled wrist actigraphy data. Experimental results using a dataset of 37 participants and covering 282 sleeping periods demonstrate the viability of the proposed approach on developing sleep/wake state detection models from unlabeled actigraphy data with a predictive performance that is comparable with the performance of models developed using some state-of-the-art supervised learning algorithms applied to labeled actigraphy data. Our results lay the groundwork for developing fully automated machine learning models for sleep/wake state prediction and sleep parameters estimations by eliminating the need for costly and labor-intensive expert annotations of PSG recordings for labeling actigraphy data.
Yasser El-Manzalawy, Orfeu M. Buxton, Vasant G. Honavar
BIBM3
2017 Self-Discrepancy Conditional Independence Test
Sanghack Lee, Vasant G. Honavar
UAI2
2017 Towards Conditional Independence Test for Relational Data
Sanghack Lee, Vasant G. Honavar
UAI2
2017 Template-based protein-protein docking exploiting pairwise interfacial residue restraints
abstract
Although many advanced and sophisticated ab initio approaches for modeling protein-protein complexes have been proposed in past decades, template-based modeling (TBM) remains the most accurate and widely used approach, given a reliable template is available. However, there are many different ways to exploit template information in the modeling process. Here, we systematically evaluate and benchmark a TBM method that uses conserved interfacial residue pairs as docking distance restraints [referred to as alpha carbon-alpha carbon (CA-CA)-guided docking]. We compare it with two other template-based protein-protein modeling approaches, including a conserved non-pairwise interfacial residue restrained docking approach [referred to as the ambiguous interaction restraint (AIR)-guided docking] and a simple superposition-based modeling approach. Our results show that, for most cases, the CA-CA-guided docking method outperforms both superposition with refinement and the AIR-guided docking method. We emphasize the superiority of the CA-CA-guided docking on cases with medium to large conformational changes, and interactions mediated through loops, tails or disordered regions. Our results also underscore the importance of a proper refinement of superimposition models to reduce steric clashes. In summary, we provide a benchmarked TBM protocol that uses conserved pairwise interface distance as restraints in generating realistic 3D protein-protein interaction models, when reliable templates are available. The described CA-CA-guided docking protocol is based on the HADDOCK platform, which allows users to incorporate additional prior knowledge of the target system to further improve the quality of the resulting models.
Li C. Xue, João P. G. L. M. Rodrigues, Drena Dobbs, Vasant G. Honavar, Alexandre M. J. J. Bonvin
Briefings Bioinform.4
2016 On Learning Causal Models from Relational Data
abstract
Many applications call for learning causal models from relational data. We investigate Relational Causal Models (RCM) under relational counterparts of adjacency-faithfulness and orientation-faithfulness, yielding a simple approach to identifying a subset of relational d-separation queries needed for determining the structure of an RCM using d-separation against an unrolled DAG representation of the RCM. We provide original theoretical analysis that offers the basis of a sound and efficient algorithm for learning the structure of an RCM from relational data. We describe RCD-Light, a sound and efficient constraint-based algorithm that is guaranteed to yield a correct partially-directed RCM structure with at least as many edges oriented as in that produced by RCD, the only other existing algorithm for learning RCM. We show that unlike RCD, which requires exponential time and space, RCD-Light requires only polynomial time and space to orient the dependencies of a sparse RCM.
Sanghack Lee, Vasant G. Honavar
AAAI2
2016 Labeling actors in multi-view social networks by integrating information from within and across multiple views
abstract
Real world social networks typically consist of actors (individuals) that are linked to other actors or different types of objects via links of multiple types. Different types of relationships induce different views of the underlying social network. We consider the problem of labeling actors in such multi-view networks based on the connections among them. Given a social network in which only a subset of the actors are labeled, our goal is to predict the labels of the rest of the actors. We introduce a new random walk kernel, namely the Inter-Graph Random Walk Kernel (IRWK), for labeling actors in multi-view social networks. IRWK combines information from within each of the views as well as the links across different views. The results of our experiments on two real-world multi-view social networks show that: (i) IRWK classifiers outperform or are competitive with several state-of-the-art methods for labeling actors in a social network; (ii) IRWKs are robust with respect to different choices of user-specified parameters; and (iii) IRWK kernel computation converges very fast within a few iterations.
Ngot Bui, Vasant G. Honavar
IEEE BigData3
2016 A Characterization of Markov Equivalence Classes of Relational Causal Models under Path Semantics
Sanghack Lee, Vasant G. Honavar
UAI2
2016 Temporal Causality Analysis of Sentiment Change in a Cancer Survivor Network
abstract
Online health communities constitute a useful source of information and social support for patients. American Cancer Society's Cancer Survivor Network (CSN), a 173,000-member community, is the largest online network for cancer patients, survivors, and caregivers. A discussion thread in CSN is often initiated by a cancer survivor seeking support from other members of CSN. Discussion threads are multi-party conversations that often provide a source of social support e.g., by bringing about a change of sentiment from negative to positive on the part of the thread originator. While previous studies regarding cancer survivors have shown that members of an online health community derive benefits from their participation in such communities, causal accounts of the factors that contribute to the observed benefits have been lacking. We introduce a novel framework to examine the temporal causality of sentiment dynamics in the CSN. We construct a Probabilistic Computation Tree Logic representation and a corresponding probabilistic Kripke structure to represent and reason about the changes in sentiments of posts in a thread over time. We use a sentiment classifier trained using machine learning on a set of posts manually tagged with sentiment labels to classify posts as expressing either positive or negative sentiment. We analyze the probabilistic Kripke structure to identify the prima facie causes of sentiment change on the part of the thread originators in the CSN forum and their significance. We find that the sentiment of replies appears to causally influence the sentiment of the thread originator. Our experiments also show that the conclusions are robust with respect to the choice of the (i) classification threshold of the sentiment classifier; (ii) and the choice of the specific sentiment classifier used. We also extend the basic framework for temporal causality analysis to incorporate the uncertainty in the states of the probabilistic Kripke structure resulting from the use of an imperfect state transducer (in our case, the sentiment classifier). Our analysis of temporal causality of CSN sentiment dynamics offers new insights that the designers, managers and moderators of an online community such as CSN can utilize to facilitate and enhance the interactions so as to better meet the social support needs of the CSN participants. The proposed methodology for analysis of temporal causality has broad applicability in a variety of settings where the dynamics of the underlying system can be modeled in terms of state variables that change in response to internal or external inputs.
Ngot Bui, John Yen, Vasant G. Honavar
IEEE Trans. Comput. Soc. Syst.3
2015 Learning classifiers from remote RDF data stores augmented with RDFS subclass hierarchies
abstract
Rapid growth of RDF data in the Linked Open Data (LOD) cloud offers unprecedented opportunities for analyzing such data using machine learning algorithms. The massive size and distributed nature of LOD cloud present a challenging machine learning problem where the data can only be accessed remotely, i.e. through a query interface such as the SPARQL end-point of the data store. Existing approaches to learning classifiers from RDF data in such a setting fail to take advantage of RDF schema (RDFS) associated with the data store that asserts subclass hierarchies which provide information that can potentially be exploited by the learner. Against this background, we present a general approach that augments an existing directed graphical model with hidden variables that encode subclass hierarchies via probabilistic constraints. We also present an algorithm ProbAVT that adopts the variational Bayesian expectation maximization approach to efficiently learn parameters in such settings. Our experiments with several synthetic and real world datasets show that: (i) ProbAVT matches or outperforms its counterpart that does not incorporate background knowledge in the form of subclass hierarchies; (ii) ProbAVT remains competitive compared to other state-of-art models that incorporate subclass hierarchies, and is able to scale up to large hierarchies consisting of over tens of thousands of nodes.
Harris T. Lin, Ngot Bui, Vasant G. Honavar
IEEE BigData3
2014 A Conceptual Framework for Secrecy-preserving Reasoning in Knowledge Bases
abstract
In many applications, Knowledge Bases (KBs) contain confidential or private information (secrets). The KB should be able to use this secret information in its reasoning process but in answering user queries care must be exercised so that secrets are not revealed to unauthorized users. We consider this problem under the Open World Assumption (OWA) in a setting with multiple querying agents M 1 ,…, M m that can pose queries against the KB K and selectively share answers that they receive from K with one or more other querying agents. We assume that for each M i , the KB has a prespecified set of secrets S i that need to be protected from M i . Communication between querying agents is modeled by a communication graph, a directed graph with self-loops. We introduce a general framework and propose an approach to secrecy-preserving query answering based on sound and complete proof systems. The idea is to hide the truthful answer from a querying agent M i by feigning ignorance without lying (i.e., to provide the answer ‘Unknown’ to a query q if it needs to be protected. Under the OWA, a querying agent cannot distinguish between the case that q is being protected (for reasons of secrecy) and the case that it cannot be inferred from K . In the pre-query stage we compute a set of envelopes E 1 , …, E m (restricted to a finite subset of the set of formulae that are entailed by K ) so that S i ⊆ E i , and a query α posed by agent M i can be answered truthfully whenever α ∉ E i and ¬ α ∉ E i . After the pre-query stage, the envelope is updated as needed. We illustrate this approach with two simple cases: the Propositional Horn KBs and the Description Logic AL KBs.
Jia Tao 0001, Giora Slutzki, Vasant G. Honavar
ACM Trans. Comput. Log.3
2013 m-Transportability: Transportability of a Causal Effect from Multiple Environments
abstract
We study m-transportability, a generalization of transportability, which offers a license to use causal information elicited from experiments and observations in m>=1 source environments to estimate a causal effect in a given targetenvironment. We provide a novel characterization of m-transportability that directly exploits the completeness of do-calculus to obtain the necessary and sufficient conditions for m-transportability. We provide an algorithm for deciding m-transportability that determines whether a causal relation is m-transportable; and if it is, produces a transport formula, that is, a recipe for estimating the desired causal effect by combining experimental information from m source environments with observational information from the target environment.
Sanghack Lee, Vasant G. Honavar
AAAI2
2013 On the utility of abstraction in labeling actors in social networks
abstract
Social networks are naturally represented as heterogeneous networks with multiple types of objects e.g., actors, items and multiple types of links e.g., links between actors that denote social ties e.g., friendship, and links that connect actors to items e.g., photos, videos, articles, etc. that denote relationships between actors and items. In this paper, we consider the task of assigning labels to the unlabeled actors (individuals) in a large heterogeneous social network in which labels are available for a subset of actors. Specifically, we seek to learn a predictive model to label actors based on the attributes of the actors themselves and/or items that are linked to them in the network. Unfortunately, the number of distinct items, represented in real-world networks such as Facebook or Flickr is quite large (in the millions) although only a small subset of them are linked to specific actors. This leads to data sparsity which causes over-fitting and hence poor performance in predicting the labels of unlabeled actors. To address this problem, we induce hierarchical taxonomies over items and use the resulting taxonomies as a basis for selecting abstract and hence parsimonious representations of network data for learning the predictive models. Our experiments using three different predictors (Iterative classification Naïve Bayes, Iterative classification Logistic Regression, and EdgeCluster) on two real-world data sets, Last.fm and Flickr, show that the predictive models that take advantage of abstract representations of network data are competitive with, and in some cases, outperform those that do not.
Ngot Bui, Vasant G. Honavar
ASONAM2
2013 Transportability from Multiple Environments with Limited Experiments
abstract
This paper considers the problem of transferring experimental findings learned from multiple heterogeneous domains to a target environment, in which only limited experiments can be performed. We reduce questions of transportability from multiple domains and with limited scope to symbolic derivations in the do-calculus, thus extending the treatment of transportability from full experiments introduced in Pearl and Bareinboim (2011). We further provide different graphical and algorithmic conditions for computing the transport formula for this setting, that is, a way of fusing the observational and experimental information scattered throughout different domains to synthesize a consistent estimate of the desired effects.
Elias Bareinboim, Sanghack Lee, Vasant G. Honavar, Judea Pearl
NIPS3
2013 Causal Transportability of Experiments on Controllable Subsets of Variables: z-Transportability
Sanghack Lee, Vasant G. Honavar
UAI2
2013 Preference Based Service Adaptation Using Service Substitution
abstract
In many applications such as service-oriented computing, users often prefer some compositions over the others based on their preferences over non-functional attributes such as security and cost. After a composition is deployed, apart from changes in the functional requirements, service-oriented architectures often have to deal with changes in the user preferences over the non-functional attributes and/or repository of available components. We formulate the problem of adaptation as iterative substitution of appropriate components in a composition, and provide two algorithms that produce a sequence of increasingly preferred adaptations with time: a fast algorithm that searches for preferred adaptations by improving the valuation of the relatively more important attributes, and another that is computationally more intensive but guaranteed to produce at least one preferred adaptation, if one exists.
Ganesh Ram Santhanam, Samik Basu 0001, Vasant G. Honavar
Web Intelligence3
2012 Unambiguity Regularization for Unsupervised Learning of Probabilistic Grammars
Kewei Tu, Vasant G. Honavar
EMNLP-CoNLL2
2012 Predicting protein-protein interface residues using local surface structural similarity
abstract
BACKGROUND: Identification of the residues in protein-protein interaction sites has a significant impact in problems such as drug discovery. Motivated by the observation that the set of interface residues of a protein tend to be conserved even among remote structural homologs, we introduce PrISE, a family of local structural similarity-based computational methods for predicting protein-protein interface residues. RESULTS: We present a novel representation of the surface residues of a protein in the form of structural elements. Each structural element consists of a central residue and its surface neighbors. The PrISE family of interface prediction methods uses a representation of structural elements that captures the atomic composition and accessible surface area of the residues that make up each structural element. Each of the members of the PrISE methods identifies for each structural element in the query protein, a collection of similar structural elements in its repository of structural elements and weights them according to their similarity with the structural element of the query protein. PrISEL relies on the similarity between structural elements (i.e. local structural similarity). PrISEG relies on the similarity between protein surfaces (i.e. general structural similarity). PrISEC, combines local structural similarity and general structural similarity to predict interface residues. These predictors label the central residue of a structural element in a query protein as an interface residue if a weighted majority of the structural elements that are similar to it are interface residues, and as a non-interface residue otherwise. The results of our experiments using three representative benchmark datasets show that the PrISEC outperforms PrISEL and PrISEG; and that PrISEC is highly competitive with state-of-the-art structure-based methods for predicting protein-protein interface residues. Our comparison of PrISEC with PredUs, a recently developed method for predicting interface residues of a query protein based on the known interface residues of its (global) structural homologs, shows that performance superior or comparable to that of PredUs can be obtained using only local surface structural similarity. PrISEC is available as a Web server at http://prise.cs.iastate.edu/ CONCLUSIONS: Local surface structural similarity based methods offer a simple, efficient, and effective approach to predict protein-protein interface residues.
Rafael A. Jordan, Yasser El-Manzalawy, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.4
2012 Protein-RNA interface residue prediction using machine learning: an assessment of the state of the art
abstract
BACKGROUND: RNA molecules play diverse functional and structural roles in cells. They function as messengers for transferring genetic information from DNA to proteins, as the primary genetic material in many viruses, as catalysts (ribozymes) important for protein synthesis and RNA processing, and as essential and ubiquitous regulators of gene expression in living organisms. Many of these functions depend on precisely orchestrated interactions between RNA molecules and specific proteins in cells. Understanding the molecular mechanisms by which proteins recognize and bind RNA is essential for comprehending the functional implications of these interactions, but the recognition 'code' that mediates interactions between proteins and RNA is not yet understood. Success in deciphering this code would dramatically impact the development of new therapeutic strategies for intervening in devastating diseases such as AIDS and cancer. Because of the high cost of experimental determination of protein-RNA interfaces, there is an increasing reliance on statistical machine learning methods for training predictors of RNA-binding residues in proteins. However, because of differences in the choice of datasets, performance measures, and data representations used, it has been difficult to obtain an accurate assessment of the current state of the art in protein-RNA interface prediction. RESULTS: We provide a review of published approaches for predicting RNA-binding residues in proteins and a systematic comparison and critical assessment of protein-RNA interface residue predictors trained using these approaches on three carefully curated non-redundant datasets. We directly compare two widely used machine learning algorithms (Naïve Bayes (NB) and Support Vector Machine (SVM)) using three different data representations in which features are encoded using either sequence- or structure-based windows. Our results show that (i) Sequence-based classifiers that use a position-specific scoring matrix (PSSM)-based representation (PSSMSeq) outperform those that use an amino acid identity based representation (IDSeq) or a smoothed PSSM (SmoPSSMSeq); (ii) Structure-based classifiers that use smoothed PSSM representation (SmoPSSMStr) outperform those that use PSSM (PSSMStr) as well as sequence identity based representation (IDStr). PSSMSeq classifiers, when tested on an independent test set of 44 proteins, achieve performance that is comparable to that of three state-of-the-art structure-based predictors (including those that exploit geometric features) in terms of Matthews Correlation Coefficient (MCC), although the structure-based methods achieve substantially higher Specificity (albeit at the expense of Sensitivity) compared to sequence-based methods. We also find that the expected performance of the classifiers on a residue level can be markedly different from that on a protein level. Our experiments show that the classifiers trained on three different non-redundant protein-RNA interface datasets achieve comparable cross-validation performance. However, we find that the results are significantly affected by differences in the distance threshold used to define interface residues. CONCLUSIONS: Our results demonstrate that protein-RNA interface residue predictors that use a PSSM-based encoding of sequence windows outperform classifiers that use other encodings of sequence windows. While structure-based methods that exploit geometric features can yield significant increases in the Specificity of protein-RNA interface residue predictions, such increases are offset by decreases in Sensitivity. These results underscore the importance of comparing alternative methods using rigorous statistical procedures, multiple performance measures, and datasets that are constructed based on several alternative definitions of interface residues and redundancy cutoffs as well as including evaluations on independent test sets into the comparisons.
Rasna R. Walia, Cornelia Caragea, Benjamin A. Lewis, Fadi Towfic, Michael Terribilini, Yasser El-Manzalawy, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.8
2012 PSPACE Tableau Algorithms for Acyclic Modalized $\boldsymbol{\mathcal{ALC}}$
Jia Tao 0001, Giora Slutzki, Vasant G. Honavar
J. Autom. Reason.3
2011 Verifying Intervention Policies to Counter Infection Propagation over Networks: A Model Checking Approach
abstract
Spread of infections (diseases, ideas, etc.) in a networkcan be modeled as the evolution of states of nodes ina graph as a function of the states of their neighbors.Given an initial configuration of a network in which asubset of the nodes have been infected, and an infectionpropagation function that specifies how the states ofthe nodes evolve over time, we show how to use modelchecking to identify, verify, and evaluate the effectivenessof intervention policies for containing the propagationof infection over such networks.
Ganesh Ram Santhanam, Yuly Suvorov, Samik Basu 0001, Vasant G. Honavar
AAAI4
2011 Multi-Instance Multi-Label Learning for Image Classification with Large Vocabularies
abstract
Multiple Instance Multiple Label learning problem has received much attention in machine learning and computer vision literature due to its applications in image classification and object detection. However, the current state-of-the-art solutions to this problem lack scalability and cannot be applied to datasets with a large number of instances and a large number of labels. In this paper we present a novel learning algorithm for Multiple Instance Multiple Label learning that is scalable for large datasets and performs comparable to the state-of-the-art algorithms. The proposed algorithm trains a set of discriminative multiple instance classifiers (one for each label in the vocabulary of all possible labels) and models the correlations among labels by finding a low rank weight matrix thus forcing the classifiers to share weights. This algorithm is a linear model unlike the state-of-the-art kernel methods which need to compute the kernel matrix. The model parameters are efficiently learned by solving an unconstrained optimization problem for which Stochastic Gradient Descent can be used to avoid storing all the data in memory. 1
Oksana Yakhnenko, Vasant G. Honavar
BMVC2
2011 On the Utility of Curricula in Unsupervised Learning of Probabilistic Grammars
abstract
Section 2 gives more details of the experimental settings and results. Section 3 discusses the related work.
Kewei Tu, Vasant G. Honavar
IJCAI2
2011 Exemplar-based Robust Coherent Biclustering
abstract
The biclustering, co-clustering, or subspace clustering problem involves simultaneously grouping the rows and columns of a data matrix to uncover biclusters or sub-matrices of the data matrix that optimize a desired objective function. In coherent biclustering, the objective function contains a coherence measure of the biclusters. We introduce a novel formulation of the coherent biclustering problem and use it to derive two algorithms. The first algorithm is based on loopy message passing; and the second relies on a greedy strategy yielding an algorithm that is significantly faster than the first. A distinguishing feature of these algorithms is that they identify an exemplar or a prototypical member of each bicluster. We note the interference from background elements in biclustering, and offer a means to circumvent such interference using additional regularization. Our experiments with synthetic as well as real-world datasets show that our algorithms are competitive with the current state-of-the-art algorithms for finding coherent biclusters.
Kewei Tu, Xixiu Ouyang, Dingyi Han, Vasant G. Honavar
SDM4
2011 Learning Relational Bayesian Classifiers from RDF Data
Harris T. Lin, Neeraj Koul, Vasant G. Honavar
ISWC (1)3
2011 Predicting RNA-Protein Interactions Using Only Sequence Information
abstract
BACKGROUND: RNA-protein interactions (RPIs) play important roles in a wide variety of cellular processes, ranging from transcriptional and post-transcriptional regulation of gene expression to host defense against pathogens. High throughput experiments to identify RNA-protein interactions are beginning to provide valuable information about the complexity of RNA-protein interaction networks, but are expensive and time consuming. Hence, there is a need for reliable computational methods for predicting RNA-protein interactions. RESULTS: We propose RPISeq, a family of classifiers for predicting RNA-protein interactions using only sequence information. Given the sequences of an RNA and a protein as input, RPIseq predicts whether or not the RNA-protein pair interact. The RNA sequence is encoded as a normalized vector of its ribonucleotide 4-mer composition, and the protein sequence is encoded as a normalized vector of its 3-mer composition, based on a 7-letter reduced alphabet representation. Two variants of RPISeq are presented: RPISeq-SVM, which uses a Support Vector Machine (SVM) classifier and RPISeq-RF, which uses a Random Forest classifier. On two non-redundant benchmark datasets extracted from the Protein-RNA Interface Database (PRIDB), RPISeq achieved an AUC (Area Under the Receiver Operating Characteristic (ROC) curve) of 0.96 and 0.92. On a third dataset containing only mRNA-protein interactions, the performance of RPISeq was competitive with that of a published method that requires information regarding many different features (e.g., mRNA half-life, GO annotations) of the putative RNA and protein partners. In addition, RPISeq classifiers trained using the PRIDB data correctly predicted the majority (57-99%) of non-coding RNA-protein interactions in NPInter-derived networks from E. coli, S. cerevisiae, D. melanogaster, M. musculus, and H. sapiens. CONCLUSIONS: Our experiments with RPISeq demonstrate that RNA-protein interactions can be reliably predicted using only sequence-derived information. RPISeq offers an inexpensive method for computational construction of RNA-protein interaction networks, and should provide useful insights into the function of non-coding RNAs. RPISeq is freely available as a web-based server at http://pridb.gdcb.iastate.edu/RPISeq/.
Usha Muppirala, Vasant G. Honavar, Drena Dobbs
BMC Bioinform.2
2011 HomPPI: A Class of Sequence Homology Based Protein-Protein Interface Prediction Methods
abstract
BACKGROUND: Although homology-based methods are among the most widely used methods for predicting the structure and function of proteins, the question as to whether interface sequence conservation can be effectively exploited in predicting protein-protein interfaces has been a subject of debate. RESULTS: We studied more than 300,000 pair-wise alignments of protein sequences from structurally characterized protein complexes, including both obligate and transient complexes. We identified sequence similarity criteria required for accurate homology-based inference of interface residues in a query protein sequence.Based on these analyses, we developed HomPPI, a class of sequence homology-based methods for predicting protein-protein interface residues. We present two variants of HomPPI: (i) NPS-HomPPI (Non partner-specific HomPPI), which can be used to predict interface residues of a query protein in the absence of knowledge of the interaction partner; and (ii) PS-HomPPI (Partner-specific HomPPI), which can be used to predict the interface residues of a query protein with a specific target protein.Our experiments on a benchmark dataset of obligate homodimeric complexes show that NPS-HomPPI can reliably predict protein-protein interface residues in a given protein, with an average correlation coefficient (CC) of 0.76, sensitivity of 0.83, and specificity of 0.78, when sequence homologs of the query protein can be reliably identified. NPS-HomPPI also reliably predicts the interface residues of intrinsically disordered proteins. Our experiments suggest that NPS-HomPPI is competitive with several state-of-the-art interface prediction servers including those that exploit the structure of the query proteins. The partner-specific classifier, PS-HomPPI can, on a large dataset of transient complexes, predict the interface residues of a query protein with a specific target, with a CC of 0.65, sensitivity of 0.69, and specificity of 0.70, when homologs of both the query and the target can be reliably identified. The HomPPI web server is available at http://homppi.cs.iastate.edu/. CONCLUSIONS: Sequence homology-based methods offer a class of computationally efficient and reliable approaches for predicting the protein-protein interface residues that participate in either obligate or transient interactions. For query proteins involved in transient interactions, the reliability of interface residue prediction can be improved by exploiting knowledge of putative interaction partners.
Li C. Xue, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.3
2011 Representing and Reasoning with Qualitative Preferences for Compositional Systems
Ganesh Ram Santhanam, Samik Basu 0001, Vasant G. Honavar
J. Artif. Intell. Res.3
2011 Predicting MHC-II Binding Affinity Using Multiple Instance Regression
abstract
Reliably predicting the ability of antigen peptides to bind to major histocompatibility complex class II (MHC-II) molecules is an essential step in developing new vaccines. Uncovering the amino acid sequence correlates of the binding affinity of MHC-II binding peptides is important for understanding pathogenesis and immune response. The task of predicting MHC-II binding peptides is complicated by the significant variability in their length. Most existing computational methods for predicting MHC-II binding peptides focus on identifying a nine amino acids core region in each binding peptide. We formulate the problems of qualitatively and quantitatively predicting flexible length MHC-II peptides as multiple instance learning and multiple instance regression problems, respectively. Based on this formulation, we introduce MHCMIR, a novel method for predicting MHC-II binding affinity using multiple instance regression. We present results of experiments using several benchmark data sets that show that MHCMIR is competitive with the state-of-the-art methods for predicting MHC-II binding peptides. An online web server that implements the MHCMIR method for MHC-II binding affinity prediction is freely accessible at http://ailab.cs.iastate.edu/mhcmir.
Yasser El-Manzalawy, Drena Dobbs, Vasant G. Honavar
IEEE ACM Trans. Comput. Biol. Bioinform.3
2010 Dominance Testing via Model Checking
abstract
Dominance testing, the problem of determining whether an outcome is preferred over another, is of fundamental importance in many applications. Hence, there is a need for algorithms and tools for dominance testing. CP-nets and TCP-nets are some of the widely studied languages for representing and reasoning with preferences. We reduce dominance testing in TCP-nets to reachability analysis in a graph of outcomes. We provide an encoding of TCP-nets in the form of a Kripke structure for CTL. We show how to compute dominance using NuSMV, a model checker for CTL. We present results of experiments that demonstrate the feasibility of our approach to dominance testing.
Ganesh Ram Santhanam, Samik Basu 0001, Vasant G. Honavar
AAAI3
2010 Scalable, updatable predictive models for sequence data
abstract
The emergence of data rich domains has led to an exponential growth in the size and number of data repositories, offering exciting opportunities to learn from the data using machine learning algorithms. In particular, sequence data is being made available at a rapid rate. In many applications, the learning algorithm may not have direct access to the entire dataset because of a variety of reasons such as massive data size or bandwidth limitation. In such settings, there is a need for techniques that can learn predictive models (e.g., classifiers) from large datasets without direct access to the data. We describe an approach to learn from massive sequence datasets using statistical queries. Specifically we show how Markov Models and Probabilistic Suffix Trees (PSTs) can be constructed from sequence databases that answer only a class of count queries. We analyze the query complexity (a measure of the number of queries needed) for constructing classifiers in such settings and outline some techniques to minimize the query complexity. We also show how some of the models can be updated in response to addition or deletion of subsets of sequences from the underlying sequence database.
Neeraj Koul, Ngot Bui, Vasant G. Honavar
BIBM3
2010 Abstraction Augmented Markov Models
abstract
High accuracy sequence classification often requires the use of higher order Markov models (MMs). However, the number of MM parameters increases exponentially with the range of direct dependencies between sequence elements, thereby increasing the risk of overfitting when the data set is limited in size. We present abstraction augmented Markov models (AAMMs) that effectively reduce the number of numeric parameters of k(th) order MMs by successively grouping strings of length k (i.e., k-grams) into abstraction hierarchies. We evaluate AAMMs on three protein subcellular localization prediction tasks. The results of our experiments show that abstraction makes it possible to construct predictive models that use significantly smaller number of features (by one to three orders of magnitude) as compared to MMs. AAMMs are competitive with and, in some cases, significantly outperform MMs. Moreover, the results show that AAMMs often perform significantly better than variable order Markov models, such as decomposed context tree weighting, prediction by partial match, and probabilistic suffix trees.
Cornelia Caragea, Adrian Silvescu, Doina Caragea, Vasant G. Honavar
ICDM4
2010 Ontology-guided Extraction of Complex Nested Relationships
abstract
Many applications call for methods to enable automatic extraction of structured information from unstructured natural language text. Due to inherent challenges of natural language processing, most of the existing methods for information extraction from text tend to be domain specific. We explore a modular ontology-based approach to information extraction that decouples domain-specific knowledge from the rules used for information extraction. We describe a framework for extraction of a subset of complex nested relationships (e.g., Joe reports that Jim is a reliable employee). The extracted relationships are output in the form of sets of RDF (resource description framework) triples, which can be queried using query languages for RDF and mined for knowledge acquisition.
Sushain Pandit, Vasant G. Honavar
ICTAI (2)2
2010 Automata-Based Verification of Security Requirements of Composite Web Services
abstract
With the increasing reliance of complex real-world applications on composite web services assembled from independently developed component services, there is a growing need for effective approaches to verifying that a composite service not only offers the required functionality but also satisfies the desired non-functional requirements (NFRs). In high-assurance applications such as traffic control, medical decision support, and coordinated response to civil emergencies, of special concern are NFRs having to do with security, safety and reliability of composite services. Current approaches to verifying NFRs of composite services (as opposed to individual services) remain largely ad-hoc and informal in nature. In this paper we develop techniques for ensuring that a composite service meets the user-specified NFRs expressible in the form of hard constraints e.g., “response time has to be less than 5 minutes.” We introduce an automata-based framework for verifying that a composite service satisfies the desired NFRs based on the known guarantees regarding the non-functional properties of the component services. We further show how to improve the efficiency of verifying that a composite service indeed satisfies a desired set of NFRs by: (i) Exploiting information about the applicability of specific NFRs (e.g., security) only to certain subsets of the component services that make up a composite service to minimize the verification effort and (ii) Identifying inconsistencies between NFRs with overlapping scopes. We illustrate how our approach can be used to verify the security requirements for an Emergency Management System. We also show how the approach can be used to verify whether a composite service satisfies any desired set of NFRs that can be expressed in the form of hard constraints of a quantitative nature.
Hongyu Sun 0001, Samik Basu 0001, Vasant G. Honavar, Robyn R. Lutz
ISSRE3
2010 Efficient Dominance Testing for Unconditional Preferences
Ganesh Ram Santhanam, Samik Basu 0001, Vasant G. Honavar
KR3
2010 Learning in Presence of Ontology Mapping Errors
abstract
The widespread use of ontologies to associate semantics with data has resulted in a growing interest in the problem of learning predictive models from data sources that use different ontologies to model the same underlying domain (world of interest). Learning from such semantically disparate data sources involves the use of a mapping to resolve semantic disparity among the ontologies used. Often, in practice, the mapping used to resolve the disparity may contain errors and as such the learning algorithms used in such a setting must be robust in presence of mapping errors. We reduce the problem of learning from semantically disparate data sources in the presence of mapping errors to a variant of the problem of learning in the presence of nasty classification noise. This reduction allows us to transfer theoretical results and algorithms from the latter to the former.
Neeraj Koul, Vasant G. Honavar
Web Intelligence2
2010 Semi-supervised prediction of protein subcellular localization using abstraction augmented Markov models
abstract
BACKGROUND: Determination of protein subcellular localization plays an important role in understanding protein function. Knowledge of the subcellular localization is also essential for genome annotation and drug discovery. Supervised machine learning methods for predicting the localization of a protein in a cell rely on the availability of large amounts of labeled data. However, because of the high cost and effort involved in labeling the data, the amount of labeled data is quite small compared to the amount of unlabeled data. Hence, there is a growing interest in developing semi-supervised methods for predicting protein subcellular localization from large amounts of unlabeled data together with small amounts of labeled data. RESULTS: In this paper, we present an Abstraction Augmented Markov Model (AAMM) based approach to semi-supervised protein subcellular localization prediction problem. We investigate the effectiveness of AAMMs in exploiting unlabeled data. We compare semi-supervised AAMMs with: (i) Markov models (MMs) (which do not take advantage of unlabeled data); (ii) an expectation maximization (EM); and (iii) a co-training based approaches to semi-supervised training of MMs (that make use of unlabeled data). CONCLUSIONS: The results of our experiments on three protein subcellular localization data sets show that semi-supervised AAMMs: (i) can effectively exploit unlabeled data; (ii) are more accurate than both the MMs and the EM based semi-supervised MMs; and (iii) are comparable in performance, and in some cases outperform, the co-training based semi-supervised MMs.
Cornelia Caragea, Doina Caragea, Adrian Silvescu, Vasant G. Honavar
BMC Bioinform.4
2010 Detection of gene orthology from gene co-expression and protein interaction networks
abstract
BACKGROUND: Ortholog detection methods present a powerful approach for finding genes that participate in similar biological processes across different organisms, extending our understanding of interactions between genes across different pathways, and understanding the evolution of gene families. RESULTS: We exploit features derived from the alignment of protein-protein interaction networks and gene-coexpression networks to reconstruct KEGG orthologs for Drosophila melanogaster, Saccharomyces cerevisiae, Mus musculus and Homo sapiens protein-protein interaction networks extracted from the DIP repository and Mus musculus and Homo sapiens and Sus scrofa gene coexpression networks extracted from NCBI's Gene Expression Omnibus using the decision tree, Naive-Bayes and Support Vector Machine classification algorithms. CONCLUSIONS: The performance of our classifiers in reconstructing KEGG orthologs is compared against a basic reciprocal BLAST hit approach. We provide implementations of the resulting algorithms as part of BiNA, an open source biomolecular network alignment toolkit.
Fadi Towfic, Susan VanderPlas, Casey A. Oliver, Oliver Couture, Christopher K. Tuggle, M. Heather West Greenlee, Vasant G. Honavar
BMC Bioinform.7
2009 Detection of Gene Orthology Based on Protein-Protein Interaction Networks
abstract
Ortholog detection methods present a powerful approach for finding genes that participate in similar biological processes across different organisms, extending our understanding of interactions between genes across different pathways, and understanding the evolution of gene families. We exploit features derived from the alignment of protein-protein interaction networks to reconstruct KEGG orthologs for Drosophila melanogaster, Saccharomyces cerevisiae, Mus musculus and Homo sapiens protein-protein interaction networks extracted from the DIP repository for protein-protein interaction data using the decision tree, naive-Bayes and support vector machine classification algorithms. The performance of our classifiers in reconstructing KEGG orthologs is compared against a basic reciprocal BLAST hit approach. We provide implementations of the resulting algorithms as part of BiNA, an open source biomolecular network alignment toolkit.
Fadi Towfic, M. Heather West Greenlee, Vasant G. Honavar
BIBM3
2009 MICCLLR: Multiple-Instance Learning Using Class Conditional Log Likelihood Ratio
Yasser El-Manzalawy, Vasant G. Honavar
Discovery Science2
2009 Combining Super-Structuring and Abstraction on Sequence Classification
abstract
We present an approach to adapting the data representation used by a learner on sequence classification tasks. Our approach that exploits the complementary strengths of super-structuring (constructing complex features by combining existing features) and abstraction (grouping of similar features to generate more abstract features), yields smaller and, at the same time, accurate models. Super-structuring provides a way to increase the predictive accuracy of the learned models by enriching the data representation (and hence, increases the complexity of the learned models) whereas abstraction helps reduce the number of model parameters by simplifying the data representation. The results of our experiments on two data sets drawn from macromolecular sequence classification applications show that adapting data representation by combining super-structuring and abstraction, makes it possible to construct predictive models that use significantly smaller number of features (by one to three orders of magnitude) than those that are obtained using super-structuring alone, without sacrificing predictive accuracy. Our experiments also show that simplifying data representation using abstraction yields better performing models than those obtained using feature selection.
Adrian Silvescu, Cornelia Caragea, Vasant G. Honavar
ICDM3
2009 Learning Link-Based Classifiers from Ontology-Extended Textual Data
abstract
Real-world data mining applications call for effective strategies for learning predictive models from richly structured relational data. In this paper, we address the problem of learning classifiers from structured relational data that are annotated with relevant meta data. Specifically, we show how to learn classifiers at different levels of abstraction in a relational setting, where the structured relational data are organized in an abstraction hierarchy that describes the semantics of the content of the data. We show how to cope with some of the challenges presented by partial specification in the case of structured data, that unavoidably results from choosing a particular level of abstraction. Our solution to partial specification is based on a statistical method, called shrinkage. We present results of experiments in the case of learning link-based Naive Bayes classifiers on a text classification task that (i) demonstrate that the choice of the level of abstraction can impact the performance of the resulting link-based classifiers and (ii) examine the effect of partially specified data.
Cornelia Caragea, Doina Caragea, Vasant G. Honavar
ICTAI3
2009 Design and Implementation of a Query Planner for Data Integration
abstract
Many applications require integrated access to multiple distributed, autonomous, and often semantically disparate data. Hence there is a need for bridging the semantic gap between the user and the data sources and for answering user queries based on the contents of multiple data sources. This paper describes a query planner that solves these two problems.
Neeraj Koul, Vasant G. Honavar
ICTAI2
2009 Multi-Modal Hierarchical Dirichlet Process Model for Predicting Image Annotation and Image-Object Label Correspondence
abstract
Many real-world applications call for learning predictive relationships from multi-modal data. In particular, in multi-media and web applications, given a dataset of images and their associated captions, one might want to construct a predictive model that not only predicts a caption for the image but also labels the individual objects in the image. We address this problem using a multi-modal hierarchical Dirichlet Process model (MoM-HDP) – a stochastic process for modeling multi-modal data. MoM-HDP is an analog of a multi-modal Latent Dirichlet Allocation (MoM-LDA) with an infinite number of mixture components. Thus MoM-HDP allows circumventing the need for a priori choice of the number of mixture components or the computational expense of model selection. During training, the model has access to an un-segmented image and its caption, but not the labels for each object in the image. The trained model is used to predict the label for each region of interest in a segmented image. The model parameters are estimated efficiently using variational inference. We use two large benchmark datasets to compare the performance of the proposed MoM-HDP model with that of MoM-LDA model as well as some simple alternatives: Naive Bayes and Logistic Regression classifiers based on the formulation of the image annotation and image-label correspondence problems as one-against-all classification. Our experimental results show that unlike MoM-LDA, the performance of MoM-HDP is invariant to the number of mixture components. Furthermore, our experimental evaluation shows that the generalization performance of MoM-HDP is superior to that of MoM-HDP as well as the one-against-all Naive Bayes and Logistic Regression classifiers.
Oksana Yakhnenko, Vasant G. Honavar
SDM2
2009 Aligning Biomolecular Networks Using Modular Graph Kernels
Fadi Towfic, M. Heather West Greenlee, Vasant G. Honavar
WABI3
2009 Mixture of experts models to exploit global sequence similarity on biomolecular sequence labeling
abstract
BACKGROUND: Identification of functionally important sites in biomolecular sequences has broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. Experimental determination of such sites lags far behind the number of known biomolecular sequences. Hence, there is a need to develop reliable computational methods for identifying functionally important sites from biomolecular sequences. RESULTS: We present a mixture of experts approach to biomolecular sequence labeling that takes into account the global similarity between biomolecular sequences. Our approach combines unsupervised and supervised learning techniques. Given a set of sequences and a similarity measure defined on pairs of sequences, we learn a mixture of experts model by using spectral clustering to learn the hierarchical structure of the model and by using bayesian techniques to combine the predictions of the experts. We evaluate our approach on two biomolecular sequence labeling problems: RNA-protein and DNA-protein interface prediction problems. The results of our experiments show that global sequence similarity can be exploited to improve the performance of classifiers trained to label biomolecular sequence data. CONCLUSION: The mixture of experts model helps improve the performance of machine learning methods for identifying functionally important sites in biomolecular sequences.
Cornelia Caragea, Jivko Sinapov, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.4
2009 Efficient Markov Network Structure Discovery Using Independence Tests
abstract
We present two algorithms for learning the structure of a Markov network from data: GSMN* and GSIMN. Both algorithms use statistical independence tests to infer the structure by successively constraining the set of structures consistent with the results of these tests. Until very recently, algorithms for structure learning were based on maximum likelihood estimation, which has been proved to be NP-hard for Markov networks due to the difficulty of estimating the parameters of the network, needed for the computation of the data likelihood. The independence-based approach does not require the computation of the likelihood, and thus both GSMN* and GSIMN can compute the structure efficiently (as shown in our experiments). GSMN* is an adaptation of the Grow-Shrink algorithm of Margaritis and Thrun for learning the structure of Bayesian networks. GSIMN extends GSMN* by additionally exploiting Pearl's well-known properties of the conditional independence relation to infer novel independences from known ones, thus avoiding the performance of statistical tests to estimate them. To accomplish this efficiently GSIMN uses the Triangle theorem, also introduced in this work, which is a simplified version of the set of Markov axioms. Experimental comparisons on artificial and real-world data sets show GSIMN can yield significant savings with respect to GSMN*, while generating a Markov network with comparable or in some cases improved quality. We also compare GSIMN to a forward-chaining implementation, called GSIMN-FCH, that produces all possible conditional independences resulting from repeatedly applying Pearl's theorems on the known conditional independence tests. The results of this comparison show that GSIMN, by the sole use of the Triangle theorem, is nearly optimal in terms of the set of independences tests that it infers.
Facundo Bromberg, Dimitris Margaritis 0001, Vasant G. Honavar
J. Artif. Intell. Res.3
2008 On the Decidability of Role Mappings between Modular Ontologies
Jie Bao 0001, George Voutsadakis, Giora Slutzki, Vasant G. Honavar
AAAI4
2008 Using Global Sequence Similarity to Enhance Biological Sequence Labeling
abstract
Identifying functionally important sites from biological sequences, formulated as a biological sequence labeling problem, has broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. In this paper, we present an approach to biological sequence labeling that takes into account the global similarity between biological sequences. Our approach combines unsupervised and supervised learning techniques. Given a set of sequences and a similarity measure defined on pairs of sequences, we learn a mixture of experts model by using spectral clustering to learn the hierarchical structure of the model and by using bayesian approaches to combine the predictions of the experts. We evaluate our approach on two important biological sequence labeling problems: RNA-protein and DNA-protein interface prediction problems. The results of our experiments show that global sequence similarity can be exploited to improve the performance of classifiers trained to label biological sequence data.
Cornelia Caragea, Jivko Sinapov, Drena Dobbs, Vasant G. Honavar
BIBM4
2008 Predicting Protective Linear B-Cell Epitopes Using Evolutionary Information
abstract
Mapping B-cell epitopes plays an important role in vaccine design, immunodiagnostic tests, and antibody production. Because the experimental determination of B-cell epitopes is time-consuming and expensive, there is an urgent need for computational methods for reliable identification of putative B-cell epitopes from antigenic sequences. In this study, we explore the utility of evolutionary profiles derived from antigenic sequences in improving the performance of machine learning methods for protective linear B-cell epitope prediction. Specifically, we compare propensity scale based methods with a Naive Bayes classifier using three different representations of the classifier input: amino acid identities, position specific scoring matrix (PSSM) profiles, and dipeptide composition. We find that in predicting protective linear B-cell epitopes, a Naive Bayes classifier trained using PSSM profiles significantly outperforms the propensity scale based methods as well as the Naive Bayes classifiers trained using the amino acid identity or dipeptide composition representations of input data.
Yasser El-Manzalawy, Drena Dobbs, Vasant G. Honavar
BIBM3
2008 TCP-Compose* - A TCP-Net Based Algorithm for Efficient Composition of Web Services Using Qualitative Preferences
Ganesh Ram Santhanam, Samik Basu 0001, Vasant G. Honavar
ICSOC3
2008 Attribute Value Taxonomy Generation through Matrix Based Adaptive Genetic Algorithm
abstract
We introduce a new adaptive genetic method for AVT generation, MCM-AVT-Learner. The MCM-AVT-Learner imports the mutation and crossover matrices which makes effective use of the fitness ranking and loci statistics information. The suggested method is not only parameter-free, but also capable of producing high quality AVTs. We describe experiments on several complete and missing benchmark data sets that compare the performance of AVT-DTL using the reslut AVTs of the MCM-AVT-Learner and existing AVT learning algorithms. Results show that the AVTs generated by MCM-AVT-Learner are competitive with human-generated AVTs or AVTs generated by HAC-AVT-Learner and GA-AVT-Learner in terms of classification accuracy and the compactness of the classifier.
Hyunsung Jo, Yong-chan Na, Byonghwa Oh, Jihoon Yang, Vasant G. Honavar
ICTAI (1)5
2008 Learning Classifiers from Large Databases Using Statistical Queries
abstract
We describe an approach to learning predictive models from large databases in settings where direct access to data is not available because of massive size of data, access restrictions, or bandwidth requirements. We outline some techniques for minimizing the number of statistical queries needed; and for efficiently coping with missing values in the data. We provide open source implementation of the decision tree and naive Bayes algorithms to demonstrate the feasibility of the proposed approach.
Neeraj Koul, Cornelia Caragea, Vasant G. Honavar, Vikas Bahirwani, Doina Caragea
Web Intelligence3
2008 Federated ALCI: Preliminary Report
abstract
We introduce F-ALCI, a federated version of the description logic ALCI. An F-ALCI ontology, like its package-based counterpart ALCIP-, consists of multiple ALCI ontologies that can import concepts or roles defined in other modules. Unlike ALCIP-which supports only contextualized negation, F-ALCI, supports contextualization of each of the logical connectives, a feature that allows more flexible reuse of knowledge from independently developed ontologies. We provide a new semantics for F-ALCI based on image domain relations and establish the conditions that need to be imposed on domain relations to ensure properties, such as preservation of unsatisfiability and monotonicity of inference, that are desirable in distributed web applications. We also establish the decidability of F-ALCI.
George Voutsadakis, Giora Slutzki, Vasant G. Honavar, Jie Bao 0001
Web Intelligence3
2008 Use of machine learning algorithms to classify binary protein sequences as highly-designable or poorly-designable
abstract
BACKGROUND: By using a standard Support Vector Machine (SVM) with a Sequential Minimal Optimization (SMO) method of training, Naïve Bayes and other machine learning algorithms we are able to distinguish between two classes of protein sequences: those folding to highly-designable conformations, or those folding to poorly- or non-designable conformations. RESULTS: First, we generate all possible compact lattice conformations for the specified shape (a hexagon or a triangle) on the 2D triangular lattice. Then we generate all possible binary hydrophobic/polar (H/P) sequences and by using a specified energy function, thread them through all of these compact conformations. If for a given sequence the lowest energy is obtained for a particular lattice conformation we assume that this sequence folds to that conformation. Highly-designable conformations have many H/P sequences folding to them, while poorly-designable conformations have few or no H/P sequences. We classify sequences as folding to either highly- or poorly-designable conformations. We have randomly selected subsets of the sequences belonging to highly-designable and poorly-designable conformations and used them to train several different standard machine learning algorithms. CONCLUSION: By using these machine learning algorithms with ten-fold cross-validation we are able to classify the two classes of sequences with high accuracy -- in some cases exceeding 95%.
Myron Peto, Andrzej Kloczkowski, Vasant G. Honavar, Robert L. Jernigan
BMC Bioinform.3
2007 A Semantic Importing Approach to Knowledge Reuse from Multiple Ontologies
Jie Bao 0001, Giora Slutzki, Vasant G. Honavar
AAAI3
2007 Assessing the Performance of Macromolecular Sequence Classifiers
abstract
Machine learning approaches offer some of the most cost-effective approaches to building predictive models (e.g., classifiers) in a broad range of applications in computational biology. Comparing the effectiveness of different algorithms requires reliable procedures for accurately assessing the performance (e.g., accuracy, sensitivity, and specificity) of the resulting predictive classifiers. The difficulty of this task is compounded by the use of different data selection and evaluation procedures and in some cases, even different definitions for the same performance measures. We explore the problem of assessing the performance of predictive classifiers trained on macromolecular sequence data, with an emphasis on cross-validation and data selection methods. Specifically, we compare sequence-based and window-based cross-validation procedures on three sequence-based prediction tasks: identification of glycosylation sites, RNA-Protein interface residues, and Protein-Protein interface residues from amino acid sequence. Our experiments with two representative classifiers (Naive Bayes and Support Vector Machine) show that sequence-based and windows-based cross-validation procedures and data selection methods can yield different estimates of commonly used performance measures such as accuracy, Matthews correlation coefficient and area under the Receiver Operating Characteristic curve. We argue that the performance estimates obtained using sequence-based cross-validation provide more realistic estimates of performance than those obtained using window-based cross-validation.
Cornelia Caragea, Jivko Sinapov, Vasant G. Honavar, Drena Dobbs
BIBE3
2007 Analysis of Protein Protein Dimeric Interfaces
abstract
We analyzed the structural properties and the local surface environment of surface amino acid residues of proteins using a large, non-redundant dataset of 2383 protein chains in dimeric complexes from PDB. We compared the interface residues and non-interface residues based on six properties: side chain orientation, surface roughness, solid angle, ex value, hydrophobicity and interface cluster size. The results of our analysis show that interface residues have side chains pointing inward; interfaces are rougher, tend to be flat, moderately convex or concave and protrude more relative to non-interface surface residues. Interface residues tend to be surrounded by hydrophobic neighbors and tend to form clusters consisting of three or more interfaces residues. These findings are consistent with previous published studies using much smaller datasets, while allowing for more qualitative conclusions due to our larger dataset. Preliminary results suggest the possibility of using the six the properties to identify putative interface residues.
Feihong Wu, Fadi Towfic, Drena Dobbs, Vasant G. Honavar
BIBM4
2007 On Context-Specific Substitutability of Web Services
abstract
Web service substitution refers to the problem of identifying a service that can replace another service in the context of a composition with a specified functionality. Existing solutions to this problem rely on detecting the functional and behavioral equivalence of a particular service to be replaced and candidate services that could replace it. We introduce the notion of context-specific substitutability, where context refers to the overall functionality of the composition that is required to be maintained after replacement of its constituents. Using the context information, we investigate two variants of the substitution problem, namely environment-independent and environment- dependent, where environment refers to the constituents of a composition and show how the substitutability criteria can be relaxed within this model. We provide a logical formulation of the resulting criteria based on model checking techniques as well as prove the soundness and completeness of the proposed approach.
Jyotishman Pathak, Samik Basu 0001, Vasant G. Honavar
ICWS3
2007 Privacy-Preserving Reasoning on the SemanticWeb
abstract
Many semantic web applications require selective sharing of ontologies between autonomous entities due to copyright, privacy or security concerns. In such cases, an agent might want to hide a part of its ontology while sharing the rest. However, prohibiting any use of the hidden part of the ontology in answering queries from other agents may be overly restrictive. We provide a framework for privacy- preserving reasoning in which an agent can safely answer queries against its knowledge base using inferences based on both the hidden and visible part of the knowledge base, without revealing the hidden knowledge. We show an application of this framework in the widely used special case of hierarchical ontologies.
Jie Bao 0001, Giora Slutzki, Vasant G. Honavar
Web Intelligence3
2007 Exploring inconsistencies in genome-wide protein function annotations: a machine learning approach
abstract
BACKGROUND: Incorrectly annotated sequence data are becoming more commonplace as databases increasingly rely on automated techniques for annotation. Hence, there is an urgent need for computational methods for checking consistency of such annotations against independent sources of evidence and detecting potential annotation errors. We show how a machine learning approach designed to automatically predict a protein's Gene Ontology (GO) functional class can be employed to identify potential gene annotation errors. RESULTS: In a set of 211 previously annotated mouse protein kinases, we found that 201 of the GO annotations returned by AmiGO appear to be inconsistent with the UniProt functions assigned to their human counterparts. In contrast, 97% of the predicted annotations generated using a machine learning approach were consistent with the UniProt annotations of the human counterparts, as well as with available annotations for these mouse protein kinases in the Mouse Kinome database. CONCLUSION: We conjecture that most of our predicted annotations are, therefore, correct and suggest that the machine learning approach developed here could be routinely used to detect potential errors in GO annotations generated by high-throughput gene annotation projects. Editors Note: Authors from the original publication (Okazaki et al.: Nature 2002, 420:563-73) have provided their response to Andorf et al, directly following the correspondence.
Carson M. Andorf, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.3
2007 Glycosylation site prediction using ensembles of Support Vector Machine classifiers
abstract
BACKGROUND: Glycosylation is one of the most complex post-translational modifications (PTMs) of proteins in eukaryotic cells. Glycosylation plays an important role in biological processes ranging from protein folding and subcellular localization, to ligand recognition and cell-cell interactions. Experimental identification of glycosylation sites is expensive and laborious. Hence, there is significant interest in the development of computational methods for reliable prediction of glycosylation sites from amino acid sequences. RESULTS: We explore machine learning methods for training classifiers to predict the amino acid residues that are likely to be glycosylated using information derived from the target amino acid residue and its sequence neighbors. We compare the performance of Support Vector Machine classifiers and ensembles of Support Vector Machine classifiers trained on a dataset of experimentally determined N-linked, O-linked, and C-linked glycosylation sites extracted from O-GlycBase version 6.00, a database of 242 proteins from several different species. The results of our experiments show that the ensembles of Support Vector Machine classifiers outperform single Support Vector Machine classifiers on the problem of predicting glycosylation sites in terms of a range of standard measures for comparing the performance of classifiers. The resulting methods have been implemented in EnsembleGly, a web server for glycosylation site prediction. CONCLUSION: Ensembles of Support Vector Machine classifiers offer an accurate and reliable approach to automated identification of putative glycosylation sites in glycoprotein sequences.
Cornelia Caragea, Jivko Sinapov, Adrian Silvescu, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.5
2007 Software fault tree and coloured Petri net-based specification, design and implementation of agent-based intrusion detection systems
abstract
The integration of Software Fault Tree (SFT), which describes intrusions and Coloured Petri Nets (CPNs) that specifies design, is examined for an Intrusion Detection System (IDS). The IDS under development is a collection of mobile agents that detect, classify, and correlate the system and network activities. SFTs, augmented with nodes that describe trust, temporal and contextual relationships, are used to describe intrusions. CPNs for intrusion detection are built using CPN templates created from the augmented SFTs. Hierarchical CPNs are created to detect critical stages of intrusions. The agentbased implementation of the IDS is then constructed from the CPNs. Examples of intrusions and descriptions of the prototype implementation are used to demonstrate how the CPN approach has been used in the development of the IDS. The main contribution of this paper is an approach to systematic specification, design and implementation of an IDS; Innovations include (1) using stages of intrusions to structure the specification and design of the IDS; (2) augmentation of SFT with trust, temporal and contextual nodes to model intrusions; (3) algorithmic construction of CPNs from augmented SFT; and (4) generation of mobile agents from CPNs.
Guy G. Helmer, Johnny S. Wong, Mark Slagell, Vasant G. Honavar, Leslie L. Miller, Natalia Stakhanova
Int. J. Inf. Comput. Secur.4
2006 Experimental Comparison of Feature Subset Selection Using GA and ACO Algorithm
Keunjoon Lee, Jinu Joo, Jihoon Yang, Vasant G. Honavar
ADMA4
2006 Learning Classifiers from Distributed, Ontology-Extended Data Sources
Doina Caragea, Jun Zhang 0002, Jyotishman Pathak, Vasant G. Honavar
DaWaK4
2006 Modeling Web Services by Iterative Reformulation of Functional and Non-functional Requirements
Jyotishman Pathak, Samik Basu 0001, Vasant G. Honavar
ICSOC3
2006 Selecting and Composing Web Services through Iterative Reformulation of Functional Specifications
abstract
We propose a specification-driven approach to Web service composition. The proposed framework allows users to start with a high-level, possibly incomplete specification of a desired (goal) service that is to be realized using a subset of the available component services. These services are represented by the system using transition systems augmented with guards over variables with infinite domains and are used to determine a strategy for their composition that would realize the goal service. In the event that the goal service cannot be realized using the available services, the system identifies the cause(s) for such failure which can then be used by the developer to reformulate the goal specification. Thus, the system supports Web service composition through iterative refinement of the functional specifications. We present a prototype implementation in tabled-logic programming environment that illustrates the key features of the proposed approach
Jyotishman Pathak, Samik Basu 0001, Robyn R. Lutz, Vasant G. Honavar
ICTAI4
2006 RNBL-MN: A Recursive Naive Bayes Learner for Sequence Classification
Dae-Ki Kang, Adrian Silvescu, Vasant G. Honavar
PAKDD3
2006 TRIPPER: Rule Learning Using Taxonomies
Flavian Vasile, Adrian Silvescu, Dae-Ki Kang, Vasant G. Honavar
PAKDD4
2006 Efficient Markov Network Structure Discovery using Independence Tests
abstract
We present two algorithms for learning the structure of a Markov network from discrete data: GSMN and GSIMN. Both algorithms use statistical conditional independence tests on data to infer the structure by successively constraining the set of structures consistent with the results of these tests. GSMN is a natural adaptation of the Grow-Shrink algorithm of Margaritis and Thrun for learning the structure of Bayesian networks. GSIMN extends GSMN by additionally exploiting Pearl's well-known properties of conditional independence relations to infer novel independencies from known independencies, thus avoiding the need to perform these tests. Experiments on artificial and real data sets show GSIMN can yield savings of up to 70% with respect to GSMN, while generating a Markov network with comparable or in several cases considerably improved quality. In addition to GSMN, we also compare GSIMN to a forward-chaining implementation, called GSIMN-FCH, that produces all possible conditional independence results by repeatedly applying Pearl's theorems on the known conditional independence tests. The results of this comparison show that GSIMN is nearly optimal in terms of the number of tests it can infer, under a fixed ordering of the tests performed.
Facundo Bromberg, Dimitris Margaritis 0001, Vasant G. Honavar
SDM3
2006 On the Semantics of Linking and Importing in Modular Ontologies
Jie Bao 0001, Doina Caragea, Vasant G. Honavar
ISWC3
2006 Package-Based Description Logics - Preliminary Results
Jie Bao 0001, Doina Caragea, Vasant G. Honavar
ISWC3
2006 A Tableau-Based Federated Reasoning Algorithm for Modular Ontologies
abstract
Many real world applications of ontologies call for reasoning with modular ontologies. We describe a tableau-based reasoning algorithm based on package-based description logics (P-DL), an modular ontology language that extends description logics. Unlike classical approaches that assume a single centralized, consistent ontology, the proposed algorithm adopts a federated approach to reasoning with modular ontologies wherein each ontology module has associated with it, a local reasoner. The local reasoners communicate with each other as needed in an asynchronous fashion. Hence, the proposed approach offers an attractive approach to reasoning with multiple, autonomously developed ontology modules, in settings where it is neither possible nor desirable to integrate all involved modules into a single centralized ontology
Jie Bao 0001, Doina Caragea, Vasant G. Honavar
Web Intelligence3
2006 Predicting DNA-binding sites of proteins from amino acid sequence
abstract
BACKGROUND: Understanding the molecular details of protein-DNA interactions is critical for deciphering the mechanisms of gene regulation. We present a machine learning approach for the identification of amino acid residues involved in protein-DNA interactions. RESULTS: We start with a Naïve Bayes classifier trained to predict whether a given amino acid residue is a DNA-binding residue based on its identity and the identities of its sequence neighbors. The input to the classifier consists of the identities of the target residue and 4 sequence neighbors on each side of the target residue. The classifier is trained and evaluated (using leave-one-out cross-validation) on a non-redundant set of 171 proteins. Our results indicate the feasibility of identifying interface residues based on local sequence information. The classifier achieves 71% overall accuracy with a correlation coefficient of 0.24, 35% specificity and 53% sensitivity in identifying interface residues as evaluated by leave-one-out cross-validation. We show that the performance of the classifier is improved by using sequence entropy of the target residue (the entropy of the corresponding column in multiple alignment obtained by aligning the target sequence with its sequence homologs) as additional input. The classifier achieves 78% overall accuracy with a correlation coefficient of 0.28, 44% specificity and 41% sensitivity in identifying interface residues. Examination of the predictions in the context of 3-dimensional structures of proteins demonstrates the effectiveness of this method in identifying DNA-binding sites from sequence information. In 33% (56 out of 171) of the proteins, the classifier identifies the interaction sites by correctly recognizing at least half of the interface residues. In 87% (149 out of 171) of the proteins, the classifier correctly identifies at least 20% of the interface residues. This suggests the possibility of using such classifiers to identify potential DNA-binding motifs and to gain potentially useful insights into sequence correlates of protein-DNA interactions. CONCLUSION: Naïve Bayes classifiers trained to identify DNA-binding residues using sequence information offer a computationally efficient approach to identifying putative DNA-binding sites in DNA-binding proteins and recognizing potential DNA-binding motifs.
Changhui Yan, Michael Terribilini, Feihong Wu, Robert L. Jernigan, Drena Dobbs, Vasant G. Honavar
BMC Bioinform.6
2006 Towards the automatic generation of mobile agents for distributed intrusion detection system
Smruti Ranjan Behera, Johnny S. Wong, Guy G. Helmer, Vasant G. Honavar, Leslie L. Miller, Robyn R. Lutz, Mark Slagell
J. Syst. Softw.5
2006 Learning accurate and concise naïve Bayes classifiers from attribute value taxonomies and data
Jun Zhang 0002, Dae-Ki Kang, Adrian Silvescu, Vasant G. Honavar
Knowl. Inf. Syst.4
2005 Learning Support Vector Machines from Distributed Data Sources
Cornelia Caragea, Doina Caragea, Vasant G. Honavar
AAAI3
2005 Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar
ALT5
2005 Algorithms and Software for Collaborative Discovery from Autonomous, Semantically Heterogeneous, Distributed Information Sources
Doina Caragea, Jun Zhang 0002, Jie Bao 0001, Jyotishman Pathak, Vasant G. Honavar
Discovery Science5
2005 Learning Ontology-Aware Classifiers
Jun Zhang 0002, Doina Caragea, Vasant G. Honavar
Discovery Science3
2005 Discriminatively Trained Markov Model for Sequence Classification
abstract
In this paper, we propose a discriminative counterpart of the directed Markov Models of order k - 1, or MM(k - 1) for sequence classification. MM(k - 1) models capture dependencies among neighboring elements of a sequence. The parameters of the classifiers are initialized to based on the maximum likelihood estimates for their generative counterparts. We derive gradient based update equations for the parameters of the sequence classifiers in order to maximize the conditional likelihood function. Results of our experiments with data sets drawn from biological sequence classification (specifically protein function and subcellular localization) and text classification applications show that the discriminatively trained sequence classifiers outperform their generative counterparts, confirming the benefits of discriminative training when the primary objective is classification. Our experiments also show that the discriminatively trained MM(k - 1) sequence classifiers are competitive with the computationally much more expensive Support Vector Machines trained using k-gram representations of sequences.
Oksana Yakhnenko, Adrian Silvescu, Vasant G. Honavar
ICDM3
2005 Learning Classifiers for Misuse Detection Using a Bag of System Calls Representation
Dae-Ki Kang, Doug Fuller, Vasant G. Honavar
ISI3
2004 Generating AVTs Using GA for Learning Decision Tree Classifiers with Missing Data
Jinu Joo, Jun Zhang 0002, Jihoon Yang, Vasant G. Honavar
Discovery Science4
2004 Generation of Attribute Value Taxonomies from Data for Data-Driven Construction of Accurate and Compact Classifiers
abstract
Attribute value taxonomies (AVT) have been shown to be useful in constructing compact, robust, and comprehensible classifiers. However, in many application domains, human-designed AVTs are unavailable. We introduce AVT-learner, an algorithm for automated construction of attribute value taxonomies from data. AVT-learner uses hierarchical agglomerative clustering (HAC) to cluster attribute values based on the distribution of classes that co-occur with the values. We describe experiments on UCI data sets that compare the performance of AVT-NBL (an AVT-guided naive Bayes learner) with that of the standard naive Bayes learner (NBL) applied to the original data set. Our results show that the AVTs generated by AVT-learner are competitive with human-gene rated AVTs (in cases where such AVTs are available). AVT-NBL using AVTs generated by AVT-learner achieves classification accuracies that are comparable to or higher than those obtained by NBL; and the resulting classifiers are significantly more compact than those generated by NBL.
Dae-Ki Kang, Adrian Silvescu, Jun Zhang 0002, Vasant G. Honavar
ICDM4
2004 AVT-NBL: An Algorithm for Learning Compact and Accurate Naïve Bayes Classifiers from Attribute Value Taxonomies and Data
abstract
In many application domains, there is a need for learning algorithms that can effectively exploit attribute value taxonomies (AVT) - hierarchical groupings of attribute values - to learn compact, comprehensible, and accurate classifiers from data - including data that are partially specified. This paper describes AVT-NBL, a natural generalization of the naive Bayes learner (NBL), for learning classifiers from AVT and data. Our experimental results show that AVT-NBL is able to generate classifiers that are substantially more compact and more accurate than those produced by NBL on a broad range of data sets with different percentages of partially specified values. We also show that AVT-NBL is more efficient in its use of training data: AVT-NBL produces classifiers that outperform those produced by NBL using substantially fewer training examples.
Jun Zhang 0002, Vasant G. Honavar
ICDM2
2004 Predicting binding sites of hydrolase-inhibitor complexes by combining several methods
abstract
BACKGROUND: Protein-protein interactions play a critical role in protein function. Completion of many genomes is being followed rapidly by major efforts to identify interacting protein pairs experimentally in order to decipher the networks of interacting, coordinated-in-action proteins. Identification of protein-protein interaction sites and detection of specific amino acids that contribute to the specificity and the strength of protein interactions is an important problem with broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. RESULTS: In order to increase the power of predictive methods for protein-protein interaction sites, we have developed a consensus methodology for combining four different methods. These approaches include: data mining using Support Vector Machines, threading through protein structures, prediction of conserved residues on the protein surface by analysis of phylogenetic trees, and the Conservatism of Conservatism method of Mirny and Shakhnovich. Results obtained on a dataset of hydrolase-inhibitor complexes demonstrate that the combination of all four methods yield improved predictions over the individual methods. CONCLUSIONS: We developed a consensus method for predicting protein-protein interface residues by combining sequence and structure-based methods. The success of our consensus approach suggests that similar methodologies can be developed to improve prediction accuracies for other bioinformatic problems.
Taner Z. Sen, Andrzej Kloczkowski, Robert L. Jernigan, Changhui Yan, Vasant G. Honavar, Kai-Ming Ho, Cai-Zhuang Wang, Yungok Ihm, Haibo Cao, Drena Dobbs
BMC Bioinform.5
2004 Identification of interface residues in protease-inhibitor and antigen-antibody complexes: a support vector machine approach
Changhui Yan, Vasant G. Honavar, Drena Dobbs
Neural Comput. Appl.2
2003 Towards Simple, Easy-to-Understand, yet Accurate Classifiers
abstract
We design a method for weighting linear support vector machine classifiers or random hyperplanes, to obtain classifiers whose accuracy is comparable to the accuracy of a nonlinear support vector machine classifier, and whose results can be readily visualized. We conduct a simulation study to examine how our weighted linear classifiers behave in the presence of known structure. The results show that the weighted linear classifiers might perform well compared to the nonlinear support vector machine classifiers, while they are more readily interpretable than the nonlinear classifiers.
Doina Caragea, Dianne Cook, Vasant G. Honavar
ICDM3
2003 Learning from Attribute Value Taxonomies and Partially Specified Instances
Jun Zhang 0002, Vasant G. Honavar
ICML2
2003 A Multi-relational Decision Tree Learning Algorithm - Implementation and Experiments
Anna Atramentov, Hector Leiva, Vasant G. Honavar
ILP3
2003 Automated data-driven discovery of motif-based protein function classifiers
Xiangyun Wang, Diane Schroeder, Drena Dobbs, Vasant G. Honavar
Inf. Sci.4
2003 Lightweight agents for intrusion detection
Guy G. Helmer, Johnny S. Wong, Vasant G. Honavar, Leslie L. Miller
J. Syst. Softw.3
2002 Automated discovery of concise predictive rules for intrusion detection
Guy G. Helmer, Johnny S. Wong, Vasant G. Honavar, Leslie L. Miller
J. Syst. Softw.3
2002 A Software Fault Tree Approach to Requirements Analysis of an Intrusion Detection System
Guy G. Helmer, Johnny S. Wong, Mark Slagell, Vasant G. Honavar, Leslie L. Miller, Robyn R. Lutz
Requir. Eng.4
2001 Detection and identification of odorants using an electronic nose
abstract
Gas sensing systems for detection and identification of odorant molecules are of crucial importance in an increasing number of applications. Such applications include environmental monitoring, food quality assessment, airport security, and detection of hazardous gases. We describe a gas sensing system for detecting and identifying volatile organic compounds (VOC), and discuss the unique problems associated with the separability of signal patterns obtained by using such a system. We then present solutions for enhancing the separability of VOC patterns to enable classification. A new incremental learning algorithm that allows new odorants to be learned is also introduced.
Robi Polikar, Ruth Shinar, Vasant G. Honavar, Lalita Udpa, Marc D. Porter
ICASSP3
2001 Gaining insights into support vector machine pattern classifiers using projection-based tour methods
abstract
This paper discusses visual methods that can be used to understand and interpret the results of classification using support vector machines (SVM) on data with continuous real-valued variables. SVM induction algorithms build pattern classifiers by identifying a maximal margin separating hyperplane from training examples in high dimensional pattern spaces or spaces induced by suitable nonlinear kernel transformations over pattern spaces. SVM have been demonstrated to be quite effective in a number of practical pattern classification tasks. Since the separating hyperplane is defined in terms of more than two variables it is necessary to use visual techniques that can navigate the viewer through high-dimensional spaces. We demonstrate the use of projection-based tour methods to gain useful insights into SVM classifiers with linear kernels on 8-dimensional data.
Doina Caragea, Dianne Cook, Vasant G. Honavar
KDD3
2001 Autonomous agents for coordinated distributed parameterized heuristic routing in large dynamic communication networks
Armin R. Mikler, Vasant G. Honavar, Johnny S. Wong
J. Syst. Softw.2
2001 SMART mobile agent facility
Johnny S. Wong, Guy G. Helmer, Venkatraman Naganathan, Sriniwas Polavarapu, Vasant G. Honavar, Leslie L. Miller
J. Syst. Softw.5
2001 Introduction
Vasant G. Honavar, Colin de la Higuera
Mach. Learn.1
2001 Learning DFA from Simple Examples
Rajesh Parekh, Vasant G. Honavar
Mach. Learn.2
2001 Learn++: an incremental learning algorithm for supervised neural networks
abstract
We introduce Learn++, an algorithm for incremental training of neural network (NN) pattern classifiers. The proposed algorithm enables supervised NN paradigms, such as the multilayer perceptron (MLP), to accommodate new data, including examples that correspond to previously unseen classes. Furthermore, the algorithm does not require access to previously used data during subsequent incremental learning sessions, yet at the same time, it does not forget previously acquired knowledge. Learn++ utilizes ensemble of classifiers by generating multiple hypotheses using training data sampled according to carefully tailored distributions. The outputs of the resulting classifiers are combined using a weighted majority voting procedure. We present simulation results on several benchmark datasets as well as a real-world classification task. Initial results indicate that the proposed algorithm works rather well in practice. A theoretical upper bound on the error of the classifiers constructed by Learn++ is also provided.
Robi Polikar, L. Upda, S. S. Upda, Vasant G. Honavar
IEEE Trans. Syst. Man Cybern. Part C4
2000 LEARN++: an incremental learning algorithm for multilayer perceptron networks
abstract
We introduce a supervised learning algorithm that gives neural network classification algorithms the capability of learning incrementally from new data without forgetting what has been learned in earlier training sessions. Schapire's (1990) boosting algorithm, originally intended for improving the accuracy of weak learners, has been modified to be used in an incremental learning setting. The algorithm is based on generating a number of hypotheses using different distributions of the training data and combining these hypotheses using a weighted majority voting. This scheme allows the classifier previously trained with a training database, to learn from new data when the original data is no longer available, even when new classes are introduced. Initial results on incremental training of multilayer perceptron networks on synthetic as well as real-world data are presented in this paper.
Robi Polikar, Lalita Udpa, Satish S. Udpa, Vasant G. Honavar
ICASSP4
2000 Constructive neural-network learning algorithms for pattern classification
abstract
Constructive learning algorithms offer an attractive approach for the incremental construction of near-minimal neural-network architectures for pattern classification. They help overcome the need for ad hoc and often inappropriate choices of network topology in algorithms that search for suitable weights in a priori fixed network architectures. Several such algorithms are proposed in the literature and shown to converge to zero classification errors (under certain assumptions) on tasks that involve learning a binary to binary mapping (i.e., classification problems involving binary-valued input attributes and two output categories). We present two constructive learning algorithms MPyramid-real and MTiling-real that extend the pyramid and tiling algorithms, respectively, for learning real to M-ary mappings (i.e., classification problems involving real-valued input attributes and multiple output classes). We prove the convergence of these algorithms and empirically demonstrate their applicability to practical pattern classification problems. Additionally, we show how the incorporation of a local pruning step can eliminate several redundant neurons from MTiling-real networks.
Rajesh Parekh, Jihoon Yang, Vasant G. Honavar
IEEE Trans. Neural Networks Learn. Syst.3
1999 Feature Selection Using a Genetic Algorithm for Intrusion Detection
Guy G. Helmer, Johnny S. Wong, Vasant G. Honavar, Leslie L. Miller
GECCO3
1999 Simple DFA are Polynomially Probably Exactly Learnable from Simple Examples
Rajesh Parekh, Vasant G. Honavar
ICML2
1999 Data-Driven Theory Refinement Using KBDistAl
Jihoon Yang, Rajesh Parekh, Vasant G. Honavar, Drena Dobbs
IDA3
1999 Data-driven theory refinement algorithms for bioinformatics
abstract
Bioinformatics and related applications call for efficient algorithms for knowledge-intensive learning and data-driven knowledge refinement. Knowledge based artificial neural networks offer an attractive approach to extending or modifying incomplete knowledge bases or domain theories. We present results of experiments with several such algorithms for data-driven knowledge discovery and theory refinement in some simple bioinformatics applications. Results of experiments on the ribosome binding site and promoter site identification problems indicate that the performance of KBDistAl and Tiling-Pyramid algorithms compares quite favorably with those of substantially more computationally demanding techniques.
Jihoon Yang, Rajesh Parekh, Vasant G. Honavar, Drena Dobbs
IJCNN3
1999 DistAl: An inter-pattern distance-based constructive learning algorithm
abstract
Multi-layer networks of threshold logic units (TLU) offer an attractive framework for the design of pattern classification systems. A new constructive neural network learning algorithm (DistAl) based on inter-pattern distance is introduced. DistAl constructs a single hidden layer of hyperspherical threshold neurons. Each neuron is designed to determine a cluster of training patterns belonging to the same class. The weights and thresholds of the hidden neurons are determined directly by comparing the inter-pattern distances of the training patterns. This offers a significant advantage over other constructive learning algorithms that use an iterative (and often time consuming) weight modification strategy to train individual neurons. The individual clusters (represented by the hidden neurons) are combined by a single output layer of threshold neurons. The speed of DistAl makes it a good candidate for datamining and knowledge acquisition from large datasets. The paper presents results of experiments using several artificial and real-world datasets. The results demonstrate that DistAl compares favorably with other learning algorithms for pattern classification.
Jihoon Yang, Rajesh Parekh, Vasant G. Honavar
Intell. Data Anal.3
1999 A neural-network architecture for syntax analysis
abstract
Artificial neural networks (ANN's), due to their inherent parallelism, offer an attractive paradigm for implementation of symbol processing systems for applications in computer science and artificial intelligence. This paper explores systematic synthesis of modular neural-network architectures for syntax analysis using a prespecified grammar--a prototypical symbol processing task which finds applications in programming language interpretation, syntax analysis of symbolic expressions, and high-performance compilers. The proposed architecture is assembled from ANN components for lexical analysis, stack, parsing and parse tree construction. Each of these modules takes advantage of parallel content-based pattern matching using a neural associative memory. The proposed neural-network architecture for syntax analysis provides a relatively efficient and high performance alternative to current computer systems for applications that involve parsing of LR grammars which constitute a widely used subset of deterministic context-free grammars. Comparison of quantitatively estimated performance of such a system [implemented using current CMOS very large scale integration (VLSI) technology] with that of conventional computers demonstrates the benefits of massively parallel neural-network architectures for symbol processing applications.
Chun-Hsien Chen, Vasant G. Honavar
IEEE Trans. Neural Networks2
1998 An object oriented approach to simulating large communication networks
Armin R. Mikler, Johnny S. K. Wong, Vasant G. Honavar
J. Syst. Softw.3
1997 Learning DFA from Simple Examples
Rajesh Parekh, Vasant G. Honavar
ALT2
1997 Quo Vadis - A Framework for Intelligent Routing in Large Communication Networks
Armin R. Mikler, Johnny S. Wong, Vasant G. Honavar
J. Syst. Softw.3
1995 A Neural Architecture for Content as well as Address-based Storage and Recall: Theory and Applications
abstract
This paper presents an approach to design of a neural architecture for both associative (content-addressed) and address-based memories. Several interesting properties of the memory module are mathematically analyzed in detail. When used as an associative memory, the proposed neural memory module supports recall from partial input patterns, (sequential) multiple recalls and fault-tolerance. When used as an address-based memory, the memory module can provide working space for dynamic representations for symbol processing and shared message-passing among neural network modules within an integrated neural network system. It also provides for real-time update of memory contents by one-shot learning without interference with other stored patterns.
Chun-Hsien Chen, Vasant G. Honavar
Connect. Sci.2
1993 Generative learning structures and processes for generalized connectionist networks
Vasant G. Honavar, Leonard Uhr
Inf. Sci.1
1992 Neural Network Design and the Complexity of Learning (Book Review)
Vasant G. Honavar
Mach. Learn.1
1990 Coordination and control structures and processes: possibilities for connectionist networks (CN)
abstract
The absence of powerful control structures and processes that synchronize, coordinate, switch between, choose among, regulate, direct, modulate interactions between, and combine distinct yet interdependent modules of large connectionist networks (CN) is probably one of the most important reasons why such networks have not yet succeeded at handling difficult tasks (e.g. complex object recognition and description, complex problem-solving, planning). This paper examines how CN built from large numbers of relatively simple neuronlike units can be given the ability to handle problems that in typical multicomputer networks and artificial intelligence programs— along with all other types of programs— are always handled using extremely elaborate and precisely worked out central control (coordination, synchronization, switching, etc.) The paper shows the several mechanisms for central control of this un-brain-like sort that CN already have built into them— albeit in hidden, often overlooked, ways. The kinds of control mechanisms found in computers, programs, fetal development, cellular function and the immune system, evolution, social organizations, and especially brains, that might be of use in CN are examined. Particularly intriguing suggestions are found in the pacemakers, oscillators, and other local sources of the brain' s complex partial synchronies; the diffuse, global effects of slow electrical waves and neurohormones; the developmental program that guides fetal development; communication and coordination within and among living cells; the working of the immune system; the evolutionary processes that operate on large populations of organisms; and the great variety of partially competing partially cooperating controls found in small groups, organizations, and larger societies. All these systems are rich in control— but typically control that emerges from complex interactions of many local and diffuse sources. This paper explores how several different kinds of plausible control mechanisms might be incorporated into CN, and assesses their potential benefits with respect to their cost.
Vasant G. Honavar, Leonard Uhr
J. Exp. Theor. Artif. Intell.1
1989 Generation, Local Receptive Fields and Global Convergence Improve Perceptual Learning in Connectionist Networks
Vasant G. Honavar, Leonard Uhr
IJCAI1