Shawn Martin

dblp:71/6817 · DBLP profile ↗
← Back
8ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Optimization for machine learning · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › drug discovery
drug-target interaction
0.112008
Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptor · Bioinform. 2008
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction
0.112008
Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptor · Bioinform. 2008
Bioinformatics and computational biology › biological network › network biology › network inference › gene regulatory network inference
boolean network inference
0.112007
Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007
Bioinformatics and computational biology
gene expression analysis
0.112007
Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference
0.112007
Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007
Bioinformatics and computational biology › gene expression analysis › time-series gene expression analysis
time-course microarray analysis
0.112007
Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007
Machine learning › Optimization for machine learning › regularized risk minimization
support vector machine training
0.112005
Training Support Vector Machines Using Gilbert's Algorithm · ICDM 2005
Bioinformatics and computational biology
protein-protein interaction prediction
0.112005
Predicting protein-protein interactions using signature products · Bioinform. 2005
Bioinformatics and computational biology
protein sequence analysis
0.012005
Predicting protein-protein interactions using signature products · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

signature molecular descriptor · 0.1machine learning · 0.1support vector regression · 0.1k-means clustering · 0.1support vector machine · 0.1signature product · 0.1sequential minimal optimization · 0.1kernel function · 0.1gilbert's algorithm · 0.1
YearPublicationVenuePosition
2015 Using Bipartite Anomaly Features for Cyber Security Applications
abstract
In this paper we use anomaly scores derived from a technique for bipartite graphs as features for a supervised machine learning algorithm for two cyber security problems: classifying Short Message Service (SMS) text messages as either spam or non-spam and detecting malicious lateral movement within a network. While disparate problems, both spam and lateral movement detection can be viewed as bipartite graphs and we can compute bipartite anomaly scores for each situation. The bipartite anomaly scores by themselves are not very predictive, but used as auxiliary features can boost the receiver operating characteristic (ROC) curve of a supervised classifier. We examine the UCI SMS Spam Collection Data Set for the SPAM problem and use an authentication graph from Los Alamos National Laboratory. We create features by dimensionality reduction through principal component analysis (PCA) on the message-term or user-computer matrix, and then augment those features with anomaly scores. By using the anomaly scores we are able to improve the area under the curve (AUC) for the receiver operating characteristic (ROC) up to 27.5% for the spam data and 21.4% for the authentication data.
Eric L. Goodman, Joe Ingram, Shawn Martin, Dirk Grunwald
ICMLA3
2011 Non-manifold surface reconstruction from high-dimensional point cloud data
Shawn Martin, Jean-Paul Watson
Comput. Geom.1
2008 Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptor
abstract
Abstract Motivation: Identifying protein enzymatic or pharmacological activities are important areas of research in biology and chemistry. Biological and chemical databases are increasingly being populated with linkages between protein sequences and chemical structures. There is now sufficient information to apply machine-learning techniques to predict interactions between chemicals and proteins at a genome scale. Current machine-learning techniques use as input either protein sequences and structures or chemical information. We propose here a method to infer protein–chemical interactions using heterogeneous input consisting of both protein sequence and chemical information. Results: Our method relies on expressing proteins and chemicals with a common cheminformatics representation. We demonstrate our approach by predicting whether proteins can catalyze reactions not present in training sets. We also predict whether a given drug can bind a target, in the absence of prior binding information for that drug and target. Such predictions cannot be made with current machine-learning techniques requiring binding information for individual reactions or individual targets. Availability and Contact: For questions, paper reprints, please contact Jean-Loup Faulon at [email protected]. Additional information on the signature molecular descriptor and codes can be downloaded at: http://www.cs.sandia.gov/~jfaulon/publication-signature.html Supplementary information: Supplementary data are available at Bioinformatics online.
Jean-Loup Faulon, Milind Misra, Shawn Martin, Kenneth L. Sale, Rajat Sapra
Bioinform.3
2007 Predicting building contamination using machine learning
abstract
Potential events involving biological or chemical contamination of buildings are of major concern in the area of homeland security. Tools are needed to provide rapid, on- site predictions of contaminant levels given only approximate measurements in limited locations throughout a building. In principal, such tools could use calculations based on physical process models to provide accurate predictions. In practice, however, physical process models are too complex and computationally costly to be used in a real-time scenario. In this paper, we investigate the feasibility of using machine learning to provide easily computed but approximate models that would be applicable in the field. We develop a machine learning method based on support vector machine regression and classification. We apply our method to problems of estimating contamination levels and contaminant source location.
Shawn Martin, Sean McKenna
ICMLA1
2007 Boolean dynamics of genetic regulatory networks inferred from microarray time series data
abstract
MOTIVATION: Methods available for the inference of genetic regulatory networks strive to produce a single network, usually by optimizing some quantity to fit the experimental observations. In this article we investigate the possibility that multiple networks can be inferred, all resulting in similar dynamics. This idea is motivated by theoretical work which suggests that biological networks are robust and adaptable to change, and that the overall behavior of a genetic regulatory network might be captured in terms of dynamical basins of attraction. RESULTS: We have developed and implemented a method for inferring genetic regulatory networks for time series microarray data. Our method first clusters and discretizes the gene expression data using k-means and support vector regression. We then enumerate Boolean activation-inhibition networks to match the discretized data. Finally, the dynamics of the Boolean networks are examined. We have tested our method on two immunology microarray datasets: an IL-2-stimulated T cell response dataset and a LPS-stimulated macrophage response dataset. In both cases, we discovered that many networks matched the data, and that most of these networks had similar dynamics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shawn Martin, Zhaoduo Zhang, Anthony Martino, Jean-Loup Faulon
Bioinform.1
2006 An Approximate Version of Kernel PCA
abstract
We propose an analog of kernel principal component analysis (kernel PCA). Our algorithm is based on an approximation of PCA which uses Gram-Schmidt orthonormalization. We combine this approximation with support vector machine kernels to obtain a nonlinear generalization of PCA. By using our approximation to PCA we are able to provide a more easily computed (in the case of many data points) and readily interpretable version of kernel PCA. After demonstrating our algorithm on some examples, we explore its use in applications to fluid flow and microarray data
Shawn Martin
ICMLA1
2005 Training Support Vector Machines Using Gilbert's Algorithm
abstract
Support vector machines are classifiers designed around the computation of an optimal separating hyperplane. This hyperplane is typically obtained by solving a constrained quadratic programming problem, but may also be located by solving a nearest point problem. Gilbert's algorithm can be used to solve this nearest point problem but is unreasonably slow. In this paper we present a modified version of Gilbert's algorithm for the fast computation of the support vector machine hyperplane. We then compare our algorithm with the nearest point algorithm and with sequential minimal optimization.
Shawn Martin
ICDM1
2005 Predicting protein-protein interactions using signature products
abstract
MOTIVATION: Proteome-wide prediction of protein-protein interaction is a difficult and important problem in biology. Although there have been recent advances in both experimental and computational methods for predicting protein-protein interactions, we are only beginning to see a confluence of these techniques. In this paper, we describe a very general, high-throughput method for predicting protein-protein interactions. Our method combines a sequence-based description of proteins with experimental information that can be gathered from any type of protein-protein interaction screen. The method uses a novel description of interacting proteins by extending the signature descriptor, which has demonstrated success in predicting peptide/protein binding interactions for individual proteins. This descriptor is extended to protein pairs by taking signature products. The signature product is implemented within a support vector machine classifier as a kernel function. RESULTS: We have applied our method to publicly available yeast, Helicobacter pylori, human and mouse datasets. We used the yeast and H.pylori datasets to verify the predictive ability of our method, achieving from 70 to 80% accuracy rates using 10-fold cross-validation. We used the human and mouse datasets to demonstrate that our method is capable of cross-species prediction. Finally, we reused the yeast dataset to explore the ability of our algorithm to predict domains. CONTACT: [email protected]
Shawn Martin, Diana C. Roe, Jean-Loup Faulon
Bioinform.1