VLDB 2026 Research / reviewers in the wild / expert
Shawn Martin
dblp:71/6817
· DBLP profile ↗
8ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Optimization for machine learning · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › drug discovery
drug-target interaction |
0.1 | 1 | 2008 | Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptor · Bioinform. 2008 |
Bioinformatics and computational biology › drug discovery
drug-target interaction prediction |
0.1 | 1 | 2008 | Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptor · Bioinform. 2008 |
Bioinformatics and computational biology › biological network › network biology › network inference › gene regulatory network inference
boolean network inference |
0.1 | 1 | 2007 | Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 1 | 2007 | Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007 |
Bioinformatics and computational biology › biological network › network biology › network inference
gene regulatory network inference |
0.1 | 1 | 2007 | Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007 |
Bioinformatics and computational biology › gene expression analysis › time-series gene expression analysis
time-course microarray analysis |
0.1 | 1 | 2007 | Boolean dynamics of genetic regulatory networks inferred from microarray time series data · Bioinform. 2007 |
Machine learning › Optimization for machine learning › regularized risk minimization
support vector machine training |
0.1 | 1 | 2005 | Training Support Vector Machines Using Gilbert's Algorithm · ICDM 2005 |
Bioinformatics and computational biology
protein-protein interaction prediction |
0.1 | 1 | 2005 | Predicting protein-protein interactions using signature products · Bioinform. 2005 |
Bioinformatics and computational biology
protein sequence analysis |
0.0 | 1 | 2005 | Predicting protein-protein interactions using signature products · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
signature molecular descriptor · 0.1machine learning · 0.1support vector regression · 0.1k-means clustering · 0.1support vector machine · 0.1signature product · 0.1sequential minimal optimization · 0.1kernel function · 0.1gilbert's algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Using Bipartite Anomaly Features for Cyber Security ApplicationsabstractIn this paper we use anomaly scores derived from a technique for bipartite graphs as features for a supervised machine learning algorithm for two cyber security problems: classifying Short Message Service (SMS) text messages as either spam or non-spam and detecting malicious lateral movement within a network. While disparate problems, both spam and lateral movement detection can be viewed as bipartite graphs and we can compute bipartite anomaly scores for each situation. The bipartite anomaly scores by themselves are not very predictive, but used as auxiliary features can boost the receiver operating characteristic (ROC) curve of a supervised classifier. We examine the UCI SMS Spam Collection Data Set for the SPAM problem and use an authentication graph from Los Alamos National Laboratory. We create features by dimensionality reduction through principal component analysis (PCA) on the message-term or user-computer matrix, and then augment those features with anomaly scores. By using the anomaly scores we are able to improve the area under the curve (AUC) for the receiver operating characteristic (ROC) up to 27.5% for the spam data and 21.4% for the authentication data. Eric L. Goodman, Joe Ingram, Shawn Martin, Dirk Grunwald |
ICMLA | 3 |
| 2011 | Non-manifold surface reconstruction from high-dimensional point cloud data
Shawn Martin, Jean-Paul Watson |
Comput. Geom. | 1 |
| 2008 | Genome scale enzyme-metabolite and drug-target interaction predictions using the signature molecular descriptorabstractAbstract Motivation: Identifying protein enzymatic or pharmacological activities are important areas of research in biology and chemistry. Biological and chemical databases are increasingly being populated with linkages between protein sequences and chemical structures. There is now sufficient information to apply machine-learning techniques to predict interactions between chemicals and proteins at a genome scale. Current machine-learning techniques use as input either protein sequences and structures or chemical information. We propose here a method to infer protein–chemical interactions using heterogeneous input consisting of both protein sequence and chemical information. Results: Our method relies on expressing proteins and chemicals with a common cheminformatics representation. We demonstrate our approach by predicting whether proteins can catalyze reactions not present in training sets. We also predict whether a given drug can bind a target, in the absence of prior binding information for that drug and target. Such predictions cannot be made with current machine-learning techniques requiring binding information for individual reactions or individual targets. Availability and Contact: For questions, paper reprints, please contact Jean-Loup Faulon at [email protected]. Additional information on the signature molecular descriptor and codes can be downloaded at: http://www.cs.sandia.gov/~jfaulon/publication-signature.html Supplementary information: Supplementary data are available at Bioinformatics online. Jean-Loup Faulon, Milind Misra, Shawn Martin, Kenneth L. Sale, Rajat Sapra |
Bioinform. | 3 |
| 2007 | Predicting building contamination using machine learningabstractPotential events involving biological or chemical contamination of buildings are of major concern in the area of homeland security. Tools are needed to provide rapid, on- site predictions of contaminant levels given only approximate measurements in limited locations throughout a building. In principal, such tools could use calculations based on physical process models to provide accurate predictions. In practice, however, physical process models are too complex and computationally costly to be used in a real-time scenario. In this paper, we investigate the feasibility of using machine learning to provide easily computed but approximate models that would be applicable in the field. We develop a machine learning method based on support vector machine regression and classification. We apply our method to problems of estimating contamination levels and contaminant source location. Shawn Martin, Sean McKenna |
ICMLA | 1 |
| 2007 | Boolean dynamics of genetic regulatory networks inferred from microarray time series dataabstractMOTIVATION: Methods available for the inference of genetic regulatory networks strive to produce a single network, usually by optimizing some quantity to fit the experimental observations. In this article we investigate the possibility that multiple networks can be inferred, all resulting in similar dynamics. This idea is motivated by theoretical work which suggests that biological networks are robust and adaptable to change, and that the overall behavior of a genetic regulatory network might be captured in terms of dynamical basins of attraction. RESULTS: We have developed and implemented a method for inferring genetic regulatory networks for time series microarray data. Our method first clusters and discretizes the gene expression data using k-means and support vector regression. We then enumerate Boolean activation-inhibition networks to match the discretized data. Finally, the dynamics of the Boolean networks are examined. We have tested our method on two immunology microarray datasets: an IL-2-stimulated T cell response dataset and a LPS-stimulated macrophage response dataset. In both cases, we discovered that many networks matched the data, and that most of these networks had similar dynamics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shawn Martin, Zhaoduo Zhang, Anthony Martino, Jean-Loup Faulon |
Bioinform. | 1 |
| 2006 | An Approximate Version of Kernel PCAabstractWe propose an analog of kernel principal component analysis (kernel PCA). Our algorithm is based on an approximation of PCA which uses Gram-Schmidt orthonormalization. We combine this approximation with support vector machine kernels to obtain a nonlinear generalization of PCA. By using our approximation to PCA we are able to provide a more easily computed (in the case of many data points) and readily interpretable version of kernel PCA. After demonstrating our algorithm on some examples, we explore its use in applications to fluid flow and microarray data Shawn Martin |
ICMLA | 1 |
| 2005 | Training Support Vector Machines Using Gilbert's AlgorithmabstractSupport vector machines are classifiers designed around the computation of an optimal separating hyperplane. This hyperplane is typically obtained by solving a constrained quadratic programming problem, but may also be located by solving a nearest point problem. Gilbert's algorithm can be used to solve this nearest point problem but is unreasonably slow. In this paper we present a modified version of Gilbert's algorithm for the fast computation of the support vector machine hyperplane. We then compare our algorithm with the nearest point algorithm and with sequential minimal optimization. Shawn Martin |
ICDM | 1 |
| 2005 | Predicting protein-protein interactions using signature productsabstractMOTIVATION: Proteome-wide prediction of protein-protein interaction is a difficult and important problem in biology. Although there have been recent advances in both experimental and computational methods for predicting protein-protein interactions, we are only beginning to see a confluence of these techniques. In this paper, we describe a very general, high-throughput method for predicting protein-protein interactions. Our method combines a sequence-based description of proteins with experimental information that can be gathered from any type of protein-protein interaction screen. The method uses a novel description of interacting proteins by extending the signature descriptor, which has demonstrated success in predicting peptide/protein binding interactions for individual proteins. This descriptor is extended to protein pairs by taking signature products. The signature product is implemented within a support vector machine classifier as a kernel function. RESULTS: We have applied our method to publicly available yeast, Helicobacter pylori, human and mouse datasets. We used the yeast and H.pylori datasets to verify the predictive ability of our method, achieving from 70 to 80% accuracy rates using 10-fold cross-validation. We used the human and mouse datasets to demonstrate that our method is capable of cross-species prediction. Finally, we reused the yeast dataset to explore the ability of our algorithm to predict domains. CONTACT: [email protected] Shawn Martin, Diana C. Roe, Jean-Loup Faulon |
Bioinform. | 1 |