EDBT 2026 Demo / reviewers in the wild / expert
Martin H. Schultz
dblp:93/3072
· DBLP profile ↗
16ranked-venue papers
1as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8Systems, architecture and hardware · 6Theory of computation · 2 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 99% Computational science and engineering · 1% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 50% Graph data management · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Parallel and multicore computing · 37% High-performance computing · 32% Interconnection networks and networks-on-chip · 32% |
Topics — the 19 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › network bioinformatics › biological network analysis › network alignment
biological network alignment |
0.1 | 1 | 2007 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networks · Bioinform. 2007 |
Bioinformatics and computational biology
proteomics |
0.1 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Bioinformatics and computational biology › protein structure analysis › protein binding site analysis
binding site detection |
0.1 | 1 | 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge · Bioinform. 2006 |
Bioinformatics and computational biology › epigenomics
ChIP-chip analysis |
0.1 | 1 | 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge · Bioinform. 2006 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
comparative interactomics |
0.1 | 1 | 2006 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networks · Bioinform. 2006 |
Bioinformatics and computational biology
genomics |
0.1 | 1 | 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge · Bioinform. 2006 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
network analysis |
0.1 | 1 | 2006 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networks · Bioinform. 2006 |
Bioinformatics and computational biology
systems biology |
0.1 | 1 | 2006 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networks · Bioinform. 2006 |
Bioinformatics and computational biology › gene expression analysis › microarray data analysis
tiling array analysis |
0.1 | 1 | 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledge · Bioinform. 2006 |
Information retrieval › query reformulation
query expansion |
0.0 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Graph data management
RDF graph |
0.0 | 1 | 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomics · Bioinform. 2007 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 1992 | Efficient Parallel Programming with Linda · SC 1992 |
Interconnection networks and networks-on-chip › network topology › hypercubic networks
hypercube |
0.0 | 1 | 1988 | Topological properties of hypercubes · IEEE Trans. Computers 1988 |
Interconnection networks and networks-on-chip
network topology |
0.0 | 1 | 1988 | Topological properties of hypercubes · IEEE Trans. Computers 1988 |
Graph algorithms and graph theory
graph theory |
0.0 | 1 | 1988 | Topological properties of hypercubes · IEEE Trans. Computers 1988 |
High-performance computing
scientific computing systems |
0.0 | 1 | 1992 | Efficient Parallel Programming with Linda · SC 1992 |
High-performance computing › supercomputing
supercomputer performance evaluation |
0.0 | 1 | 1989 | Supercomputers in computational ocean acoustics · SC 1989 |
Parallel and multicore computing › multiprocessor system
loosely coupled multiprocessor |
0.0 | 1 | 1988 | Topological properties of hypercubes · IEEE Trans. Computers 1988 |
Parallel and multicore computing
parallel architecture |
0.0 | 1 | 1988 | Topological properties of hypercubes · IEEE Trans. Computers 1988 |
Methods — techniques the papers use, named apart from their topics
inverse document frequency · 0.1cosine similarity · 0.1RDF graph · 0.1network motif detection · 0.1maximum entropy sampling · 0.1hidden markov model · 0.1generalized HMM · 0.1clustering coefficient computation · 0.1alternating direction implicit method · 0.0topology mapping · 0.0graph-theoretic analysis · 0.0message passing · 0.0linda · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Complementary ensemble clustering of biomedical data
Samah Jamal Fodeh, Cynthia Brandt, Thaibinh Luong, Ali Haddad, Martin H. Schultz, Terrence Murphy, Michael Krauthammer |
J. Biomed. Informatics | 5 |
| 2007 | Leveraging the structure of the Semantic Web to enhance information retrieval for proteomicsabstractMOTIVATION: Proteomics researchers need to be able to quickly retrieve relevant information from the web and the biomedical literature. To improve information retrieval, we leverage the structure of the semantic web, developing an approach for joining it with the largely opposing paradigm of unsupervised web search. RESULTS: Our approach uses a Resource-Description-Framework (RDF) graph that inter-relates documents through their associated biological identifiers (e.g., protein ID). A search begins with a simple query term (UniProt identifier), which is expanded with terms extracted from documents in the RDF graph surrounding the query ("the subgraph"). We re-rank documents in the full corpus (e.g. all PubMed) by their cosine-similarity scores against a composite word-weight vector created from the subgraph. This vector is a weighted sum of individual word-weight vectors for documents at each node of the subgraph, taking into account the types of relationships between the central query identifier and the nodes connected to it. The computation also uses inverse document frequency (IDF) in a novel way to rescale the local word frequencies in the query's subgraph relative to that in other subgraphs. Applying our procedure to PubMed, we optimize weights for various relationships in the subgraph and benchmark overall performance in detail. Using a subgraph containing family relationships (from PFAM) results in a significant improvement in accuracy (as compared to not considering the subgraph in the search) when assessed against known relationships in the yeast literature. Moreover, we achieve this accuracy using only relatively simple and computationally efficient methods. Andrew K. Smith, Kei-Hoi Cheung, Michael Krauthammer, Martin H. Schultz, Mark Gerstein |
Bioinform. | 4 |
| 2007 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networksabstractBioinformatics (2006) 22(23), 2968–2970 The authors would like to apologize for the omission of a reference from the below paper. N-Browse was used and referred to in the following article: Lall,S., Grun,D., Krek,A., Chen,K., Wang,Y.L., Dewey,C.N., Sood,P., Colombo,T., Bray,N., Macmenamin,P., Kao,H.L., Gunsalus,K.C., Pachter,L., Piano,F. and Rajewsky,N. (2006) A genome-wide map of conserved microRNA targets in C. elegans. Current Biol., 16, 460–471. Also, it should be noted that the URL of N-Browse given in our paper is a temporary redirect. The official URL should be http://www.gnetbrowse.org/. Kevin Y. Yip, Haiyuan Yu, Philip M. Kim, Martin H. Schultz, Mark Gerstein |
Bioinform. | 4 |
| 2007 | LinkHub: a Semantic Web system that facilitates cross-database queries and information retrieval in proteomicsabstractBACKGROUND: A key abstraction in representing proteomics knowledge is the notion of unique identifiers for individual entities (e.g. proteins) and the massive graph of relationships among them. These relationships are sometimes simple (e.g. synonyms) but are often more complex (e.g. one-to-many relationships in protein family membership). RESULTS: We have built a software system called LinkHub using Semantic Web RDF that manages the graph of identifier relationships and allows exploration with a variety of interfaces. For efficiency, we also provide relational-database access and translation between the relational and RDF versions. LinkHub is practically useful in creating small, local hubs on common topics and then connecting these to major portals in a federated architecture; we have used LinkHub to establish such a relationship between UniProt and the North East Structural Genomics Consortium. LinkHub also facilitates queries and access to information and documents related to identifiers spread across multiple databases, acting as "connecting glue" between different identifier spaces. We demonstrate this with example queries discovering "interologs" of yeast protein interactions in the worm and exploring the relationship between gene essentiality and pseudogene content. We also show how "protein family based" retrieval of documents can be achieved. LinkHub is available at hub.gersteinlab.org and hub.nesg.org with supplement, database models and full-source code. CONCLUSION: LinkHub leverages Semantic Web standards-based integrated data to provide novel information retrieval to identifier-related documents through relational graph queries, simplifies and manages connections to major hubs such as UniProt, and provides useful interactive and query interfaces for exploring the integrated data. Andrew K. Smith, Kei-Hoi Cheung, Kevin Y. Yip, Martin H. Schultz, Mark Gerstein |
BMC Bioinform. | 4 |
| 2006 | A supervised hidden markov model framework for efficiently segmenting tiling array data in transcriptional and chIP-chip experiments: systematically incorporating validated biological knowledgeabstractMOTIVATION: Large-scale tiling array experiments are becoming increasingly common in genomics. In particular, the ENCODE project requires the consistent segmentation of many different tiling array datasets into 'active regions' (e.g. finding transfrags from transcriptional data and putative binding sites from ChIP-chip experiments). Previously, such segmentation was done in an unsupervised fashion mainly based on characteristics of the signal distribution in the tiling array data itself. Here we propose a supervised framework for doing this. It has the advantage of explicitly incorporating validated biological knowledge into the model and allowing for formal training and testing. METHODOLOGY: In particular, we use a hidden Markov model (HMM) framework, which is capable of explicitly modeling the dependency between neighboring probes and whose extended version (the generalized HMM) also allows explicit description of state duration density. We introduce a formal definition of the tiling-array analysis problem, and explain how we can use this to describe sampling small genomic regions for experimental validation to build up a gold-standard set for training and testing. We then describe various ideal and practical sampling strategies (e.g. maximizing signal entropy within a selected region versus using gene annotation or known promoters as positives for transcription or ChIP-chip data, respectively). RESULTS: For the practical sampling and training strategies, we show how the size and noise in the validated training data affects the performance of an HMM applied to the ENCODE transcriptional and ChIP-chip experiments. In particular, we show that the HMM framework is able to efficiently process tiling array data as well as or better than previous approaches. For the idealized sampling strategies, we show how we can assess their performance in a simulation framework and how a maximum entropy approach, which samples sub-regions with very different signal intensities, gives the maximally performing gold-standard. This latter result has strong implications for the optimum way medium-scale validation experiments should be carried out to verify the results of the genome-scale tiling array experiments. Jiang Du 0003, Joel S. Rozowsky, Jan O. Korbel, Zhengdong D. Zhang, Thomas E. Royce, Martin H. Schultz, Michael Snyder 0001, Mark Gerstein |
Bioinform. | 6 |
| 2006 | The tYNA platform for comparative interactomics: a web tool for managing, comparing and mining multiple networksabstractUNLABELLED: Biological processes involve complex networks of interactions between molecules. Various large-scale experiments and curation efforts have led to preliminary versions of complete cellular networks for a number of organisms. To grapple with these networks, we developed TopNet-like Yale Network Analyzer (tYNA), a Web system for managing, comparing and mining multiple networks, both directed and undirected. tYNA efficiently implements methods that have proven useful in network analysis, including identifying defective cliques, finding small network motifs (such as feed-forward loops), calculating global statistics (such as the clustering coefficient and eccentricity), and identifying hubs and bottlenecks. It also allows one to manage a large number of private and public networks using a flexible tagging system, to filter them based on a variety of criteria, and to visualize them through an interactive graphical interface. A number of commonly used biological datasets have been pre-loaded into tYNA, standardized and grouped into different categories. AVAILABILITY: The tYNA system can be accessed at http://networks.gersteinlab.org/tyna. The source code, JavaDoc API and WSDL can also be downloaded from the website. tYNA can also be accessed from the Cytoscape software using a plugin. Kevin Y. Yip, Haiyuan Yu, Philip M. Kim, Martin H. Schultz, Mark Gerstein |
Bioinform. | 4 |
| 2005 | Case Report: A High Productivity/Low Maintenance Approach to High-performance Computation for Biomedicine: Four Case StudiesabstractThe rapid advances in high-throughput biotechnologies such as DNA microarrays and mass spectrometry have generated vast amounts of data ranging from gene expression to proteomics data. The large size and complexity involved in analyzing such data demand a significant amount of computing power. High-performance computation (HPC) is an attractive and increasingly affordable approach to help meet this challenge. There is a spectrum of techniques that can be used to achieve computational speedup with varying degrees of impact in terms of how drastic a change is required to allow the software to run on an HPC platform. This paper describes a high- productivity/low-maintenance (HP/LM) approach to HPC that is based on establishing a collaborative relationship between the bioinformaticist and HPC expert that respects the former's codes and minimizes the latter's efforts. The goal of this approach is to make it easy for bioinformatics researchers to continue to make iterative refinements to their programs, while still being able to take advantage of HPC. The paper describes our experience applying these HP/LM techniques in four bioinformatics case studies: (1) genome-wide sequence comparison using Blast, (2) identification of biomarkers based on statistical analysis of large mass spectrometry data sets, (3) complex genetic analysis involving ordinal phenotypes, (4) large-scale assessment of the effect of possible errors in analyzing microarray data. The case studies illustrate how the HP/LM approach can be applied to a range of representative bioinformatics applications and how the approach can lead to significant speedup of computationally intensive bioinformatics applications, while making only modest modifications to the programs themselves. Nicholas Carriero, Michael V. Osier, Kei-Hoi Cheung, Perry L. Miller, Mark Gerstein, Hongyu Zhao 0003, Baolin Wu, Scott A. Rifkin, Joseph T. Chang, Heping Zhang, Kevin P. White, Kenneth R. Williams, Martin H. Schultz |
J. Am. Medical Informatics Assoc. | 13 |
| 2000 | Graphically-enabled integration of bioinformatics tools allowing parallel execution
Kei-Hoi Cheung, Perry L. Miller, Andrew H. Sherman, Stephen B. Weston, Eric Stratmann, Martin H. Schultz, Michael Snyder 0001 |
AMIA | 6 |
| 1998 | First- and Second-Order Diffusive Methods for Rapid, Coarse, Distributed Load Balancing
S. Muthukrishnan 0001, Bhaskar Ghosh, Martin H. Schultz |
Theory Comput. Syst. | 3 |
| 1996 | First and Second Order Diffusive Methods for Rapid, Coarse, Distributed Load Balancing (Extended Abstract)abstractWe consider the following general problem modeling performs coarse load balancing rapidly and can be as such used in a number of applications.In all load balancing scenarios there is a tradeoff between the time invested in balancing the load and the speedup achieved as a result of a well-balanced load.In this paper, we focus only on coarse balancing, that is, decreasing large imbalances appreciably.To formalize this concept, we define the potenttal of the graph after step t to be & = ~,~v (wi -TD)z where the average weight u = (~z wi)/]Vl; the initial potential is denoted @o.Note that q$t = O if and only if Wi = D for all i and ~t >0 otherwise.Thus the larger q$t is, larger the imbalance.Throughout, we are concerned with c-balancing, that is, having the final potential to be no more than CC#IO, for a fraction e and for large q$o.lt is c-balancing that is of paramount importance in practice 2. Bhaskar Ghosh, S. Muthukrishnan 0001, Martin H. Schultz |
SPAA | 3 |
| 1992 | Efficient Parallel Programming with LindaabstractA number of computer scientists have contended that Linda cannot possibly be implemented efficiently on distributed memory machines because there is simply too much overhead. They believe that Linda will not be able to compete with message passing on such machines even for solving computationally intensive problems. The authors address this claim by discussing C-Linda's performance in solving a particular scientific computing problem, the shallow water equations. They have implemented and evaluated the performance of the Linda program on a variety of machines. They present results for shared memory machines (Sequent Symmetry and the Encore Multimax), for distributed memory machines (iPSC/2 and iPSC/860 hypercubes), and for a network of Sparcstations connected by an Ethernet. The same Linda program was executed on all these machines and its performance was evaluated and compared to that of implementations using alternative methods available on all machines. In the authors' experience, the Linda program has generally been easier and more convenient to write than the native versions for each machine.> Ashish Deshpande, Martin H. Schultz |
SC | 2 |
| 1989 | Supercomputers in computational ocean acousticsabstractIn this paper, we report on some computational experience in solving ocean acoustic propagation problems in three dimensions on supercomputers. The underlying Helmholtz equation is transformed into a parabolic-type equation in the Lee-Saad-Schultz model [5], which has a natural alternating direction implicit (ADI) implementation. We give estimates of the computing power required to solve problems with realistic sound velocity profiles. We then give performance results for the CRAY X-MP and for the computational kernel on the Intel hypercube (iPSC/2). We conclude with some remarks about architectural enhancements that would be beneficial to our application. Ding Lee, Martin H. Schultz, Faisal Saied |
SC | 2 |
| 1989 | Data Communication in HypercubesabstractIn this paper we consider several algorithms for exchanging data among processors in a hypercube network. The data transfer problems considered are those arising from classical numerical algorithms such as Gaussian elimination, conjugate gradient methods, and the N-body problem. We propose some estimates for the timings of the various algorithms which reveal that multiprocessors based on the hypercube topology can be very efficient in performing data exchange operations. Yousef Saad, Martin H. Schultz |
J. Parallel Distributed Comput. | 2 |
| 1989 | Data communication in parallel architectures
Yousef Saad, Martin H. Schultz |
Parallel Comput. | 2 |
| 1988 | Topological properties of hypercubesabstractThe n-dimensional hypercube is a highly concurrent loosely coupled multiprocessor based on the binary n-cube topology. Machines based on the hypercube topology have been advocated as ideal parallel architectures for their powerful interconnection features. The authors examine the hypercube from the graph-theory point of view and consider those features that make its connectivity so appealing. Among other things, they propose a theoretical characterization of the n-cube as a graph and and show how to map various other topologies into a hypercube.> Yousef Saad, Martin H. Schultz |
IEEE Trans. Computers | 2 |
| 1972 | Discrete Tchebycheff Approximation for Multivariate Splines
Martin H. Schultz |
J. Comput. Syst. Sci. | 1 |