Tatiana V. Karpinets

dblp:18/6740 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2012
0000-0002-3668-8148ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 100%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
knowledge base
0.112012
BESC knowledgebase public portal · Bioinform. 2012
Bioinformatics and computational biology › biological network › network biology
protein complex identification
0.112008
From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008
Bioinformatics and computational biology › protein analysis › protein-protein interaction › protein-protein interaction network analysis
protein-protein interaction network inference
0.112008
From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008
Bioinformatics and computational biology › genomics
toxicogenomics
0.012004
Tailored gene array databases: applications in mechanistic toxicology · Bioinform. 2004
Bioinformatics and computational biology › proteomics
mass spectrometry proteomics
0.012008
From pull-down data to protein interaction networks and complexes with biological relevance · Bioinform. 2008

Methods — techniques the papers use, named apart from their topics

data integration · 0.1knowledge-guided threshold selection · 0.1graph-theoretical clustering · 0.1co-purification pattern similarity · 0.1relational database design · 0.0
YearPublicationVenuePosition
2012 BESC knowledgebase public portal
abstract
UNLABELLED: The BioEnergy Science Center (BESC) is undertaking large experimental campaigns to understand the biosynthesis and biodegradation of biomass and to develop biofuel solutions. BESC is generating large volumes of diverse data, including genome sequences, omics data and assay results. The purpose of the BESC Knowledgebase is to serve as a centralized repository for experimentally generated data and to provide an integrated, interactive and user-friendly analysis framework. The Portal makes available tools for visualization, integration and analysis of data either produced by BESC or obtained from external resources. AVAILABILITY: http://besckb.ornl.gov.
Mustafa H. Syed, Tatiana V. Karpinets, Morey Parang, Michael R. Leuze, Doug Hyatt, Steven D. Brown, Steve Moulton, Michael D. Galloway, Edward C. Uberbacher
Bioinform.2
2008 Parallel, scalable, memory-efficient backtracking for combinatoria modeling of large-scale biological systems
abstract
Data-driven modeling of biological systems such as protein-protein interaction networks is data-intensive and combinatorially challenging. Backtracking can constrain a combinatorial search space. Yet, its recursive nature, exacerbated by data-intensity, limits its applicability for large-scale systems. Parallel, scalable, and memory-efficient backtracking is a promising approach. Parallel backtracking suffers from unbalanced loads. Load rebalancing via synchronization and data movement is prohibitively expensive. Balancing these discrepancies, while minimizing end-to-end execution time and memory requirements, is desirable. This paper introduces such a framework. Its scalability and efficiency, demonstrated on the maximal clique enumeration problem, are attributed to the proposed: (a) representation of search tree decomposition to enable parallelization; (b) depth-first parallel search to minimize memory requirement; (c) least stringent synchronization to minimize data movement; and (d) on-demand work stealing with stack splitting to minimize processors’ idle time. The applications of this framework to real biological problems related to bioethanol production are discussed.
Matthew C. Schmidt, Kevin Thomas 0002, Tatiana V. Karpinets, Nagiza F. Samatova
IPDPS4
2008 From pull-down data to protein interaction networks and complexes with biological relevance
abstract
Abstract Motivation: Recent improvements in high-throughput Mass Spectrometry (MS) technology have expedited genome-wide discovery of protein–protein interactions by providing a capability of detecting protein complexes in a physiological setting. Computational inference of protein interaction networks and protein complexes from MS data are challenging. Advances are required in developing robust and seamlessly integrated procedures for assessment of protein–protein interaction affinities, mathematical representation of protein interaction networks, discovery of protein complexes and evaluation of their biological relevance. Results: A multi-step but easy-to-follow framework for identifying protein complexes from MS pull-down data is introduced. It assesses interaction affinity between two proteins based on similarity of their co-purification patterns derived from MS data. It constructs a protein interaction network by adopting a knowledge-guided threshold selection method. Based on the network, it identifies protein complexes and infers their core components using a graph-theoretical approach. It deploys a statistical evaluation procedure to assess biological relevance of each found complex. On Saccharomyces cerevisiae pull-down data, the framework outperformed other more complicated schemes by at least 10% in F1-measure and identified 610 protein complexes with high-functional homogeneity based on the enrichment in Gene Ontology (GO) annotation. Manual examination of the complexes brought forward the hypotheses on cause of false identifications. Namely, co-purification of different protein complexes as mediated by a common non-protein molecule, such as DNA, might be a source of false positives. Protein identification bias in pull-down technology, such as the hydrophilic bias could result in false negatives. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Bing Zhang 0003, Tatiana V. Karpinets, Nagiza F. Samatova
Bioinform.3
2007 Multi-stage Framework to Infer Protein Functional Modules from Mass Spectrometry Pull-Down Data with Assessment of Biological Relevance
abstract
Protein functional modules are fundamental units in protein interaction networks. High-throughput Mass Spectrometry (MS) technology has become valuable for discovery of protein functional modules. Yet, their computational inference from MS pull-down data and biological significance evaluation are still challenging. This paper introduces an integrated multi-step framework for (1) assessing protein-protein interaction affinities, (2) constructing a genome-wide protein association map, (3) finding putative protein functional modules, and (4) evaluating their biological relevance. The protein affinity score utilizes co- purification pattern of two proteins and adopts an information theoretic-approach to build the protein affinity map. Putative protein modules are then derived using a graph-theoretical approach. A two-stage statistical procedure assesses biological relevance of identified modules. On Saccharomyces cerevisiae's pull-down data (Nature, vol. 415, pp. 141-7, 2002), the scoring scheme outperformed other methods by at least 10% in F1-measure, and statistical tests identified 489 protein modules enriched in all of three general GO categories with p-values less than 0.05.
Bing Zhang 0003, Tatiana V. Karpinets, Nagiza F. Samatova
BIBM3
2004 Tailored gene array databases: applications in mechanistic toxicology
abstract
MOTIVATION: The development of an annotated global database suitable for a wide range of investigations is a challenging and labor-intensive task. Thus, the development of databases tailored for specific applications remains necessary. For example, in the field of toxicology, no annotated gene array databases are now available that may assist in the correlation of changes in gene activity to cellular functions and processes associated with the toxic response. RESULTS: As an example of a tailored annotated database, an attempt was made to systematize available biological information on genes present on the Affymetrix Rat Toxicology U34 GeneChip, with a focus on how the gene products relate to liver cells and their response to chemical toxins. The information collected was imbedded in a local relational database to analyze data obtained in toxicological gene array experiments with hydrazine-exposed hepatocytes. The advantages and benefits of the tailored database in the biological interpretation of the results are demonstrated.
Tatiana V. Karpinets, Brent D. Foy, John M. Frazier
Bioinform.1