Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Joseph White

dblp:86/6424 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
1since 2021 · last 2024
0000-0003-4914-9140ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 83% Computational science and engineering · 17%
Software engineering, system software, and programming languages
1 paper
Requirements engineering and software design · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational science and engineering › scientific data management
data standards
0.112006
The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006
Bioinformatics and computational biology › ontology
ontology development
0.112006
The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006
Bioinformatics and computational biology › sequence analysis › sequence clustering
EST clustering
0.012003
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003
Bioinformatics and computational biology
sequence analysis
0.012003
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003
Bioinformatics and computational biology › transcriptomics
transcript assembly
0.012003
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003
Bioinformatics and computational biology › sequence analysis › sequence assembly
consensus sequence generation
0.012003
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003

Methods — techniques the papers use, named apart from their topics

ontology engineering · 0.1parallel computing · 0.0pairwise sequence similarity · 0.0
YearPublicationVenuePosition
2024 Target Permutation for Feature Significance and Applications in Neural Networks
abstract
Statistical techniques like generalized linear models have always been the indispensable tools for understanding relationships in data. However, over the last two decades, growth in data size and complexity has fueled the rise of powerful “black box” machine learning models (i.e., Deep Learning). Such models lack transparent mechanisms to evaluate the importance and contributions of individual features. This may be unacceptable ethically and/or legally. Moreover, the inability to precisely assess contributions from features complicates model validation and understanding. In this paper, we propose a permutation test for feature importance in differentiable models with a focus on neural networks. Unlike existing permutation-based methods, ours shuffles the target instead of the inputs. This change allows for simultaneous testing of all features and does not require the assumption of independence among inputs that is prevalent in current work in this area. Through extensive experiments, we empirically demonstrate that this permutation test can reveal highly nonlinear associations, is robust to multicollinearity among features, and can be used to filter unnecessary inputs, while preserving or improving models' predictive performance in both classification and regression.
Sanad Biswas, Nina Grundlingh, Jonathan Boardman, Joseph White, Linh Le
ICMLA4
2012 Breaking Weak 1024-bit RSA Keys with CUDA
abstract
An exploit involving the greatest common divisor (GCD) of RSA moduli was recently discovered [1]. This paper presents a tool that can efficiently and completely compare a large number of 1024-bit RSA public keys, and identify any keys that are susceptible to this weakness. NVIDIA's graphics processing units (GPU) and the CUDA massively-parallel programming model are powerful tools that can be used to accelerate this tool. Our method using CUDA has a measured performance speedup of 27.5 compared to a sequential CPU implementation, making it a more practical method to compare large sets of keys. A computation for finding GCDs between 200,000 keys, i.e., approximately 20 billion comparisons, was completed in 113 minutes, the equivalent of approximately 2.9 million 1024-bit GCD comparisons per second.
Kerry Scharfglass, Darrin Weng, Joseph White, Chris Lupo
PDCAT3
2012 Zen Puzzle Garden is NP-complete
Robin Houston, Joseph White, Martyn Amos
Inf. Process. Lett.2
2011 GCOD - GeneChip Oncology Database
abstract
BACKGROUND: DNA microarrays have become a nearly ubiquitous tool for the study of human disease, and nowhere is this more true than in cancer. With hundreds of studies and thousands of expression profiles representing the majority of human cancers completed and in public databases, the challenge has been effectively accessing and using this wealth of data. DESCRIPTION: To address this issue we have collected published human cancer gene expression datasets generated on the Affymetrix GeneChip platform, and carefully annotated those studies with a focus on providing accurate sample annotation. To facilitate comparison between datasets, we implemented a consistent data normalization and transformation protocol and then applied stringent quality control procedures to flag low-quality assays. CONCLUSION: The resulting resource, the GeneChip Oncology Database, is available through a publicly accessible website that provides several query options and analytical tools through an intuitive interface.
Fenglong Liu, Joseph White, Corina Antonescu, John Quackenbush
BMC Bioinform.2
2010 Annotare - a tool for annotating high-throughput biomedical investigations and resulting data
abstract
UNLABELLED: Computational methods in molecular biology will increasingly depend on standards-based annotations that describe biological experiments in an unambiguous manner. Annotare is a software tool that enables biologists to easily annotate their high-throughput experiments, biomaterials and data in a standards-compliant way that facilitates meaningful search and analysis. AVAILABILITY AND IMPLEMENTATION: Annotare is available from http://code.google.com/p/annotare/ under the terms of the open-source MIT License (http://www.opensource.org/licenses/mit-license.php). It has been tested on both Mac and Windows.
Helen E. Parkinson, Tony Burdett, Emma Hastings, Junmin Liu, Michael Miller 0001, Rashmi Srinivasa, Joseph White, Alvis Brazma, Gavin Sherlock, Christian J. Stoeckert Jr., Catherine A. Ball
Bioinform.8
2006 The MGED Ontology: a resource for semantics-based description of microarray experiments
abstract
MOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Patricia L. Whetzel, Helen E. Parkinson, Helen C. Causton, Liju Fan, Jennifer Fostel, Gilberto Fragoso, Laurence Game, Mervi Heiskanen, Norman Morrison, Philippe Rocca-Serra, Susanna-Assunta Sansone, Chris F. Taylor, Joseph White, Christian J. Stoeckert Jr.
Bioinform.13
2006 A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TAB
abstract
BACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML.
Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma
BMC Bioinform.15
2003 TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets
abstract
Abstract TGICL is a pipeline for analysis of large Expressed Sequence Tags (EST) and mRNA databases in which the sequences are first clustered based on pairwise sequence similarity, and then assembled by individual clusters (optionally with quality values) to produce longer, more complete consensus sequences. The system can run on multi-CPU architectures including SMP and PVM. Availability: http://www.tigr.org/tdb/tgi/software/ Contact: [email protected]; [email protected] * To whom correspondence should be addressed.
Geo Pertea, Xiaoqiu Huang 0001, Valentin Antonescu, Razvan Sultana, Svetlana Karamycheva, Yuandan Lee, Joseph White, Foo Cheung, Babak Parvizi, Jennifer Tsai, John Quackenbush
Bioinform.8