EDBT 2026 Demo / reviewers in the wild / expert
Joseph White
dblp:86/6424
· DBLP profile ↗
8ranked-venue papers
0as first author
1since 2021 · last 2024
0000-0003-4914-9140ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 83% Computational science and engineering · 17% | |
| Software engineering, system software, and programming languages
1 paper |
Requirements engineering and software design · 100% |
Topics — the 6 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computational science and engineering › scientific data management
data standards |
0.1 | 1 | 2006 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 |
Bioinformatics and computational biology › ontology
ontology development |
0.1 | 1 | 2006 | The MGED Ontology: a resource for semantics-based description of microarray experiments · Bioinform. 2006 |
Bioinformatics and computational biology › sequence analysis › sequence clustering
EST clustering |
0.0 | 1 | 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003 |
Bioinformatics and computational biology › transcriptomics
transcript assembly |
0.0 | 1 | 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
consensus sequence generation |
0.0 | 1 | 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets · Bioinform. 2003 |
Methods — techniques the papers use, named apart from their topics
ontology engineering · 0.1parallel computing · 0.0pairwise sequence similarity · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Target Permutation for Feature Significance and Applications in Neural NetworksabstractStatistical techniques like generalized linear models have always been the indispensable tools for understanding relationships in data. However, over the last two decades, growth in data size and complexity has fueled the rise of powerful “black box” machine learning models (i.e., Deep Learning). Such models lack transparent mechanisms to evaluate the importance and contributions of individual features. This may be unacceptable ethically and/or legally. Moreover, the inability to precisely assess contributions from features complicates model validation and understanding. In this paper, we propose a permutation test for feature importance in differentiable models with a focus on neural networks. Unlike existing permutation-based methods, ours shuffles the target instead of the inputs. This change allows for simultaneous testing of all features and does not require the assumption of independence among inputs that is prevalent in current work in this area. Through extensive experiments, we empirically demonstrate that this permutation test can reveal highly nonlinear associations, is robust to multicollinearity among features, and can be used to filter unnecessary inputs, while preserving or improving models' predictive performance in both classification and regression. Sanad Biswas, Nina Grundlingh, Jonathan Boardman, Joseph White, Linh Le |
ICMLA | 4 |
| 2012 | Breaking Weak 1024-bit RSA Keys with CUDAabstractAn exploit involving the greatest common divisor (GCD) of RSA moduli was recently discovered [1]. This paper presents a tool that can efficiently and completely compare a large number of 1024-bit RSA public keys, and identify any keys that are susceptible to this weakness. NVIDIA's graphics processing units (GPU) and the CUDA massively-parallel programming model are powerful tools that can be used to accelerate this tool. Our method using CUDA has a measured performance speedup of 27.5 compared to a sequential CPU implementation, making it a more practical method to compare large sets of keys. A computation for finding GCDs between 200,000 keys, i.e., approximately 20 billion comparisons, was completed in 113 minutes, the equivalent of approximately 2.9 million 1024-bit GCD comparisons per second. Kerry Scharfglass, Darrin Weng, Joseph White, Chris Lupo |
PDCAT | 3 |
| 2012 | Zen Puzzle Garden is NP-complete
Robin Houston, Joseph White, Martyn Amos |
Inf. Process. Lett. | 2 |
| 2011 | GCOD - GeneChip Oncology DatabaseabstractBACKGROUND: DNA microarrays have become a nearly ubiquitous tool for the study of human disease, and nowhere is this more true than in cancer. With hundreds of studies and thousands of expression profiles representing the majority of human cancers completed and in public databases, the challenge has been effectively accessing and using this wealth of data. DESCRIPTION: To address this issue we have collected published human cancer gene expression datasets generated on the Affymetrix GeneChip platform, and carefully annotated those studies with a focus on providing accurate sample annotation. To facilitate comparison between datasets, we implemented a consistent data normalization and transformation protocol and then applied stringent quality control procedures to flag low-quality assays. CONCLUSION: The resulting resource, the GeneChip Oncology Database, is available through a publicly accessible website that provides several query options and analytical tools through an intuitive interface. Fenglong Liu, Joseph White, Corina Antonescu, John Quackenbush |
BMC Bioinform. | 2 |
| 2010 | Annotare - a tool for annotating high-throughput biomedical investigations and resulting dataabstractUNLABELLED: Computational methods in molecular biology will increasingly depend on standards-based annotations that describe biological experiments in an unambiguous manner. Annotare is a software tool that enables biologists to easily annotate their high-throughput experiments, biomaterials and data in a standards-compliant way that facilitates meaningful search and analysis. AVAILABILITY AND IMPLEMENTATION: Annotare is available from http://code.google.com/p/annotare/ under the terms of the open-source MIT License (http://www.opensource.org/licenses/mit-license.php). It has been tested on both Mac and Windows. Helen E. Parkinson, Tony Burdett, Emma Hastings, Junmin Liu, Michael Miller 0001, Rashmi Srinivasa, Joseph White, Alvis Brazma, Gavin Sherlock, Christian J. Stoeckert Jr., Catherine A. Ball |
Bioinform. | 8 |
| 2006 | The MGED Ontology: a resource for semantics-based description of microarray experimentsabstractMOTIVATION: The generation of large amounts of microarray data and the need to share these data bring challenges for both data management and annotation and highlights the need for standards. MIAME specifies the minimum information needed to describe a microarray experiment and the Microarray Gene Expression Object Model (MAGE-OM) and resulting MAGE-ML provide a mechanism to standardize data representation for data exchange, however a common terminology for data annotation is needed to support these standards. RESULTS: Here we describe the MGED Ontology (MO) developed by the Ontology Working Group of the Microarray Gene Expression Data (MGED) Society. The MO provides terms for annotating all aspects of a microarray experiment from the design of the experiment and array layout, through to the preparation of the biological sample and the protocols used to hybridize the RNA and analyze the data. The MO was developed to provide terms for annotating experiments in line with the MIAME guidelines, i.e. to provide the semantics to describe a microarray experiment according to the concepts specified in MIAME. The MO does not attempt to incorporate terms from existing ontologies, e.g. those that deal with anatomical parts or developmental stages terms, but provides a framework to reference terms in other ontologies and therefore facilitates the use of ontologies in microarray data annotation. AVAILABILITY: The MGED Ontology version.1.2.0 is available as a file in both DAML and OWL formats at http://mged.sourceforge.net/ontologies/index.php. Release notes and annotation examples are provided. The MO is also provided via the NCICB's Enterprise Vocabulary System (http://nciterms.nci.nih.gov/NCIBrowser/Dictionary.do). CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Patricia L. Whetzel, Helen E. Parkinson, Helen C. Causton, Liju Fan, Jennifer Fostel, Gilberto Fragoso, Laurence Game, Mervi Heiskanen, Norman Morrison, Philippe Rocca-Serra, Susanna-Assunta Sansone, Chris F. Taylor, Joseph White, Christian J. Stoeckert Jr. |
Bioinform. | 13 |
| 2006 | A simple spreadsheet-based, MIAME-supportive format for microarray data: MAGE-TABabstractBACKGROUND: Sharing of microarray data within the research community has been greatly facilitated by the development of the disclosure and communication standards MIAME and MAGE-ML by the MGED Society. However, the complexity of the MAGE-ML format has made its use impractical for laboratories lacking dedicated bioinformatics support. RESULTS: We propose a simple tab-delimited, spreadsheet-based format, MAGE-TAB, which will become a part of the MAGE microarray data standard and can be used for annotating and communicating microarray data in a MIAME compliant fashion. CONCLUSION: MAGE-TAB will enable laboratories without bioinformatics experience or support to manage, exchange and submit well-annotated microarray data in a standard format using a spreadsheet. The MAGE-TAB format is self-contained, and does not require an understanding of MAGE-ML or XML. Tim F. Rayner, Philippe Rocca-Serra, Paul T. Spellman, Helen C. Causton, Anna Farne, Ele Holloway, Rafael A. Irizarry, Junmin Liu, Donald Maier, Michael Miller 0001, Kjell Petersen, John Quackenbush, Gavin Sherlock, Christian J. Stoeckert Jr., Joseph White, Patricia L. Whetzel, Farrell Wymore, Helen E. Parkinson, Ugis Sarkans, Catherine A. Ball, Alvis Brazma |
BMC Bioinform. | 15 |
| 2003 | TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasetsabstractAbstract TGICL is a pipeline for analysis of large Expressed Sequence Tags (EST) and mRNA databases in which the sequences are first clustered based on pairwise sequence similarity, and then assembled by individual clusters (optionally with quality values) to produce longer, more complete consensus sequences. The system can run on multi-CPU architectures including SMP and PVM. Availability: http://www.tigr.org/tdb/tgi/software/ Contact: [email protected]; [email protected] * To whom correspondence should be addressed. Geo Pertea, Xiaoqiu Huang 0001, Valentin Antonescu, Razvan Sultana, Svetlana Karamycheva, Yuandan Lee, Joseph White, Foo Cheung, Babak Parvizi, Jennifer Tsai, John Quackenbush |
Bioinform. | 8 |