Mudita Singhal

dblp:75/765 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 1 first-authorArtificial intelligence and machine learning · 1Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 61% Visualization and visual analytics · 39%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 13 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
proteomics
0.232010
Machine learning based prediction for peptide drift times in ion mobility spectrometry · Bioinform. 2010
PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006
Multimedia analysis and retrieval
visual search
0.212014
Footprints: A Visual Search Tool that Supports Discovery and Coverage Tracking · IEEE Trans. Vis. Comput. Graph. 2014
Bioinformatics and computational biology › bioinformatics infrastructure
biological data management
0.112007
Enabling high-throughput data management for systems biology: The Bioinformatics Resource Manager · Bioinform. 2007
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
pathway simulation
0.112006
COPASI - a COmplex PAthway SImulator · Bioinform. 2006
Bioinformatics and computational biology › proteomics
peptide identification
0.112006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006
Bioinformatics and computational biology › systems biology › computational systems biology
systems biology modeling
0.112006
COPASI - a COmplex PAthway SImulator · Bioinform. 2006
Visualization and visual analytics
biological data visualization
0.112006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006
Visualization and visual analytics › visual analytics
visual analytics system
0.112014
Footprints: A Visual Search Tool that Supports Discovery and Coverage Tracking · IEEE Trans. Vis. Comput. Graph. 2014
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
0.012010
Machine learning based prediction for peptide drift times in ion mobility spectrometry · Bioinform. 2010
Bioinformatics and computational biology › genome annotation
functional annotation
0.012007
Enabling high-throughput data management for systems biology: The Bioinformatics Resource Manager · Bioinform. 2007
Bioinformatics and computational biology
genomics
0.012007
PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007
Bioinformatics and computational biology › genome annotation
prokaryotic genome annotation
0.012007
PQuad - a visual analysis platform for proteomic data exploration of microbial organisms · Bioinform. 2007
High-performance computing
high-throughput computing
0.012006
Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge · SC 2006

Methods — techniques the papers use, named apart from their topics

topic extraction · 0.4coordinated visualization · 0.4support vector regression · 0.1partial least squares regression · 0.1machine learning · 0.1visual analytics integration · 0.1multiresolution visualization · 0.1data reformatting · 0.1stochastic simulation algorithm · 0.1random number generation · 0.1hybrid deterministic-stochastic simulation · 0.1
YearPublicationVenuePosition
2014 Footprints: A Visual Search Tool that Supports Discovery and Coverage Tracking
abstract
Searching a large document collection to learn about a broad subject involves the iterative process of figuring out what to ask, filtering the results, identifying useful documents, and deciding when one has covered enough material to stop searching. We are calling this activity "discoverage," discovery of relevant material and tracking coverage of that material. We built a visual analytic tool called Footprints that uses multiple coordinated visualizations to help users navigate through the discoverage process. To support discovery, Footprints displays topics extracted from documents that provide an overview of the search space and are used to construct searches visuospatially. Footprints allows users to triage their search results by assigning a status to each document (To Read, Read, Useful), and those status markings are shown on interactive histograms depicting the user's coverage through the documents across dates, sources, and topics. Coverage histograms help users notice biases in their search and fill any gaps in their analytic process. To create Footprints, we used a highly iterative, user-centered approach in which we conducted many evaluations during both the design and implementation stages and continually modified the design in response to feedback.
Ellen Isaacs, Kelly Domico, Shane Ahern, Eugene Bart, Mudita Singhal
IEEE Trans. Vis. Comput. Graph.5
2010 Machine learning based prediction for peptide drift times in ion mobility spectrometry
abstract
MOTIVATION: Ion mobility spectrometry (IMS) has gained significant traction over the past few years for rapid, high-resolution separations of analytes based upon gas-phase ion structure, with significant potential impacts in the field of proteomic analysis. IMS coupled with mass spectrometry (MS) affords multiple improvements over traditional proteomics techniques, such as in the elucidation of secondary structure information, identification of post-translational modifications, as well as higher identification rates with reduced experiment times. The high throughput nature of this technique benefits from accurate calculation of cross sections, mobilities and associated drift times of peptides, thereby enhancing downstream data analysis. Here, we present a model that uses physicochemical properties of peptides to accurately predict a peptide's drift time directly from its amino acid sequence. This model is used in conjunction with two mathematical techniques, a partial least squares regression and a support vector regression setting. RESULTS: When tested on an experimentally created high confidence database of 8675 peptide sequences with measured drift times, both techniques statistically significantly outperform the intrinsic size parameters-based calculations, the currently held practice in the field, on all charge states (+2, +3 and +4). AVAILABILITY: The software executable, imPredict, is available for download from http:/omics.pnl.gov/software/imPredict.php CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Anuj R. Shah, Khushbu Agarwal, Erin S. Baker, Mudita Singhal, Anoop M. Mayampurath, Yehia M. Ibrahim, Lars J. Kangas, Matthew E. Monroe, Mikhail E. Belov, Gordon A. Anderson, Richard D. Smith
Bioinform.4
2008 An Extensible, Scalable Architecture for Managing Bioinformatics Data and Analyses
abstract
Systems biology research demands the availability of tools and technologies that span a comprehensive range of computational capabilities, including data management, transfer, processing, integration, and interpretation. To address these needs, we have created the bioinformatics resource manager (BRM), a scalable, flexible, and easy to use tool for biologists to undertake complex analyses. This paper describes the underlying software architecture of the BRM that integrates multiple commodity platforms to provide a highly extensible and scalable software infrastructure for bioinformatics. The architecture integrates a J2EE 3-tier application with an archival experimental data management system, the GAGGLE framework for desktop tool integration, and the MeDICi integration framework for high-throughput data analysis workflows. This architecture facilitates a systems biology software solution that enables the entire spectrum of scientific activities, from experimental data access to high throughput processing and analysis of data for biologists and experimental scientists.
Anuj R. Shah, Mudita Singhal, Tara D. Gibson, Chandrika Sivaramakrishnan, Katrina M. Waters, Ian Gorton
eScience2
2008 Network Inference Algorithms Elucidate Nrf2 Regulation of Mouse Lung Oxidative Stress
abstract
A variety of cardiovascular, neurological, and neoplastic conditions have been associated with oxidative stress, i.e., conditions under which levels of reactive oxygen species (ROS) are elevated over significant periods. Nuclear factor erythroid 2-related factor (Nrf2) regulates the transcription of several gene products involved in the protective response to oxidative stress. The transcriptional regulatory and signaling relationships linking gene products involved in the response to oxidative stress are, currently, only partially resolved. Microarray data constitute RNA abundance measures representing gene expression patterns. In some cases, these patterns can identify the molecular interactions of gene products. They can be, in effect, proxies for protein-protein and protein-DNA interactions. Traditional techniques used for clustering coregulated genes on high-throughput gene arrays are rarely capable of distinguishing between direct transcriptional regulatory interactions and indirect ones. In this study, newly developed information-theoretic algorithms that employ the concept of mutual information were used: the Algorithm for the Reconstruction of Accurate Cellular Networks (ARACNE), and Context Likelihood of Relatedness (CLR). These algorithms captured dependencies in the gene expression profiles of the mouse lung, allowing the regulatory effect of Nrf2 in response to oxidative stress to be determined more precisely. In addition, a characterization of promoter sequences of Nrf2 regulatory targets was conducted using a Support Vector Machine classification algorithm to corroborate ARACNE and CLR predictions. Inferred networks were analyzed, compared, and integrated using the Collective Analysis of Biological Interaction Networks (CABIN) plug-in of Cytoscape. Using the two network inference algorithms and one machine learning algorithm, a number of both previously known and novel targets of Nrf2 transcriptional activation were identified. Genes predicted as novel Nrf2 targets include Atf1, Srxn1, Prnp, Sod2, Als2, Nfkbib, and Ppp1r15b. Furthermore, microarray and quantitative RT-PCR experiments following cigarette-smoke-induced oxidative stress in Nrf2(+/+) and Nrf2(-/-) mouse lung affirmed many of the predictions made. Several new potential feed-forward regulatory loops involving Nrf2, Nqo1, Srxn1, Prdx1, Als2, Atf1, Sod1, and Park7 were predicted. This work shows the promise of network inference algorithms operating on high-throughput gene expression data in identifying transcriptional regulatory and other signaling relationships implicated in mammalian disease.
Ronald C. Taylor, George K. Acquaah-Mensah, Mudita Singhal, Deepti Malhotra, Shyam Biswal
PLoS Comput. Biol.3
2007 SEBINI-CABIN: An Analysis Pipeline for Biological Network Inference, with a Case Study in Protein-Protein Interaction Network Reconstruction
abstract
The Software Environment for Biological Network Inference (SEBINI) has been created to provide an interactive environment for the deployment and testing of network inference algorithms that use high-throughput expression data. Networks inferred from the SEBINI software platform can be further analyzed using the Collective Analysis of Biological Interaction Networks (CABIN), software that allows integration and analysis of protein- protein interaction and gene-to-gene regulatory evidence obtained from multiple sources. In this paper, we present a case study on the SEBINI and CABIN tools for protein-protein interaction network reconstruction. Incorporating the Bayesian Estimator of Protein-Protein Association Probabilities (BEPro) algorithm into the SEBINI toolkit, we have created a pipeline for structural inference and supplemental analysis of protein- protein interaction networks from sets of mass spectrometry bait-prey experiment data.
Ronald C. Taylor, Mudita Singhal, Don Simone Daly, Kelly Domico, Amanda M. White, Deanna L. Auberry, Kenneth J. Auberry, Brian Hooker, Gregory B. Hurst, Jason E. McDermott, W. Hayes McDonald, Dale Pelletier, Denise Schmoyer, William R. Cannon
ICMLA2
2007 Enabling high-throughput data management for systems biology: The Bioinformatics Resource Manager
abstract
UNLABELLED: The Bioinformatics Resource Manager (BRM) is a software environment that provides the user with data management, retrieval and integration capabilities. Designed in collaboration with biologists, BRM simplifies mundane analysis tasks of merging microarray and proteomic data across platforms, facilitates integration of users' data with functional annotation and interaction data from public sources and provides connectivity to visual analytic tools through reformatting of the data for easy import or dynamic launching capability. BRM is developed using Java and other open-source technologies for free distribution. AVAILABILITY: BRM, sample data sets and a user manual can be downloaded from http://www.sysbio.org/dataresources/brm.stm.
Anuj R. Shah, Mudita Singhal, Kyle R. Klicker, Eric G. Stephan, H. Steven Wiley, Katrina M. Waters
Bioinform.2
2007 PQuad - a visual analysis platform for proteomic data exploration of microbial organisms
abstract
UNLABELLED: The visual Platform for Proteomics Peptide and Protein data exploration (PQuad) is a multi-resolution environment that visually integrates genomic and proteomic data for prokaryotic systems, overlays categorical annotation and compares differential expression experiments. PQuad requires Java 1.5 and has been tested to run across different operating systems. AVAILABILITY: http://ncrr.pnl.gov/software.
Bobbie-Jo M. Webb-Robertson, Elena S. Peterson, Mudita Singhal, Kyle R. Klicker, Christopher S. Oehmen, Joshua N. Adkins, Susan L. Havre
Bioinform.3
2007 A domain-based approach to predict protein-protein interactions
abstract
BACKGROUND: Knowing which proteins exist in a certain organism or cell type and how these proteins interact with each other are necessary for the understanding of biological processes at the whole cell level. The determination of the protein-protein interaction (PPI) networks has been the subject of extensive research. Despite the development of reasonably successful methods, serious technical difficulties still exist. In this paper we present DomainGA, a quantitative computational approach that uses the information about the domain-domain interactions to predict the interactions between proteins. RESULTS: DomainGA is a multi-parameter optimization method in which the available PPI information is used to derive a quantitative scoring scheme for the domain-domain pairs. Obtained domain interaction scores are then used to predict whether a pair of proteins interacts. Using the yeast PPI data and a series of tests, we show the robustness and insensitivity of the DomainGA method to the selection of the parameter sets, score ranges, and detection rules. Our DomainGA method achieves very high explanation ratios for the positive and negative PPIs in yeast. Based on our cross-verification tests on human PPIs, comparison of the optimized scores with the structurally observed domain interactions obtained from the iPFAM database, and sensitivity and specificity analysis; we conclude that our DomainGA method shows great promise to be applicable across multiple organisms. CONCLUSION: We envision the DomainGA as a first step of a multiple tier approach to constructing organism specific PPIs. As it is based on fundamental structural information, the DomainGA approach can be used to create potential PPIs and the accuracy of the constructed interaction template can be further improved using complementary methods. Explanation ratios obtained in the reported test case studies clearly show that the false prediction rates of the template networks constructed using the DomainGA scores are reasonably low, and the erroneous predictions can be filtered further using supplementary approaches such as those based on literature search or other prediction methods.
Mudita Singhal, Haluk Resat
BMC Bioinform.1
2006 Analytics challenge - High-throughput visual analytics biological sciences: turning data into knowledge
abstract
For the SC|06 analytics challenge, we demonstrate an end-to-end solution for processing data produced by high-throughput mass spectrometry (MS)-based proteomics so biological hypotheses can be explored. This approach is based on a tool called the Bioinformatics Resource Manager (BRM) which will interact with high-performance architecture and experimental data sources to provide high-throughput analytics to a specific experimental dataset. Peptide identification was achieved by a high-performance code, Polygraph, which has been shown to scale well beyond 1000 processors. Visual analytics applications such as PQuad, Cytoscape, or others may be used to visualize protein identities in the context of pathways using data from public repositories such as Kyoto Encyclopedia of Genes and Genomes (KEGG). The end result was that a user can go from experimental spectra to pathway data in a single workflow reducing time-to-solution for analyzing biological data from weeks to minutes.
Christopher S. Oehmen, Lee Ann McCue, Joshua N. Adkins, Katrina M. Waters, Tim Carlson, William R. Cannon, Bobbie-Jo M. Webb-Robertson, Douglas J. Baxter, Elena S. Peterson, Mudita Singhal, Anuj R. Shah, Kyle R. Klicker
SC10
2006 COPASI - a COmplex PAthway SImulator
abstract
MOTIVATION: Simulation and modeling is becoming a standard approach to understand complex biochemical processes. Therefore, there is a big need for software tools that allow access to diverse simulation and modeling methods as well as support for the usage of these methods. RESULTS: Here, we present COPASI, a platform-independent and user-friendly biochemical simulator that offers several unique features. We discuss numerical issues with these features; in particular, the criteria to switch between stochastic and deterministic simulation methods, hybrid deterministic-stochastic methods, and the importance of random number generator numerical resolution in stochastic simulation. AVAILABILITY: The complete software is available in binary (executable) for MS Windows, OS X, Linux (Intel) and Sun Solaris (SPARC), as well as the full source code under an open source license from http://www.copasi.org.
Stefan Hoops, Sven Sahle, Ralph Gauges, Christine Lee, Jürgen Pahle, Natalia Simus, Mudita Singhal, Pedro Mendes 0001, Ursula Kummer
Bioinform.7
2004 PQuad: Visualization of Predicted Peptides and Proteins
Susan L. Havre, Mudita Singhal, Deborah A. Payne, Bobbie-Jo M. Webb-Robertson
IEEE Visualization2