EDBT 2026 Demo / reviewers in the wild / expert
Linda R. Petzold
dblp:03/3669 · also Linda Ruth Petzold
· DBLP profile ↗
40ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0001-6251-6078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 8 since 2021Artificial intelligence and machine learning · 12 · 8 since 2021Systems, architecture and hardware · 6Databases, data management, data science and information retrieval · 4 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models ReasoningabstractInstruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs). While coding data is known to boost LLM reasoning abilities during pretraining, its role in activating internal reasoning capacities during IFT remains understudied. This paper investigates a key question: How does coding data impact LLMs' reasoning capacities during IFT stage? To explore this, we thoroughly examine the impact of coding data across different coding data proportions, model families, sizes, and reasoning domains, from various perspectives. Specifically, we create three IFT datasets with increasing coding data proportions, fine-tune six LLM backbones across different families and scales on these datasets, evaluate the tuned models' performance across twelve tasks in three reasoning domains, and analyze the outcomes from three broad-to-granular perspectives: overall, domain-level, and task-specific. Our holistic analysis provides valuable insights into each perspective. First, coding data tuning enhances the overall reasoning capabilities of LLMs across different model families and scales. Moreover, while the impact of coding data varies by domain, it shows consistent trends within each domain across different model families and scales. Additionally, coding data generally provides comparable task-specific benefits across model families, with optimal proportions in IFT datasets being task-dependent. Xinlu Zhang, Zhiyu Chen 0002, Xi Ye 0003, Xianjun Yang, Lichang Chen, William Yang Wang, Linda R. Petzold |
AAAI | 7 |
| 2024 | DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextabstractLarge language models (LLMs) have notably enhanced the fluency and diversity of machine-generated text. However, this progress also presents a significant challenge in detecting the origin of a given text, and current research on detection methods lags behind the rapid evolution of LLMs. Conventional training-based methods have limitations in flexibility, particularly when adapting to new domains, and they often lack explanatory power. To address this gap, we propose a novel training-free detection strategy called Divergent N-Gram Analysis (DNA-GPT). Given a text, we first truncate it in the middle and then use only the preceding portion as input to the LLMs to regenerate the new remaining parts. By analyzing the differences between the original and new remaining parts through N-gram analysis in black-box or probability divergence in white-box, we can clearly illustrate significant discrepancies between machine-generated and human-written text. We conducted extensive experiments on the most advanced LLMs from OpenAI, including text-davinci-003, GPT-3.5-turbo, and GPT-4, as well as open-source models such as GPT-NeoX-20B and LLaMa-13B. Results show that our zero-shot approach exhibits state-of-the-art performance in distinguishing between human and GPT-generated text on four English and one German dataset, outperforming OpenAI's own classifier, which is trained on millions of text. Additionally, our methods provide reasonable explanations and evidence to support our claim, which is a unique feature of explainable detection. Our method is also robust under the revised text attack and can additionally solve model sourcing. Xianjun Yang, Wei Cheng 0002, Linda R. Petzold, William Yang Wang |
ICLR | 4 |
| 2024 | Enhancing Small Medical Learners with Privacy-preserving Contextual PromptingabstractLarge language models (LLMs) demonstrate remarkable medical expertise, but data privacy concerns impede their direct use in healthcare environments. Although offering improved data privacy protection, domain-specific small language models (SLMs) often underperform LLMs, emphasizing the need for methods that reduce this performance gap while alleviating privacy concerns. In this paper, we present a simple yet effective method that harnesses LLMs' medical proficiency to boost SLM performance in medical tasks under $privacy-restricted$ scenarios. Specifically, we mitigate patient privacy issues by extracting keywords from medical data and prompting the LLM to generate a medical knowledge-intensive context by simulating clinicians' thought processes. This context serves as additional input for SLMs, augmenting their decision-making capabilities. Our method significantly enhances performance in both few-shot and full training settings across three medical knowledge-intensive tasks, achieving up to a 22.57% increase in absolute accuracy compared to SLM fine-tuning without context, and sets new state-of-the-art results in two medical tasks within privacy-restricted scenarios. Further out-of-domain testing and experiments in two general domain datasets showcase its generalizability and broad applicability. Xinlu Zhang, Xianjun Yang, Chenxin Tian, Yao Qin 0001, Linda R. Petzold |
ICLR | 6 |
| 2024 | Bayesian polynomial neural networks and polynomial neural ordinary differential equationsabstractSymbolic regression with polynomial neural networks and polynomial neural ordinary differential equations (ODEs) are two recent and powerful approaches for equation recovery of many science and engineering problems. However, these methods provide point estimates for the model parameters and are currently unable to accommodate noisy data. We address this challenge by developing and validating the following Bayesian inference methods: the Laplace approximation, Markov Chain Monte Carlo (MCMC) sampling methods, and variational inference. We have found the Laplace approximation to be the best method for this class of problems. Our work can be easily extended to the broader class of symbolic neural networks to which the polynomial neural network belongs. Colby Fronk, Jaewoong Yun, Linda R. Petzold |
PLoS Comput. Biol. | 4 |
| 2024 | An empirical study on the robustness of the segment anything model (SAM)abstractThe Segment Anything Model (SAM) is a foundation model for general image segmentation. Although it exhibits impressive performance predominantly on natural images, understanding its robustness against various image perturbations and domains is critical for real-world applications where such challenges frequently arise. In this study we conduct a comprehensive robustness investigation of SAM under diverse real-world conditions. Our experiments encompass a wide range of image perturbations. Our experimental results demonstrate that SAM’s performance generally declines under perturbed images, with varying degrees of vulnerability across different perturbations. By customizing prompting techniques and leveraging domain knowledge based on the unique characteristics of each dataset, the model’s resilience to these perturbations can be enhanced, addressing dataset-specific challenges. This work sheds light on the limitations and strengths of SAM in real-world applications, promoting the development of more robust and versatile image segmentation solutions. Our code is available at https://github.com/EternityYW/SAM-Robustness/. Yuqing Wang 0004, Yun Zhao 0001, Linda R. Petzold |
Pattern Recognit. | 3 |
| 2023 | Few-Shot Document-Level Event Argument ExtractionabstractEvent argument extraction (EAE) has been well studied at the sentence level but under-explored at the document level.In this paper, we study to capture event arguments that actually spread across sentences in documents.Prior works usually assume full access to rich document supervision, ignoring the fact that the available argument annotation is usually limited.To fill this gap, we present FewDocAE, a Few-Shot Document-Level Event Argument Extraction benchmark, based on the existing documentlevel event extraction dataset.We first define the new problem and reconstruct the corpus by a novel N -Way-D-Doc sampling instead of the traditional N -Way-K-Shot strategy.Then we adjust the current document-level neural models into the few-shot setting to provide baseline results under in-and cross-domain settings.Since the argument extraction depends on the context from multiple sentences and the learning process is limited to very few examples, we find this novel task to be very challenging with substantively low performance.Considering FewDocAE is closely related to practical use under low-resource regimes, we hope this benchmark encourages more research in this direction.Our data and codes will be available online 1 . Xianjun Yang, Linda R. Petzold |
ACL (1) | 3 |
| 2023 | Improving Medical Predictions by Irregular Multimodal Electronic Health Records ModelingabstractHealth conditions among patients in intensive care units (ICUs) are monitored via electronic health records (EHRs), composed of numerical time series and lengthy clinical note sequences, both taken at $\textit{irregular}$ time intervals. Dealing with such irregularity in every modality, and integrating irregularity into multimodal representations to improve medical predictions, is a challenging problem. Our method first addresses irregularity in each single modality by (1) modeling irregular time series by dynamically incorporating hand-crafted imputation embeddings into learned interpolation embeddings via a gating mechanism, and (2) casting a series of clinical note representations as multivariate irregular time series and tackling irregularity via a time attention mechanism. We further integrate irregularity in multimodal fusion with an interleaved attention mechanism across temporal steps. To the best of our knowledge, this is the first work to thoroughly model irregularity in multimodalities for improving medical predictions. Our proposed methods for two medical prediction tasks consistently outperforms state-of-the-art (SOTA) baselines in each single modality and multimodal fusion scenarios. Specifically, we observe relative improvements of 6.5%, 3.6%, and 4.3% in F1 for time series, clinical notes, and multimodal fusion, respectively. These results demonstrate the effectiveness of our methods and the importance of considering irregularity in multimodal EHRs. Xinlu Zhang, Zhiyu Chen 0002, Xifeng Yan, Linda R. Petzold |
ICML | 5 |
| 2022 | Robust and integrative Bayesian neural networks for likelihood-free parameter inferenceabstractState-of-the-art neural network-based methods for learning summary statistics have delivered promising results for simulation-based likelihood-free parameter inference. Existing approaches for learning summarizing networks are mainly based on deterministic neural networks, and do not take network prediction uncertainty into account. This work proposes a robust integrated approach that learns summary statistics using Bayesian neural networks, and produces a proposal posterior density using categorical distributions. An adaptive sampling scheme selects simulation locations to efficiently and iteratively refine the predictive proposal posterior of the network conditioned on observations. This allows for more efficient and robust convergence on comparatively large prior spaces. The approximated proposal posterior can then either be processed through a correction mechanism, or be used in conjunction with a density estimator to arrive at the true posterior. We demonstrate our approach on benchmark examples. Fredrik Wrede, Robin Eriksson, Richard M. Jiang 0002, Linda R. Petzold, Stefan Engblom, Andreas Hellander |
IJCNN | 4 |
| 2022 | Identification of dynamic mass-action biochemical reaction networks using sparse Bayesian methodsabstractIdentifying the reactions that govern a dynamical biological system is a crucial but challenging task in systems biology. In this work, we present a data-driven method to infer the underlying biochemical reaction system governing a set of observed species concentrations over time. We formulate the problem as a regression over a large, but limited, mass-action constrained reaction space and utilize sparse Bayesian inference via the regularized horseshoe prior to produce robust, interpretable biochemical reaction networks, along with uncertainty estimates of parameters. The resulting systems of chemical reactions and posteriors inform the biologist of potentially several reaction systems that can be further investigated. We demonstrate the method on two examples of recovering the dynamics of an unknown reaction system, to illustrate the benefits of improved accuracy and information obtained. Richard M. Jiang 0002, Fredrik Wrede, Andreas Hellander, Linda R. Petzold |
PLoS Comput. Biol. | 5 |
| 2021 | Domain Adaptation for Trauma Mortality Prediction in EHRs with Feature DisparityabstractTrauma mortality prediction from electronic health records (EHRs) with machine learning models has received growing attention in medical fields, but EHRs in different hospitals and sub-medical domain populations are often scarce due to expensive collection processes or privacy issues. Domain Adaptation (DA) has emerged as a promising approach in computer vision and natural language processing to improve model performance in small data regimes by leveraging domain-invariant knowledge learned from a different yet related large source dataset. However, its applicability in trauma mortality prediction is challenging since EHRs collected from different hospital systems encounter feature disparity, i.e. distinct features between the source and target domain data. This paper demonstrates the effectiveness of three DA techniques in trauma mortality prediction, with a private encoding strategy that maps EHRs in both source and target domains with different raw features into the same latent space to alleviate feature disparity issues. Our experimental results on two real-world EHR datasets with various training data scenarios show that DA can improve mortality prediction consistently and significantly with private encoding. Finally, an ablation study manifests the importance of modeling feature disparity in DA, and 2-d t-SNE analysis explains its effectiveness. Xinlu Zhang, Shiyang Liy, Zhuowei Cheng, Rachael Callcut, Linda R. Petzold |
BIBM | 5 |
| 2021 | An Analysis of Relation Extraction within Sentences from Wet Lab ProtocolsabstractWet lab protocols (WLPs) are sets of instructions written in domain-specific natural language for step-by-step biological experimental processes. There have been efforts to annotate WLPs for shallow semantic parsing to enable reproducible procedures, text mining, and automatic conversion into a machine-readable format. However, current methods have not fully exploited the relation extraction sub-task on the protocol corpus. Neural approaches have the potential to deal with the various noise and in-domain jargon in the texts. To explore the viability of neural methods for this task, we perform a thorough analysis of both graph and nongraph neural approaches. We find that both graph neural networks with generated parameters (GP-GNNs) and Context-Aware models show advantages in relation extraction and are well suited to our goal. Specifically, the GP-GNNs and Context-Aware models demonstrate similar performance on all three WLPs datasets when the full training set is used, both outperforming the previous best results significantly. This can be explained by the observation that considering multiple relations in a sentence enhances the predictive ability. In addition, our extensive experiments demonstrate that the Context-Aware approach in particular can achieve good results even with a limited amount of training data, providing new insights for low-resource scenarios. Xianjun Yang, Xinlu Zhang, Julia Zuo, Stephen D. Wilson, Linda R. Petzold |
IEEE BigData | 5 |
| 2021 | Epidemiological modeling in StochSS Live!abstractSUMMARY: We present StochSS Live!, a web-based service for modeling, simulation and analysis of a wide range of mathematical, biological and biochemical systems. Using an epidemiological model of COVID-19, we demonstrate the power of StochSS Live! to enable researchers to quickly develop a deterministic or a discrete stochastic model, infer its parameters and analyze the results. AVAILABILITY AND IMPLEMENTATION: StochSS Live! is freely available at https://live.stochss.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard M. Jiang 0002, Bruno Jacob, Matthew Geiger, Sean Matthew, Bryan Rumsey, Fredrik Wrede, Tau-Mu Yi, Brian Drawert, Andreas Hellander, Linda R. Petzold |
Bioinform. | 11 |
| 2021 | Associations of longitudinal D-Dimer and Factor II on early trauma survival riskabstractBACKGROUND: Trauma-induced coagulopathy (TIC) is a disorder that occurs in one-third of severely injured trauma patients, manifesting as increased bleeding and a 4X risk of mortality. Understanding the mechanisms driving TIC, clinical risk factors are essential to mitigating this coagulopathic bleeding and is therefore essential for saving lives. In this retrospective, single hospital study of 891 trauma patients, we investigate and quantify how two prominently described phenotypes of TIC, consumptive coagulopathy and hyperfibrinolysis, affect survival odds in the first 25 h, when deaths from TIC are most prevalent. METHODS: We employ a joint survival model to estimate the longitudinal trajectories of the protein Factor II (% activity) and the log of the protein fragment D-Dimer ([Formula: see text]g/ml), representative biomarkers of consumptive coagulopathy and hyperfibrinolysis respectively, and tie them together with patient outcomes. Joint models have recently gained popularity in medical studies due to the necessity to simultaneously track continuously measured biomarkers as a disease evolves, as well as to associate them with patient outcomes. In this work, we estimate and analyze our joint model using Bayesian methods to obtain uncertainties and distributions over associations and trajectories. RESULTS: We find that a unit increase in log D-Dimer increases the risk of mortality by 2.22 [1.57, 3.28] fold while a unit increase in Factor II only marginally decreases the risk of mortality by 0.94 [0.91,0.96] fold. This suggests that, while managing consumptive coagulopathy and hyperfibrinolysis both seem to affect survival odds, the effect of hyperfibrinolysis is much greater and more sensitive. Furthermore, we find that the longitudinal trajectories, controlling for many fixed covariates, trend differently for different patients. Thus, a more personalized approach is necessary when considering treatment and risk prediction under these phenotypes. CONCLUSION: This study reinforces the finding that hyperfibrinolysis is linked with poor patient outcomes regardless of factor consumption levels. Furthermore, it quantifies the degree to which measured D-Dimer levels correlate with increased risk. The single hospital, retrospective nature can be understood to specify the results to this particular hospital's patients and protocol in treating trauma patients. Expanding to a multi-hospital setting would result in better estimates about the underlying nature of consumptive coagulopathy and hyperfibrinolysis with survival, regardless of protocol. Individual trajectories obtained with these estimates can be used to provide personalized dynamic risk prediction when making decisions regarding management of blood factors. Richard M. Jiang 0002, Arya A. Pourzanjani, Mitchell J. Cohen, Linda R. Petzold |
BMC Bioinform. | 4 |
| 2021 | Accelerated regression-based summary statistics for discrete stochastic systems via approximate simulatorsabstractBACKGROUND: Approximate Bayesian Computation (ABC) has become a key tool for calibrating the parameters of discrete stochastic biochemical models. For higher dimensional models and data, its performance is strongly dependent on having a representative set of summary statistics. While regression-based methods have been demonstrated to allow for the automatic construction of effective summary statistics, their reliance on first simulating a large training set creates a significant overhead when applying these methods to discrete stochastic models for which simulation is relatively expensive. In this τ work, we present a method to reduce this computational burden by leveraging approximate simulators of these systems, such as ordinary differential equations and τ-Leaping approximations. RESULTS: We have developed an algorithm to accelerate the construction of regression-based summary statistics for Approximate Bayesian Computation by selectively using the faster approximate algorithms for simulations. By posing the problem as one of ratio estimation, we use state-of-the-art methods in machine learning to show that, in many cases, our algorithm can significantly reduce the number of simulations from the full resolution model at a minimal cost to accuracy and little additional tuning from the user. We demonstrate the usefulness and robustness of our method with four different experiments. CONCLUSIONS: We provide a novel algorithm for accelerating the construction of summary statistics for stochastic biochemical systems. Compared to the standard practice of exclusively training from exact simulator samples, our method is able to dramatically reduce the number of required calls to the stochastic simulator at a minimal loss in accuracy. This can immediately be implemented to increase the overall speed of the ABC workflow for estimating parameters in complex systems. Richard M. Jiang 0002, Fredrik Wrede, Andreas Hellander, Linda R. Petzold |
BMC Bioinform. | 5 |
| 2021 | Coordinating cell polarization and morphogenesis through mechanical feedbackabstractMany cellular processes require cell polarization to be maintained as the cell changes shape, grows or moves. Without feedback mechanisms relaying information about cell shape to the polarity molecular machinery, the coordination between cell polarization and morphogenesis, movement or growth would not be possible. Here we theoretically and computationally study the role of a genetically-encoded mechanical feedback (in the Cell Wall Integrity pathway) as a potential coordination mechanism between cell morphogenesis and polarity during budding yeast mating projection growth. We developed a coarse-grained continuum description of the coupled dynamics of cell polarization and morphogenesis as well as 3D stochastic simulations of the molecular polarization machinery in the evolving cell shape. Both theoretical approaches show that in the absence of mechanical feedback (or in the presence of weak feedback), cell polarity cannot be maintained at the projection tip during growth, with the polarization cap wandering off the projection tip, arresting morphogenesis. In contrast, for mechanical feedback strengths above a threshold, cells can robustly maintain cell polarization at the tip and simultaneously sustain mating projection growth. These results indicate that the mechanical feedback encoded in the Cell Wall Integrity pathway can provide important positional information to the molecular machinery in the cell, thereby enabling the coordination of cell polarization and morphogenesis. Samhita P. Banavar, Michael Trogdon, Brian Drawert, Tau-Mu Yi, Linda R. Petzold, Otger Campàs |
PLoS Comput. Biol. | 5 |
| 2019 | A Minimum Free Energy Model of Motor LearningabstractEven highly trained behaviors demonstrate variability, which is correlated with performance on current and future tasks. An objective of motor learning that is general enough to explain these phenomena has not been precisely formulated. In this six-week longitudinal learning study, participants practiced a set of motor sequences each day, and neuroimaging data were collected on days 1, 14, 28, and 42 to capture the neural correlates of the learning process. In our analysis, we first modeled the underlying neural and behavioral dynamics during learning. Our results demonstrate that the densities of whole-brain response, task-active regional response, and behavioral performance evolve according to a Fokker-Planck equation during the acquisition of a motor skill. We show that this implies that the brain concurrently optimizes the entropy of a joint density over neural response and behavior (as measured by sampling over multiple trials and subjects) and the expected performance under this density; we call this formulation of learning minimum free energy learning (MFEL). This model provides an explanation as to how behavioral variability can be tuned while simultaneously improving performance during learning. We then develop a novel variant of inverse reinforcement learning to retrieve the cost function optimized by the brain during the learning process, as well as the parameter used to tune variability. We show that this population-level analysis can be used to derive a learning objective that each subject optimizes during his or her study. In this way, MFEL effectively acts as a unifying principle, allowing users to precisely formulate learning objectives and infer their structure. Brian A. Mitchell, Nina Lauharatanahirun, Javier O. Garcia, Nicholas F. Wymbs, Scott T. Grafton, Jean M. Vettel, Linda R. Petzold |
Neural Comput. | 7 |
| 2018 | Graph-based semi-supervised learning with genomic data integration using condition-responsive genes applied to phenotype classificationabstractObjective: Data integration methods that combine data from different molecular levels such as genome, epigenome, transcriptome, etc., have received a great deal of interest in the past few years. It has been demonstrated that the synergistic effects of different biological data types can boost learning capabilities and lead to a better understanding of the underlying interactions among molecular levels. Methods: In this paper we present a graph-based semi-supervised classification algorithm that incorporates latent biological knowledge in the form of biological pathways with gene expression and DNA methylation data. The process of graph construction from biological pathways is based on detecting condition-responsive genes, where 3 sets of genes are finally extracted: all condition responsive genes, high-frequency condition-responsive genes, and P-value-filtered genes. Results: The proposed approach is applied to ovarian cancer data downloaded from the Human Genome Atlas. Extensive numerical experiments demonstrate superior performance of the proposed approach compared to other state-of-the-art algorithms, including the latest graph-based classification techniques. Conclusions: Simulation results demonstrate that integrating various data types enhances classification performance and leads to a better understanding of interrelations between diverse omics data types. The proposed approach outperforms many of the state-of-the-art data integration algorithms. Abolfazl Doostparast Torshizi, Linda R. Petzold |
J. Am. Medical Informatics Assoc. | 2 |
| 2018 | Mechanical feedback coordinates cell wall expansion and assembly in yeast mating morphogenesisabstractThe shaping of individual cells requires a tight coordination of cell mechanics and growth. However, it is unclear how information about the mechanical state of the wall is relayed to the molecular processes building it, thereby enabling the coordination of cell wall expansion and assembly during morphogenesis. Combining theoretical and experimental approaches, we show that a mechanical feedback coordinating cell wall assembly and expansion is essential to sustain mating projection growth in budding yeast (Saccharomyces cerevisiae). Our theoretical results indicate that the mechanical feedback provided by the Cell Wall Integrity pathway, with cell wall stress sensors Wsc1 and Mid2 increasingly activating membrane-localized cell wall synthases Fks1/2 upon faster cell wall expansion, stabilizes mating projection growth without affecting cell shape. Experimental perturbation of the osmotic pressure and cell wall mechanics, as well as compromising the mechanical feedback through genetic deletion of the stress sensors, leads to cellular phenotypes that support the theoretical predictions. Our results indicate that while the existence of mechanical feedback is essential to stabilize mating projection growth, the shape and size of the cell are insensitive to the feedback. Samhita P. Banavar, Carlos Gomez, Michael Trogdon, Linda R. Petzold, Tau-Mu Yi, Otger Campàs |
PLoS Comput. Biol. | 4 |
| 2018 | The effect of cell geometry on polarization in budding yeastabstractThe localization (or polarization) of proteins on the membrane during the mating of budding yeast (Saccharomyces cerevisiae) is an important model system for understanding simple pattern formation within cells. While there are many existing mathematical models of polarization, for both budding and mating, there are still many aspects of this process that are not well understood. In this paper we set out to elucidate the effect that the geometry of the cell can have on the dynamics of certain models of polarization. Specifically, we look at several spatial stochastic models of Cdc42 polarization that have been adapted from published models, on a variety of tip-shaped geometries, to replicate the shape change that occurs during the growth of the mating projection. We show here that there is a complex interplay between the dynamics of polarization and the shape of the cell. Our results show that while models of polarization can generate a stable polarization cap, its localization at the tip of mating projections is unstable, with the polarization cap drifting away from the tip of the projection in a geometry dependent manner. We also compare predictions from our computational results to experiments that observe cells with projections of varying lengths, and track the stability of the polarization cap. Lastly, we examine one model of actin polarization and show that it is unlikely, at least for the models studied here, that actin dynamics and vesicle traffic are able to overcome this effect of geometry. Michael Trogdon, Brian Drawert, Carlos Gomez, Samhita P. Banavar, Tau-Mu Yi, Otger Campàs, Linda R. Petzold |
PLoS Comput. Biol. | 7 |
| 2018 | Sparse Pathway-Induced Dynamic Network Biomarker Discovery for Early Warning Signal Detection in Complex DiseasesabstractIn many complex diseases, the transition process from the healthy stage to the catastrophic stage does not occur gradually. Recent studies indicate that the initiation and progression of such diseases are comprised of three steps including healthy stage, pre-disease stage, and disease stage. It has been demonstrated that a certain set of trajectories can be observed in the genetic signatures at the molecular level, which might be used to detect the pre-disease stage and to take necessary medical interventions. In this paper, we propose two optimization-based algorithms for extracting the dynamic network biomarkers responsible for catastrophic transition into the disease stage, and to open new horizons to reverse the disease progression at an early stage through pinpointing molecular signatures provided by high-throughput microarray data. The first algorithm relies on meta-heuristic intelligent search to characterize dynamic network biomarkers represented as a complete graph. The second algorithm induces sparsity on the adjacency matrix of the genes by taking into account the biological signaling and metabolic pathways, since not all the genes in the ineractome are biologically linked. Comprehensive numerical and meta-analytical experiments verify the effectiveness of the results of the proposed approaches in terms of network size, biological meaningfulness, and verifiability. Abolfazl Doostparast Torshizi, Linda R. Petzold |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2017 | Survival Topic Models for Predicting Outcomes for Trauma PatientsabstractData mining techniques have been proposed to predict mortality for ICU patients using their demographic data, measurements and notes from doctors and nurses. Most of these techniques suffer from two main drawbacks. First, they model the mortality prediction problem as a binary classification problem, while ignoring the time of death as continuous values. Second, they use topic models to analyze the notes, while ignoring the relationship between measurements, notes and mortality/discharge outcomes. In this paper we propose a novel model called the survival topic model (SVTM), which models patients' measurements, notes and mortality/discharge jointly, and predicts the probability of mortality/discharge as functions of time. The idea is that each patient has a latent distribution of disease conditions, which we call topics. These conditions generate the measurements and notes and determine the patients' mortality. We derive a mean-field variational inference algorithm for this model. We fitted the SVTM with two outcomes on Medical Information Mart for Intensive Care III (MIMIC III) trauma patients data and obtained some important topics. Also, we demonstrated the relationships between these topics. Yuanyang Zhang, Richard M. Jiang 0002, Linda R. Petzold |
ICDE | 3 |
| 2017 | Multivariate soft repulsive system identification for constructing rule-based classification systems: Application to trauma clinical data
Abolfazl Doostparast Torshizi, Linda R. Petzold, Mitchell J. Cohen |
Neurocomputing | 2 |
| 2016 | Stochastic Simulation Service: Bridging the Gap between the Computational Expert and the BiologistabstractWe present StochSS: Stochastic Simulation as a Service, an integrated development environment for modeling and simulation of both deterministic and discrete stochastic biochemical systems in up to three dimensions. An easy to use graphical user interface enables researchers to quickly develop and simulate a biological model on a desktop or laptop, which can then be expanded to incorporate increasing levels of complexity. StochSS features state-of-the-art simulation engines. As the demand for computational power increases, StochSS can seamlessly scale computing resources in the cloud. In addition, StochSS can be deployed as a multi-user software environment where collaborators share computational resources and exchange models via a public model repository. We demonstrate the capabilities and ease of use of StochSS with an example of model development and simulation at increasing levels of complexity. Brian Drawert, Andreas Hellander, Benjamin B. Bales, Debjani Banerjee, Giovanni Bellesia, Bernie J. Daigle Jr., Geoffrey Douglas, Mengyuan Gu, Anand Gupta, Stefan Hellander, Christopher B. Horuk, Dibyendu Nath, Aviral Takkar, Sheng Wu 0002, Per Lötstedt, Chandra Krintz, Linda R. Petzold |
PLoS Comput. Biol. | 17 |
| 2016 | Macromolecular Crowding Regulates the Gene Expression Profile by Limiting DiffusionabstractWe seek to elucidate the role of macromolecular crowding in transcription and translation. It is well known that stochasticity in gene expression can lead to differential gene expression and heterogeneity in a cell population. Recent experimental observations by Tan et al. have improved our understanding of the functional role of macromolecular crowding. It can be inferred from their observations that macromolecular crowding can lead to robustness in gene expression, resulting in a more homogeneous cell population. We introduce a spatial stochastic model to provide insight into this process. Our results show that macromolecular crowding reduces noise (as measured by the kurtosis of the mRNA distribution) in a cell population by limiting the diffusion of transcription factors (i.e. removing the unstable intermediate states), and that crowding by large molecules reduces noise more efficiently than crowding by small molecules. Finally, our simulation results provide evidence that the local variation in chromatin density as well as the total volume exclusion of the chromatin in the nucleus can induce a homogenous cell population. Mahdi Golkaram, Stefan Hellander, Brian Drawert, Linda R. Petzold |
PLoS Comput. Biol. | 4 |
| 2015 | Direct higher order fuzzy rule-based classification system: Application in mortality predictionabstractTrauma is one of the leading causes of death in the U.S. and is ranked third among death causes across all age groups. This paper presents a novel fuzzy rule-based classification approach based on the concept of General Type-2 Fuzzy sets to predict mortality for trauma patients. In this approach each rule in the rule-base has an IF and a THEN part and parameters of the IF part (antecedents) are automatically extracted using powerful general type-2 fuzzy clustering algorithms which enables the model to deal with noisy and/or missing data. To verify efficacy of the proposed model, it has been implemented on several publicly available datasets. Finally, it is used to predict mortality among patients having traumatic injuries based on a large clinical dataset. Accuracy results demonstrate superior capabilities of the proposed approach compared to crisp and fuzzy classification methods in the literature. Abolfazl Doostparast Torshizi, Linda R. Petzold, Mitchell J. Cohen |
BIBM | 2 |
| 2015 | A Cure Time Model for Joint Prediction of Outcome and Time-to-OutcomeabstractThe Cox model has been widely used in time-to-outcome predictions, particularly in studies of medical patients, where prediction of the time of death is desired. In addition, the cure model has been proposed to model times of death for discharged patients. However, neither the Cox model nor the cure model allow explicit cure information and prediction of patient cure times (discharge times). In this paper we propose a new model, the "cure time model", which models the static data for dying patients, surviving patients, and their death/cure times jointly. It models (1) mortality via logistic regression and (2) death and discharge times via Cox models. We extend the cure time model to situations with censored data, where neither time of death nor discharge time are known, as well as to multiple (>2) outcomes. In addition, we propose a joint log-odds ratio which can predict the mortality of patients using the information from both the logistic regression and Cox models. We compare our model with the Cox and cure models on a trauma patient dataset from UCSF/San Francisco General Hospital. Our results show that the cure time model more accurately predicts both mortality and time-to-mortality for patients from these datasets. Yuanyang Zhang, Bernie J. Daigle Jr., Mitchell J. Cohen, Linda R. Petzold |
ICDM | 4 |
| 2015 | Inferring single-cell gene expression mechanisms using stochastic simulationabstractMOTIVATION: Stochastic promoter switching between transcriptionally active (ON) and inactive (OFF) states is a major source of noise in gene expression. It is often implicitly assumed that transitions between promoter states are memoryless, i.e. promoters spend an exponentially distributed time interval in each of the two states. However, increasing evidence suggests that promoter ON/OFF times can be non-exponential, hinting at more complex transcriptional regulatory architectures. Given the essential role of gene expression in all cellular functions, efficient computational techniques for characterizing promoter architectures are critically needed. RESULTS: We have developed a novel model reduction for promoters with arbitrary numbers of ON and OFF states, allowing us to approximate complex promoter switching behavior with Weibull-distributed ON/OFF times. Using this model reduction, we created bursty Monte Carlo expectation-maximization with modified cross-entropy method ('bursty MCEM(2)'), an efficient parameter estimation and model selection technique for inferring the number and configuration of promoter states from single-cell gene expression data. Application of bursty MCEM(2) to data from the endogenous mouse glutaminase promoter reveals nearly deterministic promoter OFF times, consistent with a multi-step activation mechanism consisting of 10 or more inactive states. Our novel approach to modeling promoter fluctuations together with bursty MCEM(2) provides powerful tools for characterizing transcriptional bursting across genes under different environmental conditions. AVAILABILITY AND IMPLEMENTATION: R source code implementing bursty MCEM(2) is available upon request. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bernie J. Daigle Jr., Mohammad Soltani, Linda R. Petzold, Abhyudai Singh |
Bioinform. | 3 |
| 2014 | BiP Clustering Facilitates Protein Folding in the Endoplasmic ReticulumabstractThe chaperone BiP participates in several regulatory processes within the endoplasmic reticulum (ER): translocation, protein folding, and ER-associated degradation. To facilitate protein folding, a cooperative mechanism known as entropic pulling has been proposed to demonstrate the molecular-level understanding of how multiple BiP molecules bind to nascent and unfolded proteins. Recently, experimental evidence revealed the spatial heterogeneity of BiP within the nuclear and peripheral ER of S. cerevisiae (commonly referred to as 'clusters'). Here, we developed a model to evaluate the potential advantages of accounting for multiple BiP molecules binding to peptides, while proposing that BiP's spatial heterogeneity may enhance protein folding and maturation. Scenarios were simulated to gauge the effectiveness of binding multiple chaperone molecules to peptides. Using two metrics: folding efficiency and chaperone cost, we determined that the single binding site model achieves a higher efficiency than models characterized by multiple binding sites, in the absence of cooperativity. Due to entropic pulling, however, multiple chaperones perform in concert to facilitate the resolubilization and ultimate yield of folded proteins. As a result of cooperativity, multiple binding site models used fewer BiP molecules and maintained a higher folding efficiency than the single binding site model. These insilico investigations reveal that clusters of BiP molecules bound to unfolded proteins may enhance folding efficiency through cooperative action via entropic pulling. Marc Griesemer, Carissa Young, Anne S. Robinson, Linda R. Petzold |
PLoS Comput. Biol. | 4 |
| 2013 | I act, therefore I judge: network sentiment dynamics based on user activity changeabstractThe study of influence, persuasion, and user sentiment dynamics within online communities has recently emerged as a highly active area of research. In this paper, we focus on analyzing and modeling user sentiment dynamics within a real-world social media such as Twitter. Beyond text and connectivity, we are interested in exploring the level of topical user posting activity and its effect on sentiment change. We perform topic-wise analysis of tweeting behavior that reveals a strong relationship between users' activity acceleration and topic sentiment change. Inspired by this empirical observation, we develop a new generative and predictive model that extends classical neighborhood-based influence propagation with the notion of user activation. We fit the parameters of our model to a large, real-world Twitter dataset and evaluate its utility to predict future sentiment change. Our model outperforms significantly (1 order of magnitude in accuracy) existing alternatives in identifying the individuals who are most likely to change sentiment based on past information. When predicting the next sentiment of users who actually change their opinion (a relatively rare event), our model is twice more accurate than alternatives, while its overall network accuracy is 94% on average. We also study the effect of inactive users on consensus efficiency in the opinion dynamics process both analytically and in simulation within the context of our model. Kathy Macropol, Petko Bogdanov, Ambuj K. Singh, Linda R. Petzold, Xifeng Yan |
ASONAM | 4 |
| 2013 | Spatial Stochastic Dynamics Enable Robust Cell PolarizationabstractAlthough cell polarity is an essential feature of living cells, it is far from being well-understood. Using a combination of computational modeling and biological experiments we closely examine an important prototype of cell polarity: the pheromone-induced formation of the yeast polarisome. Focusing on the role of noise and spatial heterogeneity, we develop and investigate two mechanistic spatial models of polarisome formation, one deterministic and the other stochastic, and compare the contrasting predictions of these two models against experimental phenotypes of wild-type and mutant cells. We find that the stochastic model can more robustly reproduce two fundamental characteristics observed in wild-type cells: a highly polarized phenotype via a mechanism that we refer to as spatial stochastic amplification, and the ability of the polarisome to track a moving pheromone input. Moreover, we find that only the stochastic model can simultaneously reproduce these characteristics of the wild-type phenotype and the multi-polarisome phenotype of a deletion mutant of the scaffolding protein Spa2. Significantly, our analysis also demonstrates that higher levels of stochastic noise results in increased robustness of polarization to parameter variation. Furthermore, our work suggests a novel role for a polarisome protein in the stabilization of actin cables. These findings elucidate the intricate role of spatial stochastic effects in cell polarity, giving support to a cellular model where noise and spatial heterogeneity combine to achieve robust biological function. Michael J. Lawson, Brian Drawert, Mustafa Khammash, Linda R. Petzold, Tau-Mu Yi |
PLoS Comput. Biol. | 4 |
| 2012 | Reducing Complexity in Management of eScience ComputationsabstractIn this paper we address reduction of complexity in management of scientific computations in distributed computing environments. We explore an approach based on separation of computation design (application development) and distributed execution of computations, and investigate best practices for construction of virtual infrastructures for computational science - software systems that abstract and virtualize the processes of managing scientific computations on heterogeneous distributed resource systems. As a result we present StratUm, a toolkit for management of eScience computations. To illustrate use of the toolkit, we present it in the context of a case study where we extend the capabilities of an existing kinetic Monte Carlo software framework to utilize distributed computational resources. The case study illustrates a viable design pattern for construction of virtual infrastructures for distributed scientific computing. The resulting infrastructure is evaluated using a computational experiment from molecular systems biology. Per-Olov Östberg, Andreas Hellander, Brian Drawert, Erik Elmroth, Sverker Holmgren, Linda R. Petzold |
CCGRID | 6 |
| 2012 | Accelerated maximum likelihood parameter estimation for stochastic biochemical systemsabstractBACKGROUND: A prerequisite for the mechanistic simulation of a biochemical system is detailed knowledge of its kinetic parameters. Despite recent experimental advances, the estimation of unknown parameter values from observed data is still a bottleneck for obtaining accurate simulation results. Many methods exist for parameter estimation in deterministic biochemical systems; methods for discrete stochastic systems are less well developed. Given the probabilistic nature of stochastic biochemical models, a natural approach is to choose parameter values that maximize the probability of the observed data with respect to the unknown parameters, a.k.a. the maximum likelihood parameter estimates (MLEs). MLE computation for all but the simplest models requires the simulation of many system trajectories that are consistent with experimental data. For models with unknown parameters, this presents a computational challenge, as the generation of consistent trajectories can be an extremely rare occurrence. RESULTS: We have developed Monte Carlo Expectation-Maximization with Modified Cross-Entropy Method (MCEM(2)): an accelerated method for calculating MLEs that combines advances in rare event simulation with a computationally efficient version of the Monte Carlo expectation-maximization (MCEM) algorithm. Our method requires no prior knowledge regarding parameter values, and it automatically provides a multivariate parameter uncertainty estimate. We applied the method to five stochastic systems of increasing complexity, progressing from an analytically tractable pure-birth model to a computationally demanding model of yeast-polarization. Our results demonstrate that MCEM(2) substantially accelerates MLE computation on all tested models when compared to a stand-alone version of MCEM. Additionally, we show how our method identifies parameter values for certain classes of models more accurately than two recently proposed computationally efficient methods. CONCLUSIONS: This work provides a novel, accelerated version of a likelihood-based parameter estimation method that can be readily applied to stochastic biochemical systems. In addition, our results suggest opportunities for added efficiency improvements that will further enhance our ability to mechanistically simulate biological processes. Bernie J. Daigle Jr., Min K. Roh, Linda R. Petzold, Jarad Niemi |
BMC Bioinform. | 3 |
| 2012 | Core module biomarker identification with network exploration for breast cancer metastasisabstractBACKGROUND: In a complex disease, the expression of many genes can be significantly altered, leading to the appearance of a differentially expressed "disease module". Some of these genes directly correspond to the disease phenotype, (i.e. "driver" genes), while others represent closely-related first-degree neighbours in gene interaction space. The remaining genes consist of further removed "passenger" genes, which are often not directly related to the original cause of the disease. For prognostic and diagnostic purposes, it is crucial to be able to separate the group of "driver" genes and their first-degree neighbours, (i.e. "core module") from the general "disease module". RESULTS: We have developed COMBINER: COre Module Biomarker Identification with Network ExploRation. COMBINER is a novel pathway-based approach for selecting highly reproducible discriminative biomarkers. We applied COMBINER to three benchmark breast cancer datasets for identifying prognostic biomarkers. COMBINER-derived biomarkers exhibited 10-fold higher reproducibility than other methods, with up to 30-fold greater enrichment for known cancer-related genes, and 4-fold enrichment for known breast cancer susceptible genes. More than 50% and 40% of the resulting biomarkers were cancer and breast cancer specific, respectively. The identified modules were overlaid onto a map of intracellular pathways that comprehensively highlighted the hallmarks of cancer. Furthermore, we constructed a global regulatory network intertwining several functional clusters and uncovered 13 confident "driver" genes of breast cancer metastasis. CONCLUSIONS: COMBINER can efficiently and robustly identify disease core module genes and construct their associated regulatory network. In the same way, it is potentially applicable in the characterization of any disease that can be probed with microarrays. Ruoting Yang, Bernie J. Daigle Jr., Linda R. Petzold, Francis J. Doyle III |
BMC Bioinform. | 3 |
| 2012 | Language and Runtime Support for Automatic Configuration and Deployment of Scientific Computing Software over Cloud Fabrics
Chris Bunch, Brian Drawert, Navraj Chohan, Chandra Krintz, Linda R. Petzold, Khawaja S. Shams |
J. Grid Comput. | 5 |
| 2011 | StochKit2: software for discrete stochastic simulation of biochemical systems with eventsabstractSUMMARY: StochKit2 is the first major upgrade of the popular StochKit stochastic simulation software package. StochKit2 provides highly efficient implementations of several variants of Gillespie's stochastic simulation algorithm (SSA), and tau-leaping with automatic step size selection. StochKit2 features include automatic selection of the optimal SSA method based on model properties, event handling, and automatic parallelism on multicore architectures. The underlying structure of the code has been completely updated to provide a flexible framework for extending its functionality. AVAILABILITY: StochKit2 runs on Linux/Unix, Mac OS X and Windows. It is freely available under GPL version 3 and can be downloaded from http://sourceforge.net/projects/stochkit/. CONTACT: [email protected]. Kevin R. Sanft, Sheng Wu 0002, Min K. Roh, Jin Fu 0001, Rone Kwei Lim, Linda R. Petzold |
Bioinform. | 6 |
| 2009 | Parallel simulation for a fish schooling model on a general-purpose graphics processing unitabstractAbstract We consider an individual‐based model for fish schooling, which incorporates a tendency for each fish to align its position and orientation with an appropriate average of its neighbors' positions and orientations, in addition to a tendency for each fish to avoid collisions. To accurately determine the statistical properties of the collective motion of fish whose dynamics are described by such a model, many realizations are typically required. This carries a very high computational cost. The current generation of graphics processing units is well suited to this task. We describe our implementation and present computational experiments illustrating the power of this technology for this important and challenging class of problems. Copyright © 2008 John Wiley & Sons, Ltd. Allison Kolpas, Linda R. Petzold, Jeff Moehlis |
Concurr. Comput. Pract. Exp. | 3 |
| 2004 | Parallel Simulation of Fluid Slip in a MicrochannelabstractSummary form only given. We investigate the parallel simulation of fluid slip along microchannel walls using the multicomponent lattice Boltzmann method (LBM) with domain decomposition. Because of the high complexity for microscale simulation, even a parallel computation of fluid slip can take days or weeks. Any slowness in the participating nodes in a cluster can drag the entire computation substantially, due to frequent node synchronization involved in each computational phase of the algorithm. We augment the parallel LBM algorithm with filtered dynamic remapping for lattice points. This filtered scheme uses lazy remapping and over-redistribution strategies to balance the computational speed of participating nodes and to minimize the performance impact of slow nodes on synchronized phases. Our experimental results indicate that the proposed technique can greatly speed up fluid slip simulation on a nondedicated cluster over a long period of execution time. Jingyu Zhou, Luoding Zhu, Linda R. Petzold, Tao Yang 0009 |
IPDPS | 3 |
| 2003 | Deriving User Interface Requirements from Densely Interleaved Scientific Computing ApplicationsabstractDeriving user interface requirements is a key step in user interface generation and maintenance. For single purpose numeric routines, user interface requirements are relatively simple to derive. However, general numeric packages, which are solvers for entire classes of problems, are densely interleaved with strands shared and mixed among user options. This complexity forms a significant barrier to the derivation of user interface requirements and therefore to user interface generation and maintainance. Our methodology uses a graph representation to find potential user decision points implied by the control structure of the code. This graph is then iteratively refined to form a decision point diagram, a state machine representation of all possible user traversals through a user interface for the underlying code. Andrew Strelzoff, Linda R. Petzold |
ASE | 2 |
| 1999 | Parallel sensitivity analysis for DAEs with many parametersabstractIn this paper, we discuss the parallel computation of the sensitivity analysis of systems of differential-algebraic equations (DAEs) with a moderate number of state variables and a large number of sensitivity parameters. Several parallel implementations based on DASSLSO are explored and their performance when using the Message Passing Interface (MPI) on an SGI Origin 2000 is compared. Copyright © 1999 John Wiley & Sons, Ltd. Linda R. Petzold |
Concurr. Pract. Exp. | 2 |
| 1995 | Parallel solution of large-scale differential-algebraic systemsabstractAbstract DASPK solves large‐scale systems of differential‐algebraic equations. It is based on the integration method in DASSL, but instead of a direct method for the associated linear systems which arise at each time step, the preconditioned GMRES iteration is applied in combination with an inexact Newton method. Two parallel versions of DASPK have been developed: DASPKF90, a Fortran 90 data parallel implementation, and DASPKMP, a message‐passing implementation written in Fortran 77 with extended BLAS. The parallel versions have been implemented for the Thinking Machines Corporation (TMC) CM‐5, a massively parallel multiprocessor, keeping the user interface relatively simple while allowing for portability to other massively parallel architectures. The codes have been demonstrated on several large‐scale test problems, including three‐dimensional formulations of the heat equation, the Cahn‐Hilliard equation and a multi‐species reaction‐diffusion problem. The formulations are described, including detail on preconditioning the Krylov iteration, timing results and performance analysis. Robert S. Maier 0002, W. Rath, Linda R. Petzold |
Concurr. Pract. Exp. | 3 |