Henrik Nielsen

dblp:15/1197 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-author
YearPublicationVenuePosition
2026 Whole-genome prediction of bacterial pathogenic capacity on novel bacteria using protein language models with PathogenFinder2
abstract
MOTIVATION: Infectious diseases continue to be a leading cause of mortality and pose a significant global health threat. Thus, the development of tools for surveillance and early detection of emerging pathogens is needed. RESULTS: We introduce PathogenFinder2, a novel, alignment-free, taxonomy-agnostic model for predicting bacterial pathogenic capacity in humans using protein language models. It outperforms previous methods, particularly for novel taxa, and provides interpretable outputs by highlighting proteins most relevant to pathogenic potential. These insights aid the identification of virulence factors, vaccine targets, and infection-related metabolic pathways. Furthermore, we introduce the Bacterial Pathogenic Capacity Landscape, which reveals patterns linked to host condition, infection site, microbial antagonism, and environmental origin. AVAILABILITY: The model is freely available online at https://genepi.dk/pathogenfinder2, or as a standalone program (https://github.com/genomicepidemiology/PathogenFinder2).
Alfred Ferrer Florensa, José Juan Almagro Armenteros, Rolf S. Kaas, Philip Thomas Lanken Conradsen Clausen, Henrik Nielsen, Burkhard Rost, Frank Møller Aarestrup
Bioinform.5
2025 NetStart 2.0: prediction of eukaryotic translation initiation sites using a protein language model
abstract
BACKGROUND: Accurate identification of translation initiation sites is essential for the proper translation of mRNA into functional proteins. In eukaryotes, the choice of the translation initiation site is influenced by multiple factors, including its proximity to the 5[Formula: see text] end and the local start codon context. Translation initiation sites mark the transition from non-coding to coding regions. This fact motivates the expectation that the upstream sequence, if translated, would assemble a nonsensical order of amino acids, while the downstream sequence would correspond to the structured beginning of a protein. This distinction suggests potential for predicting translation initiation sites using a protein language model. RESULTS: We present NetStart 2.0, a deep learning-based model that integrates the ESM-2 protein language model with the local sequence context to predict translation initiation sites across a broad range of eukaryotic species. NetStart 2.0 was trained as a single model across multiple species, and despite the broad phylogenetic diversity represented in the training data, it consistently relied on features marking the transition from non-coding to coding regions. CONCLUSION: By leveraging "protein-ness", NetStart 2.0 achieves state-of-the-art performance in predicting translation initiation sites across a diverse range of eukaryotic species. This success underscores the potential of protein language models to bridge transcript- and peptide-level information in complex biological prediction tasks. The NetStart 2.0 webserver is available at: https://services.healthtech.dtu.dk/services/NetStart-2.0/ .
Line Sandvad Nielsen, Anders Gorm Pedersen, Ole Winther, Henrik Nielsen
BMC Bioinform.4
2024 Predicting the subcellular location of prokaryotic proteins with DeepLocPro
abstract
MOTIVATION: Protein subcellular location prediction is a widely explored task in bioinformatics because of its importance in proteomics research. We propose DeepLocPro, an extension to the popular method DeepLoc, tailored specifically to archaeal and bacterial organisms. RESULTS: DeepLocPro is a multiclass subcellular location prediction tool for prokaryotic proteins, trained on experimentally verified data curated from UniProt and PSORTdb. DeepLocPro compares favorably to the PSORTb 3.0 ensemble method, surpassing its performance across multiple metrics in our benchmark experiment. AVAILABILITY AND IMPLEMENTATION: The DeepLocPro prediction tool is available online at https://ku.biolib.com/deeplocpro and https://services.healthtech.dtu.dk/services/DeepLocPro-1.0/.
Jaime Moreno 0005, Henrik Nielsen, Ole Winther, Felix Teufel
Bioinform.2
2022 NetSolP: predicting protein solubility in Escherichia coli using language models
abstract
MOTIVATION: Solubility and expression levels of proteins can be a limiting factor for large-scale studies and industrial production. By determining the solubility and expression directly from the protein sequence, the success rate of wet-lab experiments can be increased. RESULTS: In this study, we focus on predicting the solubility and usability for purification of proteins expressed in Escherichia coli directly from the sequence. Our model NetSolP is based on deep learning protein language models called transformers and we show that it achieves state-of-the-art performance and improves extrapolation across datasets. As we find current methods are built on biased datasets, we curate existing datasets by using strict sequence-identity partitioning and ensure that there is minimal bias in the sequences. AVAILABILITY AND IMPLEMENTATION: The predictor and data are available at https://services.healthtech.dtu.dk/service.php?NetSolP and the open-sourced code is available at https://github.com/tvinet/NetSolP-1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vineet Thumuluri, Hannah-Marie Martiny, José Juan Almagro Armenteros, Jesper Salomon, Henrik Nielsen, Alexander Rosenberg Johansen
Bioinform.5
2017 DeepLoc: prediction of protein subcellular localization using deep learning
abstract
MOTIVATION: The prediction of eukaryotic protein subcellular localization is a well-studied topic in bioinformatics due to its relevance in proteomics research. Many machine learning methods have been successfully applied in this task, but in most of them, predictions rely on annotation of homologues from knowledge databases. For novel proteins where no annotated homologues exist, and for predicting the effects of sequence variants, it is desirable to have methods for predicting protein properties from sequence information only. RESULTS: Here, we present a prediction algorithm using deep neural networks to predict protein subcellular localization relying only on sequence information. At its core, the prediction model uses a recurrent neural network that processes the entire protein sequence and an attention mechanism identifying protein regions important for the subcellular localization. The model was trained and tested on a protein dataset extracted from one of the latest UniProt releases, in which experimentally annotated proteins follow more stringent criteria than previously. We demonstrate that our model achieves a good accuracy (78% for 10 categories; 92% for membrane-bound or soluble), outperforming current state-of-the-art algorithms, including those relying on homology information. AVAILABILITY AND IMPLEMENTATION: The method is available as a web server at http://www.cbs.dtu.dk/services/DeepLoc. Example code is available at https://github.com/JJAlmagro/subcellular_localization. The dataset is available at http://www.cbs.dtu.dk/services/DeepLoc/data.php. CONTACT: [email protected].
José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, Ole Winther
Bioinform.4
2017 DeepLoc: prediction of protein subcellular localization using deep learning
abstract
Bioinformatics (2017) doi: 10.1093/bioinformatics/btx431 The authors of the above article wish to inform readers that the following change has been made post-publication: Sentence “If the Gorodkin measure is squared, the ‘generalized squared correlation’ (GC2) of Baldi et al. (2000) is obtained” has now been corrected to: “For K = 2, the Gorodkin measure squared is the ‘generalized squared correlation’ (GC2) of Baldi et al. (2000).”
José Juan Almagro Armenteros, Casper Kaae Sønderby, Søren Kaae Sønderby, Henrik Nielsen, Ole Winther
Bioinform.4
2017 An introduction to deep learning on biological sequence data: examples and solutions
abstract
MOTIVATION: Deep neural network architectures such as convolutional and long short-term memory networks have become increasingly popular as machine learning tools during the recent years. The availability of greater computational resources, more data, new algorithms for training deep models and easy to use libraries for implementation and training of neural networks are the drivers of this development. The use of deep learning has been especially successful in image recognition; and the development of tools, applications and code examples are in most cases centered within this field rather than within biology. RESULTS: Here, we aim to further the development of deep learning methods within biology by providing application examples and ready to apply and adapt code templates. Given such examples, we illustrate how architectures consisting of convolutional and long short-term memory neural networks can relatively easily be designed and trained to state-of-the-art performance on three biological sequence problems: prediction of subcellular localization, protein secondary structure and the binding of peptides to MHC Class II molecules. AVAILABILITY AND IMPLEMENTATION: All implementations and datasets are available online to the scientific community at https://github.com/vanessajurtz/lasagne4bio. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Vanessa Isabell Jurtz, Alexander Rosenberg Johansen, Morten Nielsen 0001, José Juan Almagro Armenteros, Henrik Nielsen, Casper Kaae Sønderby, Ole Winther, Søren Kaae Sønderby
Bioinform.5
2005 Prediction of twin-arginine signal peptides
abstract
BACKGROUND: Proteins carrying twin-arginine (Tat) signal peptides are exported into the periplasmic compartment or extracellular environment independently of the classical Sec-dependent translocation pathway. To complement other methods for classical signal peptide prediction we here present a publicly available method, TatP, for prediction of bacterial Tat signal peptides. RESULTS: We have retrieved sequence data for Tat substrates in order to train a computational method for discrimination of Sec and Tat signal peptides. The TatP method is able to positively classify 91% of 35 known Tat signal peptides and 84% of the annotated cleavage sites of these Tat signal peptides were correctly predicted. This method generates far less false positive predictions on various datasets than using simple pattern matching. Moreover, on the same datasets TatP generates less false positive predictions than a complementary rule based prediction method. CONCLUSION: The method developed here is able to discriminate Tat signal peptides from cytoplasmic proteins carrying a similar motif, as well as from Sec signal peptides, with high accuracy. The method allows filtering of input sequences based on Perl syntax regular expressions, whereas hydrophobicity discrimination of Tat- and Sec-signal peptides is carried out by an artificial neural network. A potential cleavage site of the predicted Tat signal peptide is also reported. The TatP prediction server is available as a public web server at http://www.cbs.dtu.dk/services/TatP/.
Jannick Dyrløv Bendtsen, Henrik Nielsen, David Widdick, Tracy Palmer, Søren Brunak
BMC Bioinform.2
2000 Assessing the accuracy of prediction algorithms for classification: an overview
abstract
Abstract 4 Also at the Department of Biological Sciences, University of California, Irvine, USA, to whom all correspondence should be addressed. We provide a unified overview of methods that currently are widely used to assess the accuracy of prediction algorithms, from raw percentages, quadratic error measures and other distances, and correlation coefficients, and to information theoretic measures such as relative entropy and mutual information. We briefly discuss the advantages and disadvantages of each approach. For classification tasks, we derive new learning algorithms for the design of prediction systems by directly optimising the correlation coefficient. We observe and prove several results relating sensitivity and specificity of optimal systems. While the principles are general, we illustrate the applicability on specific problems such as protein secondary structure and signal peptide prediction. Contact: [email protected]
Pierre Baldi, Søren Brunak, Yves Chauvin, Claus A. F. Andersen, Henrik Nielsen
Bioinform.5
1999 Metrics and Similarity Measures for Hidden Markov Models
Rune B. Lyngsø, Christian N. S. Pedersen, Henrik Nielsen
ISMB3
1998 Prediction of Signal Peptides and Signal Anchors by a Hidden Markov Model
Henrik Nielsen, Anders Krogh
ISMB1
1997 Neural Network Prediction of Translation Initiation Sites in Eukaryotes: Perspectives for EST and Genome Analysis
Anders Gorm Pedersen, Henrik Nielsen
ISMB2
1997 A Neural Network Method for Identification of Prokaryotic and Eukaryotic Signal Peptides and Prediction of their Cleavage Sites
abstract
We have developed a new method for the identification of signal peptides and their cleavage sites based on neural networks trained on separate sets of prokaryotic and eukaryotic sequences. The method performs significantly better than previous prediction schemes, and can easily be applied to genome-wide data sets. Discrimination between cleaved signal peptides and uncleaved N-terminal signal-anchor sequences is also possible, though with lower precision. Predictions can be made on a publicly available WWW server: http://www.cbs.dtu.dk/services/SignalP/.
Henrik Nielsen, Jacob Engelbrecht, Søren Brunak, Gunnar von Heijne
Int. J. Neural Syst.1
1995 Using Mirror Cameras for Estimating Depth
Jens Arnspang, Henrik Nielsen, Morten Christensen, Knud Henriksen
CAIP2
1995 Quantization of non-linear predictors in speech coding
abstract
Discusses how to exploit the nonlinearities in speech with the main purpose of improving the prediction in speech coders. If non-linearities are absent from speech the linear technique is sufficient, but if non-linearities are present the technique is inadequate and more sophisticated predictors are called for. Thyssen et al. (1994) gave evidence for non-linearities in speech and presented two non-linear short-term predictors that both were superior to the linear predictor without quantization. The present authors give methods to design vector quantizers for the non-linear predictors and investigate how vector quantization of the non-linear predictors affects prediction. Furthermore, they compare the performance of the quantized non-linear predictors to the performance of traditional quantized linear predictors. The experiments show that 10-bit VQ of the non-linear predictor leads to similar performance as 20-bit state-of-the-art split VQ of the LSP-parameters.
Jes Thyssen, Henrik Nielsen, Steffen Duus Hansen
ICASSP2
1994 Non-linear short-term prediction in speech coding
abstract
Addresses the question of how to extract the nonlinearities in speech with the prime purpose of facilitating coding of the residual signal in residual excited coders. The short-term prediction of speech in speech coders is extensively based on linear models, e.g. the linear predictive coding technique (LPC), which is one of the most basic elements in modern speech coders. This technique does not allow extraction of nonlinear dependencies. If nonlinearities are absent from speech the technique is sufficient, but if the speech contains nonlinearities the technique is inadequate. The authors give evidence for nonlinearities in speech and propose nonlinear short-term predictors that can substitute the LPC technique. The technique, called nonlinear predictive coding, is shown to be superior to the LPC technique. Two different nonlinear predictors are presented. The first is based on a second-order Volterra filter, and the second is based on a time delay neural network. The latter is shown to be the more suitable for speech coding applications.>
Jes Thyssen, Henrik Nielsen, Steffen Duus Hansen
ICASSP (1)2
1993 Improving the spectral balance of digital speech synthesis applied to a female, synthetic voice
Ida Frehr, Marianne Elmlund, Henrik Nielsen
EUROSPEECH3
1992 Improvements in 2.4 kbps high-quality speech coding
abstract
An algorithm for 2.4 kb/s speech coding is described. The main problem addressed is the coding of voiced speech. A way of coding the pitch structure is introduced. Compared with traditional coding schemes, it results in a better compromise between bit allocation for short-term quantization and residual coding. The coder uses vector quantization of the short-term parameters (line spectrum frequencies). The residual is lowpass filtered to obtain the baseband signal. Unvoiced frames are coded by means of a method based on repetition and interpolation of pitch pulses. The method exploits the high correlation between pitch pulses. Harmonic postfiltering is applied to obtain an improved high-frequency regeneration.>
Jesper Haagen, Henrik Nielsen, Steffen Duus Hansen
ICASSP2
1992 Comparative study of error correction coding schemes for the GSM half-rate channel
abstract
The problem of achieving robust speech transmission on a digital mobile radio channel subjected to severe fading is addressed. The study includes a performance comparison between the Reed-Solomon and the punctured convolutional types of channel coders, which were used in conjunction with speech coders with bit rates ranging from 6.6 kb/s down to 5.4 kb/s. Speech coding within this range is of particular interest for the standardization of the GSM half-rate system, where the gross bit rate is limited to 11.4 kb/s. The properties of the speech coders and the channel coders are described, as well as results of a series of experiments comparing speech quality and robustness against the transmission errors of a typical GSM half-rate channel.>
Henrik Nielsen, Keld B. Mikkelsen, Henrik B. Hansen, Yuhang Wu 0010, Knud J. Larsen, John Aasted Sørensen
ICASSP1
1991 A 2.4 kbps high-quality speech coder
abstract
An algorithm for coding speech at 2.4 kbps is presented. By combining in a new way well-known coding strategies from both medium and low bit rate speech coding, a very promising configuration is developed. The coder is fundamentally a baseband coder where short-term correlation is predicted by LPC (linear predictive coding) analysis. Coding of the short-term parameters is performed by vector quantization. An open-loop long-term predictor is applied to the lowpass filtered short-term residual to reduce the quasi-periodic pitch structure before the signal is down-sampled. A novel method for coding the down-sampled residual based on voiced/unvoiced classification is proposed. For unvoiced frames a simple white Gaussian codebook is applied, and for voiced frames a new pulse codebook has been designed. The coding algorithm is demonstrated to be robust and to produce high speech quality at 2.4 kbps.>
Jesper Haagen, Henrik Nielsen, Steffen Duus Hansen
ICASSP2
1991 High performance coder: a possible candidate for the GSM half-rate system
abstract
The authors describe their contribution to the ETSI (European Telecommunications Standards Institute) effort for a definition of a GSM (Group Special Mobile) standard half-rate coding system for the pan-European digital mobile cellular radio system. The system consists of a speech coder and a channel coder. An overview of the speech coding algorithm is given. The results from a study of the different versions of the speech coder are described. A two-stage channel coding scheme is presented, and the performance test results are shown. An overview of the DSP implementation, computational requirements, and memory usage is also presented.>
Yuhang Wu 0010, Henrik B. Hansen, Knud J. Larsen, Henrik Nielsen, John Aasted Sørensen
ICASSP4