VLDB 2026 Research / reviewers in the wild / expert
Lukas Käll
dblp:06/6643
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-5689-9797ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Semi-supervised Learning While Controlling the FDR with an Application to Tandem Mass Spectrometry Analysis
Jack Freestone, Lukas Käll, William Stafford Noble, Uri Keich |
RECOMB | 2 |
| 2024 | Pathway analysis through mutual informationabstractMOTIVATION: In pathway analysis, we aim to establish a connection between the activity of a particular biological pathway and a difference in phenotype. There are many available methods to perform pathway analysis, many of them rely on an upstream differential expression analysis, and many model the relations between the abundances of the analytes in a pathway as linear relationships. RESULTS: Here, we propose a new method for pathway analysis, MIPath, that relies on information theoretical principles and, therefore, does not model the association between pathway activity and phenotype, resulting in relatively few assumptions. For this, we construct a graph of the data points for each pathway using a nearest-neighbor approach and score the association between the structure of this graph and the phenotype of these same samples using Mutual Information while adjusting for the effects of random chance in each score. The initial nearest neighbor approach evades individual gene-level comparisons, hence making the method scalable and less vulnerable to missing values. These properties make our method particularly useful for single-cell data. We benchmarked our method on several single-cell datasets, comparing it to established and new methods, and found that it produces robust, reproducible, and meaningful scores. AVAILABILITY AND IMPLEMENTATION: Source code is available at https://github.com/statisticalbiotechnology/mipath, or through Python Package Index as "mipathway." Gustavo S. Jeuken, Lukas Käll |
Bioinform. | 2 |
| 2024 | ECCB2024: The 23rd European Conference on Computational BiologyabstractThis volume of Bioinformatics includes the proceedings papers of the 23rd European Conference on Computational Biology (ECCB2024) to be held in Turku, Finland, from 16 September to 20 September 2024, under the theme Data and Algorithms for Health and Science. More information on the ECCB2024 conference is available at the conference website https://eccb2024.fi/. ECCB is one of the main international conferences in the field of computational biology and bioinformatics together with the Intelligent Systems for Molecular Biology (ISMB) and the Research in Computational Molecular Biology (RECOMB). It is held jointly with the ISMB conference in odd-numbered years and independently in even-numbered years. ECCB attracts scientists and industry professionals from diverse disciplines, including mathematics, statistics, computer science, biology, and medicine. Rapid technological advancements enable life scientists to gain increasingly detailed insights in complex biological systems. These improvements in measurement technologies, however, introduce new challenges for data interpretation and necessitate the development of advanced computational techniques to manage the complexity and volume of the data. Consequently, the field of computational biology is rapidly evolving with new algorithms, software, and databases. The ECCB2024 conference showcases cutting-edge developments in systems biology, artificial intelligence, single-cell and spatial technologies, data integration, and more, addressing the growing demand for sophisticated algorithms to enhance the analysis of large-scale biological and biomedical datasets. This is highlighted by the most frequent keywords accompanying all the accepted submissions in ECCB2024 (tutorials/workshops, proceedings, highlight talks, posters), with the most popular keywords including ‘machine learning’, ‘deep learning’, ‘single-cell rnaseq’, and ‘multiomics’ (Fig. 1). Most frequent keywords among all accepted submissions in ECCB2024. The barplot shows the frequencies of top keywords appearing in at least ten submissions, while the word cloud illustrates the relative frequencies of all keywords appearing in at least five submissions. The ECCB2024 edition features five keynote lectures by distinguished speakers: Sarah Teichmann (Cambridge Stem Cell Institute, University of Cambridge, UK), Peer Bork (EMBL—European Molecular Biology Laboratory, Germany), Ileana Cristea (Princeton University, US), Jussi Taipale (University of Cambridge, UK), and Fabian Theis (Helmholtz Munich Computational Health Center, Germany). In addition, ECCB2024 hosts a scientific debate on data sharing and privacy protection by Melissa Haendel (University of North Carolina, US) and Yves Moreau (KU Leuven, Belgium). The ECCB2024 conference covers a wide range of topics, focusing on methodological advancements in computational biology as well as innovative application of computational techniques to life sciences and medicine. To provide a cohesive overview of recent scientific progress, the conference presentations are organized under six broad themes: (i) Genomes, (ii) Proteins, (iii) Systems biology and multiomics, (iv) Single-cell omics, (v) Microbiomes and planetary health, and (vi) Digital health. The proceedings talks present new scientific contributions, while the highlight talks showcase already published cutting-edge science in computational biology and related further developments. Poster presentations provide an opportunity for the participants to discuss their recent work with other researchers in the field. Workshops and tutorials on specialized topics prior to the main conference program are platforms to share practical experiences and learn new skills. In addition to scientific presentation tracks, ECCB2024 has a separate ELIXIR track, overseen by ELIXIR, focusing on advancements in infrastructure and services within ELIXIR nodes in support of the expert groups known as ‘ELIXIR Communities’. ELIXIR is a distributed pan-European life science infrastructure for biological data that coordinates, integrates, and sustains bioinformatics resources across its member states. This enables academic and industry users to access data, tools, standards, computing, and training services for life science research. ELIXIR has selected ECCB as a primary dissemination platform, serving as a co-organizing sponsor. Apart from Community activities the track will also introduce three scientific areas as outlined in the 2024–2028 ELIXIR Scientific Programme (https://elixir-europe.org), enabling scientists to access and analyse life science data across ‘Cellular and Molecular Research’, ‘Biodiversity, food security & pathogens’, and ‘Human data & translational research’. ECCB2024 received an impressive number of 200 submissions of full manuscripts for the proceedings call, highlighting the growing interest and engagement within the community. These submissions were organized under the six conference themes. Each submission was subjected to a peer-review process, with at least two reviews per manuscript, managed by the members of the ECCB2024 Programme Committee. The Programme Committee was chaired by Laura Elo (University of Turku, Finland) and each conference theme had an area chair (Table 1). The Proceedings Review Committee had over 100 reviewers. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. Thematic areas of ECCB2024 proceedings talks.a The table lists the area chairs for each theme, the number of reviewed papers, the number of accepted papers, and the acceptance rate for each theme. The review process focused on the impact, reproducibility, and scientific quality of the submitted research, as well as its relevance, interest, and value for the ECCB2024 audience. Upon completion of the review, the Program Committee chairs selected 24 papers to be included in the ECCB2024 proceedings, with an acceptance rate of 12% across the different areas (Table 1). All proceedings papers are published open access in this Proceedings issue of September 2024 of Bioinformatics. The ECCB2024 call for Highlight Talks invited presentations of studies recently published in scientific journals since 1 March 2023, or accepted for publication. The existence of further developments related to the published paper was also positively considered. The call received a total of 212 proposals, which were ranked according to their relevance and impact on computational biology, as well as their suitability to be presented to the large and diverse audience. Ultimately, the Program Committee selected 23% of the submissions for presentation as a Highlight talk at the conference. ECCB2024 hosts two poster sessions where researchers can introduce and discuss their work under the conference themes. In total, nearly 700 posters were accepted to be presented. The poster submissions were reviewed by the Posters Committee: Sini Junttila, Asta Laiho, Tomi Suomi, and Sampsa Hautaniemi (late posters). In addition, the ELIXIR track selected posters to be presented at ECCB2024. Overall, the ECCB2024 tracks attracted participation of presenting authors from 48 countries. The countries with the highest number of presenting authors were Germany, Finland, the USA, the United Kingdom, and Spain, jointly contributing to half of the overall number (Fig. 2). Following these were Turkey, France, and China, each contributing 4%–5% of the presenting authors. Proportion of countries where the presenting author of each accepted submission had their primary affiliation. Countries with the proportion below 2% were grouped under ‘Other’. Exhibitor booths are open throughout the conference, including ELIXIR, ISCB, Oxford University Press, Royal Society Publishing and eLife Sciences Publications Ltd, as well as the two organizing institutions: University of Turku and CSC—IT Center for Science. The booths showcase the latest scientific literature in computational biology and bioinformatics, data stewardship, as well as new developments in hardware, software, and technology. In addition, the Visit Turku Archipelago (tourist office) is also represented. Before the main ECCB2024 conference, a satellite meeting, seven workshops, and nine tutorials are organized. The satellite meeting is organized by the Student Council of the International Society for Computational Biology (ISCB) by and for early-stage researchers. This 8th European Student Council Symposium (ESCS) continues the successful collaboration between ESCS and ECCB, building the future of research in computational biology. The 16 workshops/tutorials preceding the ECCB2024 main conference were selected out of a total of 24 proposals by the workshops and tutorials chair Bengt Persson (Uppsala University, Sweden), each running for half a day. The workshops foster discussions and exchange of ideas on a range of specialized or emerging topics in computational biology and provide opportunities to share practical experiences. The tutorials provide participants with lectures and hands-on training to enable learning about new areas of computational biology or important established topics. Following the tradition of previous ECCB conferences, the ISCB and ECCB sponsored a number of travel fellowships for students and postdoctoral fellows. These fellowships were primarily awarded to members presenting talks and those from low- or middle-income countries, facilitating their participation in the conference. A total of 68 applications were received from scientists across 26 countries. After a careful review, the ECCB2024 Organizing Committee in collaboration with the ISCB and ECCB awarded 10 and 18 fellowships, respectively. The ECCB2024 Code of Conduct is designed to ensure a safe and respectful environment for all attendees, outlining clear standards of behaviour to foster an inclusive conference atmosphere. In addition, the code provides specific procedures for attendees to follow if they feel these standards have been violated. This ensures that all participants have access to support and resources to address any concerns, reinforcing the commitment of ECCB2024 to uphold the highest ethical and professional standards during the conference. The ECCB2024 Organizing Committee is strongly committed to promote gender equity throughout the conference planning and execution. In particular, an important effort was made to ensure that the review panels included a balanced composition of female and male professionals across the world. In addition, three out of the seven distinguished keynote speakers are women. The conference program also features a collaborative workshop with the Bioinfo4Women initiative, discussing sex and gender bias in artificial intelligence. We would like to thank everyone who has contributed to the success of ECCB2024, ensuring it meets high standards of excellence. Special recognition is due to the Program Committee, including theme area chairs and reviewers, whose critical and dedicated efforts have been fundamental to the conference organization in selecting excellent tutorials, workshops, manuscripts, talks, and posters. We are equally thankful to the ECCB steering committee for their invaluable support and advice, especially ECCB2022 organizers for providing detailed information about the organization of the previous ECCB conference. The collaboration with the ISCB has been crucial in offering travel fellowships and promoting ECCB2024 globally, for which we are deeply grateful. Our gratitude also goes to all our financial sponsors, including our co-organizing sponsor ELIXIR. We also acknowledge the Oxford University Press production team for their work on the ECCB2024 Proceedings issue. Many individuals have played significant roles in the local organization, often exceeding their responsibilities, and we are immensely thankful for their dedication. Lastly, the conference would not be what it is without the diverse participants from around the world, whose scientific contributions, presentations, and discussions enrich ECCB2024. Thank you all for being there and for allowing us to enjoy science at ECCB2024 in Turku! None declared. This paper was published as part of a supplement financially supported by ECCB2024. All data are incorporated into the article. Additional information is available at the conference website https://eccb2024.fi. Anu Kukkonen-Macchi, Sampsa Hautaniemi, Katharina F. Heil, Merja Heinäniemi, Lars Juhl Jensen, Sini Junttila, Lukas Käll, Asta Laiho, Peter Maccallum, Matti Nykter, Bengt Persson, Tomi Suomi, Tim Van Den Bossche, Tommi H. Nyrönen, Laura Elo |
Bioinform. | 7 |
| 2022 | Survival analysis of pathway activity as a prognostic determinant in breast cancerabstractHigh throughput biology enables the measurements of relative concentrations of thousands of biomolecules from e.g. tissue samples. The process leaves the investigator with the problem of how to best interpret the potentially large number of differences between samples. Many activities in a cell depend on ordered reactions involving multiple biomolecules, often referred to as pathways. It hence makes sense to study differences between samples in terms of altered pathway activity, using so-called pathway analysis. Traditional pathway analysis gives significance to differences in the pathway components' concentrations between sample groups, however, less frequently used methods for estimating individual samples' pathway activities have been suggested. Here we demonstrate that such a method can be used for pathway-based survival analysis. Specifically, we investigate the pathway activities' association with patients' survival time based on the transcription profiles of the METABRIC dataset. Our implementation shows that pathway activities are better prognostic markers for survival time in METABRIC than the individual transcripts. We also demonstrate that we can regress out the effect of individual pathways on other pathways, which allows us to estimate the other pathways' residual pathway activity on survival. Furthermore, we illustrate how one can visualize the often interdependent measures over hierarchical pathway databases using sunburst plots. Gustavo S. Jeuken, Nicholas P. Tobin, Lukas Käll |
PLoS Comput. Biol. | 3 |
| 2021 | Parallelized calculation of permutation testsabstractMOTIVATION: Permutation tests offer a straightforward framework to assess the significance of differences in sample statistics. A significant advantage of permutation tests are the relatively few assumptions about the distribution of the test statistic are needed, as they rely on the assumption of exchangeability of the group labels. They have great value, as they allow a sensitivity analysis to determine the extent to which the assumed broad sample distribution of the test statistic applies. However, in this situation, permutation tests are rarely applied because the running time of naïve implementations is too slow and grows exponentially with the sample size. Nevertheless, continued development in the 1980s introduced dynamic programming algorithms that compute exact permutation tests in polynomial time. Albeit this significant running time reduction, the exact test has not yet become one of the predominant statistical tests for medium sample size. Here, we propose a computational parallelization of one such dynamic programming-based permutation test, the Green algorithm, which makes the permutation test more attractive. RESULTS: Parallelization of the Green algorithm was found possible by non-trivial rearrangement of the structure of the algorithm. A speed-up-by orders of magnitude-is achievable by executing the parallelized algorithm on a GPU. We demonstrate that the execution time essentially becomes a non-issue for sample sizes, even as high as hundreds of samples. This improvement makes our method an attractive alternative to, e.g. the widely used asymptotic Mann-Whitney U-test. AVAILABILITYAND IMPLEMENTATION: In Python 3 code from the GitHub repository https://github.com/statisticalbiotechnology/parallelPermutationTest under an Apache 2.0 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Markus Ekvall, Michael Höhle, Lukas Käll |
Bioinform. | 3 |
| 2019 | CoExpresso: assess the quantitative behavior of protein complexes in human cellsabstractBACKGROUND: Translational and post-translational control mechanisms in the cell result in widely observable differences between measured gene transcription and protein abundances. Herein, protein complexes are among the most tightly controlled entities by selective degradation of their individual proteins. They furthermore act as control hubs that regulate highly important processes in the cell and exhibit a high functional diversity due to their ability to change their composition and their structure. Better understanding and prediction of these functional states demands methods for the characterization of complex composition, behavior, and abundance across multiple cell states. Mass spectrometry provides an unbiased approach to directly determine protein abundances across different cell populations and thus to profile a comprehensive abundance map of proteins. RESULTS: We provide a tool to investigate the behavior of protein subunits in known complexes by comparing their abundance profiles across up to 140 cell types available in ProteomicsDB. Thorough assessment of different randomization methods and statistical scoring algorithms allows determining the significance of concurrent profiles within a complex, therefore providing insights into the conservation of their composition across human cell types as well as the identification of intrinsic structures in complex behavior to determine which proteins orchestrate complex function. This analysis can be extended to investigate common profiles within arbitrary protein groups. CoExpresso can be accessed through http://computproteomics.bmb.sdu.dk/Apps/CoExpresso . CONCLUSIONS: With the CoExpresso web service, we offer a potent scoring scheme to assess proteins for their co-regulation and thereby offer insight into their potential for forming functional groups like protein complexes. Morteza Chalabi Hajkarim, Vasileios Tsiamis, Lukas Käll, Fabio Vandin, Veit Schwämmle |
BMC Bioinform. | 3 |
| 2017 | Uncertainty estimation of predictions of peptides' chromatographic retention times in shotgun proteomicsabstractMotivation: Liquid chromatography is frequently used as a means to reduce the complexity of peptide-mixtures in shotgun proteomics. For such systems, the time when a peptide is released from a chromatography column and registered in the mass spectrometer is referred to as the peptide's retention time . Using heuristics or machine learning techniques, previous studies have demonstrated that it is possible to predict the retention time of a peptide from its amino acid sequence. In this paper, we are applying Gaussian Process Regression to the feature representation of a previously described predictor E lude . Using this framework, we demonstrate that it is possible to estimate the uncertainty of the prediction made by the model. Here we show how this uncertainty relates to the actual error of the prediction. Results: In our experiments, we observe a strong correlation between the estimated uncertainty provided by Gaussian Process Regression and the actual prediction error. This relation provides us with new means for assessment of the predictions. We demonstrate how a subset of the peptides can be selected with lower prediction error compared to the whole set. We also demonstrate how such predicted standard deviations can be used for designing adaptive windowing strategies. Contact: [email protected]. Availability and Implementation: Our software and the data used in our experiments is publicly available and can be downloaded from https://github.com/statisticalbiotechnology/GPTime . Heydar Maboudi Afkham, Xuanbin Qiu, Matthew The, Lukas Käll |
Bioinform. | 4 |
| 2012 | A cross-validation scheme for machine learning algorithms in shotgun proteomicsabstractPeptides are routinely identified from mass spectrometry-based proteomics experiments by matching observed spectra to peptides derived from protein databases. The error rates of these identifications can be estimated by target-decoy analysis, which involves matching spectra to shuffled or reversed peptides. Besides estimating error rates, decoy searches can be used by semi-supervised machine learning algorithms to increase the number of confidently identified peptides. As for all machine learning algorithms, however, the results must be validated to avoid issues such as overfitting or biased learning, which would produce unreliable peptide identifications. Here, we discuss how the target-decoy method is employed in machine learning for shotgun proteomics, focusing on how the results can be validated by cross-validation, a frequently used validation scheme in machine learning. We also use simulated data to demonstrate the proposed cross-validation scheme's ability to detect overfitting. Viktor Granholm, William Stafford Noble, Lukas Käll |
BMC Bioinform. | 3 |
| 2011 | Computational Mass Spectrometry-Based ProteomicsabstractDOAJ is a unique and extensive index of diverse open access journals from around the world, driven by a growing community, committed to ensuring quality content is freely available online for everyone. Lukas Käll, Olga Vitek |
PLoS Comput. Biol. | 1 |
| 2009 | QVALITY: non-parametric estimation of q-values and posterior error probabilitiesabstractUNLABELLED: Qvality is a C++ program for estimating two types of standard statistical confidence measures: the q-value, which is an analog of the p-value that incorporates multiple testing correction, and the posterior error probability (PEP, also known as the local false discovery rate), which corresponds to the probability that a given observation is drawn from the null distribution. In computing q-values, qvality employs a standard bootstrap procedure to estimate the prior probability of a score being from the null distribution; for PEP estimation, qvality relies upon non-parametric logistic regression. Relative to other tools for estimating statistical confidence measures, qvality is unique in its ability to estimate both types of scores directly from a null distribution, without requiring the user to calculate p-values. AVAILABILITY: A web server, C++ source code and binaries are available under MIT license at http://noble.gs.washington.edu/proj/qvality. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lukas Käll, John D. Storey, William Stafford Noble |
Bioinform. | 1 |
| 2008 | Transmembrane Topology and Signal Peptide Prediction Using Dynamic Bayesian NetworksabstractHidden Markov models (HMMs) have been successfully applied to the tasks of transmembrane protein topology prediction and signal peptide prediction. In this paper we expand upon this work by making use of the more powerful class of dynamic Bayesian networks (DBNs). Our model, Philius, is inspired by a previously published HMM, Phobius, and combines a signal peptide submodel with a transmembrane submodel. We introduce a two-stage DBN decoder that combines the power of posterior decoding with the grammar constraints of Viterbi-style decoding. Philius also provides protein type, segment, and topology confidence metrics to aid in the interpretation of the predictions. We report a relative improvement of 13% over Phobius in full-topology prediction accuracy on transmembrane proteins, and a sensitivity and specificity of 0.96 in detecting signal peptides. We also show that our confidence metrics correlate well with the observed precision. In addition, we have made predictions on all 6.3 million proteins in the Yeast Resource Center (YRC) database. This large-scale study provides an overall picture of the relative numbers of proteins that include a signal-peptide and/or one or more transmembrane segments as well as a valuable resource for the scientific community. All DBNs are implemented using the Graphical Models Toolkit. Source code for the models described here is available at http://noble.gs.washington.edu/proj/philius. A Philius Web server is available at http://www.yeastrc.org/philius, and the predictions on the YRC database are available at http://www.yeastrc.org/pdr. Sheila M. Reynolds, Lukas Käll, Michael Riffle, Jeff A. Bilmes, William Stafford Noble |
PLoS Comput. Biol. | 2 |