EDBT 2026 Demo / reviewers in the wild / expert
Simone Zaccaria
dblp:140/3270
· DBLP profile ↗
11ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-5265-7392ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
cancer genomics |
1.6 | 3 | 2025 | Characterizing the Solution Space of Migration Histories of Metastatic Cancers with MACH2 · RECOMB 2025 Single-Cell Tumor Phylogeny Inference with Copy-Number Constrained Mutation Losses · RECOMB 2020 The Copy-Number Tree Mixture Deconvolution Problem and Applications to Multi-sample Bulk Sequencing Tumor Data · RECOMB 2017 |
Bioinformatics and computational biology › cancer genomics › tumor evolution
tumor phylogenetics |
0.9 | 1 | 2025 | Characterizing the Solution Space of Migration Histories of Metastatic Cancers with MACH2 · RECOMB 2025 |
Bioinformatics and computational biology › cancer genomics › copy number analysis
copy number aberration analysis |
0.4 | 1 | 2020 | Single-Cell Tumor Phylogeny Inference with Copy-Number Constrained Mutation Losses · RECOMB 2020 |
Bioinformatics and computational biology › single-cell analysis
single-cell sequencing |
0.4 | 1 | 2020 | Identifying tumor clones in sparse single-cell mutation data · Bioinform. 2020 |
Bioinformatics and computational biology › cancer genomics
somatic mutation analysis |
0.4 | 1 | 2020 | Identifying tumor clones in sparse single-cell mutation data · Bioinform. 2020 |
Bioinformatics and computational biology › cancer genomics › tumor evolution
tumor phylogeny inference |
0.4 | 1 | 2020 | Single-Cell Tumor Phylogeny Inference with Copy-Number Constrained Mutation Losses · RECOMB 2020 |
Bioinformatics and computational biology › genomics
haplotype inference |
0.2 | 1 | 2016 | HapCol: accurate and memory-efficient haplotype assembly from long reads · Bioinform. 2016 |
Bioinformatics and computational biology
population genetics |
0.2 | 1 | 2016 | HapCol: accurate and memory-efficient haplotype assembly from long reads · Bioinform. 2016 |
Methods — techniques the papers use, named apart from their topics
markov chain monte carlo · 0.9stochastic block model · 0.4phylogenetic inference · 0.4copy-number constraint modeling · 0.4exact algorithm · 0.2error-correction minimization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Characterizing the Solution Space of Migration Histories of Metastatic Cancers with MACH2
Mrinmoy Saha Roddur, Vikram Ramavarapu, Abigail Bunkum, Ariana Huebner, Roman Mineyev, Nicholas McGranahan, Simone Zaccaria, Mohammed El-Kebir |
RECOMB | 7 |
| 2022 | CNAViz: An interactive webtool for user-guided segmentation of tumor DNA sequencing dataabstractCopy-number aberrations (CNAs) are genetic alterations that amplify or delete the number of copies of large genomic segments. Although they are ubiquitous in cancer and, thus, a critical area of current cancer research, CNA identification from DNA sequencing data is challenging because it requires partitioning of the genome into complex segments with the same copy-number states that may not be contiguous. Existing segmentation algorithms address these challenges either by leveraging the local information among neighboring genomic regions, or by globally grouping genomic regions that are affected by similar CNAs across the entire genome. However, both approaches have limitations: overclustering in the case of local segmentation, or the omission of clusters corresponding to focal CNAs in the case of global segmentation. Importantly, inaccurate segmentation will lead to inaccurate identification of CNAs. For this reason, most pan-cancer research studies rely on manual procedures of quality control and anomaly correction. To improve copy-number segmentation, we introduce CNAViz, a web-based tool that enables the user to simultaneously perform local and global segmentation, thus overcoming the limitations of each approach. Using simulated data, we demonstrate that by several metrics, CNAViz allows the user to obtain more accurate segmentation relative to existing local and global segmentation methods. Moreover, we analyze six bulk DNA sequencing samples from three breast cancer patients. By validating with parallel single-cell DNA sequencing data from the same samples, we show that by using CNAViz, our user was able to obtain more accurate segmentation and improved accuracy in downstream copy-number calling. Zubair Lalani, Gillian Chu, Silas Hsu, Shaw Kagawa, Michael Xiang, Simone Zaccaria, Mohammed El-Kebir |
PLoS Comput. Biol. | 6 |
| 2021 | Parsimonious Clone Tree Reconciliation in CancerabstractEvery tumor is composed of heterogeneous clones, each corresponding to a distinct subpopulation of cells that accumulated different types of somatic mutations, ranging from single-nucleotide variants (SNVs) to copy-number aberrations (CNAs). As the analysis of this intra-tumor heterogeneity has important clinical applications, several computational methods have been introduced to identify clones from DNA sequencing data. However, due to technological and methodological limitations, current analyses are restricted to identifying tumor clones only based on either SNVs or CNAs, preventing a comprehensive characterization of a tumor’s clonal composition. To overcome these challenges, we formulate the identification of clones in terms of both SNVs and CNAs as a reconciliation problem while accounting for uncertainty in the input SNV and CNA proportions. We thus characterize the computational complexity of this problem and we introduce a mixed integer linear programming formulation to solve it exactly. On simulated data, we show that tumor clones can be identified reliably, especially when further taking into account the ancestral relationships that can be inferred from the input SNVs and CNAs. On 49 tumor samples from 10 prostate cancer patients, our reconciliation approach provides a higher resolution view of tumor evolution than previous studies. Palash Sashittal, Simone Zaccaria, Mohammed El-Kebir |
WABI | 2 |
| 2020 | Single-Cell Tumor Phylogeny Inference with Copy-Number Constrained Mutation Losses
Gryte Satas, Simone Zaccaria, Geoffrey Mon, Benjamin J. Raphael |
RECOMB | 2 |
| 2020 | Identifying tumor clones in sparse single-cell mutation dataabstractMOTIVATION: Recent single-cell DNA sequencing technologies enable whole-genome sequencing of hundreds to thousands of individual cells. However, these technologies have ultra-low sequencing coverage (<0.5× per cell) which has limited their use to the analysis of large copy-number aberrations (CNAs) in individual cells. While CNAs are useful markers in cancer studies, single-nucleotide mutations are equally important, both in cancer studies and in other applications. However, ultra-low coverage sequencing yields single-nucleotide mutation data that are too sparse for current single-cell analysis methods. RESULTS: We introduce SBMClone, a method to infer clusters of cells, or clones, that share groups of somatic single-nucleotide mutations. SBMClone uses a stochastic block model to overcome sparsity in ultra-low coverage single-cell sequencing data, and we show that SBMClone accurately infers the true clonal composition on simulated datasets with coverage at low as 0.2×. We applied SBMClone to single-cell whole-genome sequencing data from two breast cancer patients obtained using two different sequencing technologies. On the first patient, sequenced using the 10X Genomics CNV solution with sequencing coverage ≈0.03×, SBMClone recovers the major clonal composition when incorporating a small amount of additional information. On the second patient, where pre- and post-treatment tumor samples were sequenced using DOP-PCR with sequencing coverage ≈0.5×, SBMClone shows that tumor cells are present in the post-treatment sample, contrary to published analysis of this dataset. AVAILABILITY AND IMPLEMENTATION: SBMClone is available on the GitHub repository https://github.com/raphael-group/SBMClone. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Matthew A. Myers, Simone Zaccaria, Benjamin J. Raphael |
Bioinform. | 2 |
| 2018 | HapCHAT: adaptive haplotype assembly for efficiently leveraging high coverage in long readsabstractBACKGROUND: Haplotype assembly is the process of assigning the different alleles of the variants covered by mapped sequencing reads to the two haplotypes of the genome of a human individual. Long reads, which are nowadays cheaper to produce and more widely available than ever before, have been used to reduce the fragmentation of the assembled haplotypes since their ability to span several variants along the genome. These long reads are also characterized by a high error rate, an issue which may be mitigated, however, with larger sets of reads, when this error rate is uniform across genome positions. Unfortunately, current state-of-the-art dynamic programming approaches designed for long reads deal only with limited coverages. RESULTS: Here, we propose a new method for assembling haplotypes which combines and extends the features of previous approaches to deal with long reads and higher coverages. In particular, our algorithm is able to dynamically adapt the estimated number of errors at each variant site, while minimizing the total number of error corrections necessary for finding a feasible solution. This allows our method to significantly reduce the required computational resources, allowing to consider datasets composed of higher coverages. The algorithm has been implemented in a freely available tool, HapCHAT: Haplotype Assembly Coverage Handling by Adapting Thresholds. An experimental analysis on sequencing reads with up to 60 × coverage reveals improvements in accuracy and recall achieved by considering a higher coverage with lower runtimes. CONCLUSIONS: Our method leverages the long-range information of sequencing reads that allows to obtain assembled haplotypes fragmented in a lower number of unphased haplotype blocks. At the same time, our method is also able to deal with higher coverages to better correct the errors in the original reads and to obtain more accurate haplotypes as a result. AVAILABILITY: HapCHAT is available at http://hapchat.algolab.eu under the GNU Public License (GPL). Stefano Beretta 0001, Murray Patterson, Simone Zaccaria, Gianluca Della Vedova, Paola Bonizzoni |
BMC Bioinform. | 3 |
| 2017 | The Copy-Number Tree Mixture Deconvolution Problem and Applications to Multi-sample Bulk Sequencing Tumor Data
Simone Zaccaria, Mohammed El-Kebir, Gunnar W. Klau, Benjamin J. Raphael |
RECOMB | 1 |
| 2016 | Copy-Number Evolution Problems: Complexity and Algorithms
Mohammed El-Kebir, Benjamin J. Raphael, Ron Shamir, Roded Sharan, Simone Zaccaria, Meirav Zehavi, Ron Zeira |
WABI | 5 |
| 2016 | HapCol: accurate and memory-efficient haplotype assembly from long readsabstractMOTIVATION: Haplotype assembly is the computational problem of reconstructing haplotypes in diploid organisms and is of fundamental importance for characterizing the effects of single-nucleotide polymorphisms on the expression of phenotypic traits. Haplotype assembly highly benefits from the advent of 'future-generation' sequencing technologies and their capability to produce long reads at increasing coverage. Existing methods are not able to deal with such data in a fully satisfactory way, either because accuracy or performances degrade as read length and sequencing coverage increase or because they are based on restrictive assumptions. RESULTS: By exploiting a feature of future-generation technologies-the uniform distribution of sequencing errors-we designed an exact algorithm, called HapCol, that is exponential in the maximum number of corrections for each single-nucleotide polymorphism position and that minimizes the overall error-correction score. We performed an experimental analysis, comparing HapCol with the current state-of-the-art combinatorial methods both on real and simulated data. On a standard benchmark of real data, we show that HapCol is competitive with state-of-the-art methods, improving the accuracy and the number of phased positions. Furthermore, experiments on realistically simulated datasets revealed that HapCol requires significantly less computing resources, especially memory. Thanks to its computational efficiency, HapCol can overcome the limits of previous approaches, allowing to phase datasets with higher coverage and without the traditional all-heterozygous assumption. AVAILABILITY AND IMPLEMENTATION: Our source code is available under the terms of the GNU General Public License at http://hapcol.algolab.eu/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuri Pirola, Simone Zaccaria, Riccardo Dondi, Gunnar W. Klau, Nadia Pisanti, Paola Bonizzoni |
Bioinform. | 2 |
| 2015 | On the Fixed Parameter Tractability and Approximability of the Minimum Error Correction Problem
Paola Bonizzoni, Riccardo Dondi, Gunnar W. Klau, Yuri Pirola, Nadia Pisanti, Simone Zaccaria |
CPM | 6 |
| 2013 | On the inversion-indel distanceabstractBACKGROUND: The inversion distance, that is the distance between two unichromosomal genomes with the same content allowing only inversions of DNA segments, can be computed thanks to a pioneering approach of Hannenhalli and Pevzner in 1995. In 2000, El-Mabrouk extended the inversion model to allow the comparison of unichromosomal genomes with unequal contents, thus insertions and deletions of DNA segments besides inversions. However, an exact algorithm was presented only for the case in which we have insertions alone and no deletion (or vice versa), while a heuristic was provided for the symmetric case, that allows both insertions and deletions and is called the inversion-indel distance. In 2005, Yancopoulos, Attie and Friedberg started a new branch of research by introducing the generic double cut and join (DCJ) operation, that can represent several genome rearrangements (including inversions). Among others, the DCJ model gave rise to two important results. First, it has been shown that the inversion distance can be computed in a simpler way with the help of the DCJ operation. Second, the DCJ operation originated the DCJ-indel distance, that allows the comparison of genomes with unequal contents, considering DCJ, insertions and deletions, and can be computed in linear time. RESULTS: In the present work we put these two results together to solve an open problem, showing that, when the graph that represents the relation between the two compared genomes has no bad components, the inversion-indel distance is equal to the DCJ-indel distance. We also give a lower and an upper bound for the inversion-indel distance in the presence of bad components. Eyla Willing, Simone Zaccaria, Marília D. V. Braga, Jens Stoye |
BMC Bioinform. | 2 |