Simon Anders

dblp:39/7064 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
1since 2021 · last 2022
0000-0003-4868-1805ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
5 papers
Bioinformatics and computational biology · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › sequence analysis
high-throughput sequencing data analysis
0.932022
Analysing high-throughput sequencing data in Python with HTSeq 2.0 · Bioinform. 2022
HTSeq - a Python framework to work with high-throughput sequencing data · Bioinform. 2015
ShortRead: a bioconductor package for input, quality assessment and exploration of high-throughput sequence data · Bioinform. 2009
Bioinformatics and computational biology › genomics
genomic data analysis
0.612022
Analysing high-throughput sequencing data in Python with HTSeq 2.0 · Bioinform. 2022
Bioinformatics and computational biology › epigenomics
chromatin conformation analysis
0.212015
FourCSeq: analysis of 4C sequencing data · Bioinform. 2015
Bioinformatics and computational biology
read counting
0.212015
HTSeq - a Python framework to work with high-throughput sequencing data · Bioinform. 2015
Bioinformatics and computational biology › transcriptomics
RNA-seq analysis
0.212015
HTSeq - a Python framework to work with high-throughput sequencing data · Bioinform. 2015
Bioinformatics and computational biology
genomics
0.112009
Visualization of genomic data with the Hilbert curve · Bioinform. 2009
Visualization and visual analytics
biological data visualization
0.112009
Visualization of genomic data with the Hilbert curve · Bioinform. 2009
Visualization and visual analytics › biological data visualization
genomic data visualization
0.112009
Visualization of genomic data with the Hilbert curve · Bioinform. 2009
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics
0.112015
FourCSeq: analysis of 4C sequencing data · Bioinform. 2015

Methods — techniques the papers use, named apart from their topics

Python API · 0.6z-score · 0.2python library · 0.2monotone curve fitting · 0.2DESeq2 · 0.2hilbert curve · 0.2
YearPublicationVenuePosition
2022 Analysing high-throughput sequencing data in Python with HTSeq 2.0
abstract
SUMMARY: HTSeq 2.0 provides a more extensive application programming interface including a new representation for sparse genomic data, enhancements for htseq-count to suit single-cell omics, a new script for data using cell and molecular barcodes, improved documentation, testing and deployment, bug fixes and Python 3 support. AVAILABILITY AND IMPLEMENTATION: HTSeq 2.0 is released as an open-source software under the GNU General Public License and is available from the Python Package Index at https://pypi.python.org/pypi/HTSeq. The source code is available on Github at https://github.com/htseq/htseq. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Givanna H. Putri, Simon Anders, Paul Theodor Pyl, John E. Pimanda, Fabio Zanini
Bioinform.2
2019 Focused multidimensional scaling: interactive visualization for exploration of high-dimensional data
abstract
BACKGROUND: Visualization is an important tool for generating meaning from scientific data, but the visualization of structures in high-dimensional data (such as from high-throughput assays) presents unique challenges. Dimension reduction methods are key in solving this challenge, but these methods can be misleading- especially when apparent clustering in the dimension-reducing representation is used as the basis for reasoning about relationships within the data. RESULTS: We present two interactive visualization tools, distnet and focusedMDS, that help in assessing the validity of a dimension-reducing plot and in interactively exploring relationships between objects in the data. The distnet tool is used to examine discrepancies between the placement of points in a two dimensional visualization and the points' actual similarities in feature space. The focusedMDS tool is an intuitive, interactive multidimensional scaling tool that is useful for exploring the relationships of one particular data point to the others, that might be useful in a personalized medicine framework. CONCLUSIONS: We introduce here two freely available tools for visually exploring and verifying the validity of dimension-reducing visualizations and biological information gained from these. The use of such tools can confirm that conclusions drawn from dimension-reducing visualizations are not simply artifacts of the visualization method, but are real biological insights.
Lea M. Urpa, Simon Anders
BMC Bioinform.2
2015 HTSeq - a Python framework to work with high-throughput sequencing data
abstract
MOTIVATION: A large choice of tools exists for many standard tasks in the analysis of high-throughput sequencing (HTS) data. However, once a project deviates from standard workflows, custom scripts are needed. RESULTS: We present HTSeq, a Python library to facilitate the rapid development of such scripts. HTSeq offers parsers for many common data formats in HTS projects, as well as classes to represent data, such as genomic coordinates, sequences, sequencing reads, alignments, gene model information and variant calls, and provides data structures that allow for querying via genomic coordinates. We also present htseq-count, a tool developed with HTSeq that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes. AVAILABILITY AND IMPLEMENTATION: HTSeq is released as an open-source software under the GNU General Public Licence and available from http://www-huber.embl.de/HTSeq or from the Python Package Index at https://pypi.python.org/pypi/HTSeq.
Simon Anders, Paul Theodor Pyl, Wolfgang Huber
Bioinform.1
2015 FourCSeq: analysis of 4C sequencing data
abstract
MOTIVATION: Circularized Chromosome Conformation Capture (4C) is a powerful technique for studying the spatial interactions of a specific genomic region called the 'viewpoint' with the rest of the genome, both in a single condition or comparing different experimental conditions or cell types. Observed ligation frequencies typically show a strong, regular dependence on genomic distance from the viewpoint, on top of which specific interaction peaks are superimposed. Here, we address the computational task to find these specific peaks and to detect changes between different biological conditions. RESULTS: We model the overall trend of decreasing interaction frequency with genomic distance by fitting a smooth monotonically decreasing function to suitably transformed count data. Based on the fit, z-scores are calculated from the residuals, and high z-scores are interpreted as peaks providing evidence for specific interactions. To compare different conditions, we normalize fragment counts between samples, and call for differential contact frequencies using the statistical method DESEQ2: adapted from RNA-Seq analysis. AVAILABILITY AND IMPLEMENTATION: A full end-to-end analysis pipeline is implemented in the R package FourCSeq available at www.bioconductor.org. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Felix A. Klein, Tibor Pakozdi, Simon Anders, Yad Ghavi-Helm, Eileen E. M. Furlong, Wolfgang Huber
Bioinform.3
2009 Visualization of genomic data with the Hilbert curve
abstract
UNLABELLED: In many genomic studies, one works with genome-position-dependent data, e.g. ChIP-chip or ChIP-Seq scores. Using conventional tools, it can be difficult to get a good feel for the data, especially the distribution of features. This article argues that the so-called Hilbert curve visualization can complement genome browsers and help to get further insights into the structure of one's data. This is demonstrated with examples from different use cases. An open-source application, called HilbertVis, is presented that allows the user to produce and interactively explore such plots. AVAILABILITY: http://www.ebi.ac.uk/huber-srv/hilbert/.
Simon Anders
Bioinform.1
2009 ShortRead: a bioconductor package for input, quality assessment and exploration of high-throughput sequence data
abstract
UNLABELLED: ShortRead is a package for input, quality assessment, manipulation and output of high-throughput sequencing data. ShortRead is provided in the R and Bioconductor environments, allowing ready access to additional facilities for advanced statistical analysis, data transformation, visualization and integration with diverse genomic resources. AVAILABILITY AND IMPLEMENTATION: This package is implemented in R and available at the Bioconductor web site; the package contains a 'vignette' outlining typical work flows.
Martin Morgan, Simon Anders, Michael F. Lawrence, Patrick Aboyoun, Hervé Pagès, Robert Gentleman
Bioinform.2