EDBT 2026 Demo / reviewers in the wild / expert
Mark Howison
dblp:63/9083
· DBLP profile ↗
12ranked-venue papers
6as first author
1since 2021 · last 2025
0000-0002-0764-4090ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 75% Parallel and multicore computing · 16% Performance modeling and evaluation · 5% | |
| Computer graphics and multimedia
1 paper |
Rendering · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Learning and educational technologies · 77% Wearable and physiological sensing · 23% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
genomics |
0.5 | 2 | 2019 | Measurement error and variant-calling in deep Illumina sequencing of HIV · Bioinform. 2019 Toward a statistically explicit understanding of de novo sequence assembly · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis
sequencing error correction |
0.4 | 1 | 2019 | Measurement error and variant-calling in deep Illumina sequencing of HIV · Bioinform. 2019 |
Bioinformatics and computational biology › genomics
viral genomics |
0.4 | 1 | 2019 | Measurement error and variant-calling in deep Illumina sequencing of HIV · Bioinform. 2019 |
High-performance computing
parallel i/o |
0.3 | 2 | 2012 | Parallel I/O, analysis, and visualization of a trillion particle simulation · SC 2012 Parallel index and query for large scale data analysis · SC 2011 |
Bioinformatics and computational biology › sequence analysis › sequence assembly
genome assembly |
0.2 | 1 | 2013 | Toward a statistically explicit understanding of de novo sequence assembly · Bioinform. 2013 |
Bioinformatics and computational biology › biostatistics › statistical bioinformatics
statistical genomics |
0.2 | 1 | 2013 | Toward a statistically explicit understanding of de novo sequence assembly · Bioinform. 2013 |
Rendering › volume rendering
parallel volume rendering |
0.1 | 1 | 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core Systems · IEEE Trans. Vis. Comput. Graph. 2012 |
Rendering › volume rendering
ray casting |
0.1 | 1 | 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core Systems · IEEE Trans. Vis. Comput. Graph. 2012 |
Rendering
volume rendering |
0.1 | 1 | 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core Systems · IEEE Trans. Vis. Comput. Graph. 2012 |
Parallel and multicore computing › parallelization strategies
hybrid parallelism |
0.1 | 1 | 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core Systems · IEEE Trans. Vis. Comput. Graph. 2012 |
High-performance computing
particle simulation |
0.1 | 1 | 2012 | Parallel I/O, analysis, and visualization of a trillion particle simulation · SC 2012 |
High-performance computing
scientific data analysis |
0.1 | 1 | 2011 | Parallel index and query for large scale data analysis · SC 2011 |
Performance modeling and evaluation › parallel system performance
strong and weak scaling |
0.0 | 1 | 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core Systems · IEEE Trans. Vis. Comput. Graph. 2012 |
Wearable and physiological sensing › motion capture
hand motion tracking |
0.0 | 1 | 2011 | The mathematical imagery trainer: from embodied interaction to conceptual learning · CHI 2011 |
Distributed systems
distributed data processing |
0.0 | 1 | 2011 | Parallel index and query for large scale data analysis · SC 2011 |
Methods — techniques the papers use, named apart from their topics
variant calling pipeline · 0.4Primer ID consensus · 0.4shared-memory parallelism · 0.3distributed-memory parallelism · 0.3statistical modeling · 0.2query processing · 0.1indexing · 0.1design-based research · 0.1clinical interviews · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | T-shaped alignments integrating HIV-1 near full-length genome and partial pol sequences can improve phylogenetic inference of transmission clustersabstractMolecular epidemiology and HIV-1 transmission networks reconstruction can provide insights into transmission dynamics and inform public health strategies. Long HIV sequences, such as near full-length (nFL) genomes, can improve the accuracy of phylogenetic inference. However, relatively short pol sequences are still broadly used for inferring molecular HIV clusters. Whether a mix of long and short HIV-1 sequences can improve phylogenetic inference of molecular HIV clusters remains unknown. We propose a flexible approach called T-shaped alignments that incorporates both nFL HIV-1 genomes and partial pol sequences, and investigate whether this approach improves phylogenetic reconstruction of molecular clusters. Under the assumption that clustering from 100% of long sequences is the most accurate, we obtained 1196 subtype B nFL HIV-1 sequences from the Los Alamos National Laboratory Database and a single-study subset, varied the proportion of long and short sequences in our T-shape alignments, systematically masked all non-pol regions with missing characters in proportional increments, and compared tree similarity and cluster inference among datasets. With the full dataset, we found that when more than 50% of available sequences are nFL, the T-shaped alignment gradually yields results closer to the 100% n, with more and larger clusters identified. However, below the 50% threshold accuracy did not increase. Stringent bootstrap thresholds decreased cluster accuracy gaps but also decreased number of clusters found and mean cluster size. For the subset dataset, we found that the introduction of nFL sequences to the T-shaped alignment improves accuracy in clustering either after a 30% threshold or immediately depending on bootstrap choice. Our new approach and results suggest that using T-shape alignments to mix HIV-1 sequences of different lengths can improve phylogenetic and clustering accuracy, with needed nFL proportion depending on analysis goals. The T-shape alignment provides a straightforward method for utilizing all available sequences to improve phylogenetic analysis. August Guang, Casey W. Dunn, Vlad Novitsky, Mark Howison, Rami Kantor |
PLoS Comput. Biol. | 4 |
| 2019 | Measurement error and variant-calling in deep Illumina sequencing of HIVabstractMOTIVATION: Next-generation deep sequencing of viral genomes, particularly on the Illumina platform, is increasingly applied in HIV research. Yet, there is no standard protocol or method used by the research community to account for measurement errors that arise during sample preparation and sequencing. Correctly calling high and low-frequency variants while controlling for erroneous variants is an important precursor to downstream interpretation, such as studying the emergence of HIV drug-resistance mutations, which in turn has clinical applications and can improve patient care. RESULTS: We developed a new variant-calling pipeline, hivmmer, for Illumina sequences from HIV viral genomes. First, we validated hivmmer by comparing it to other variant-calling pipelines on real HIV plasmid datasets. We found that hivmmer achieves a lower rate of erroneous variants, and that all methods agree on the frequency of correctly called variants. Next, we compared the methods on an HIV plasmid dataset that was sequenced using Primer ID, an amplicon-tagging protocol, which is designed to reduce errors and amplification bias during library preparation. We show that the Primer ID consensus exhibits fewer erroneous variants compared to the variant-calling pipelines, and that hivmmer more closely approaches this low error rate compared to the other pipelines. The frequency estimates from the Primer ID consensus do not differ significantly from those of the variant-calling pipelines. AVAILABILITY AND IMPLEMENTATION: hivmmer is freely available for non-commercial use from https://github.com/kantorlab/hivmmer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mark Howison, Mia Coetzer, Rami Kantor |
Bioinform. | 1 |
| 2016 | Mining and Visualizing Sequential Patterns in the Electronic Health Record: A Case Study for Asthma With and Without Mental Disorders
Elizabeth S. Chen, Genevieve B. Melton, Mark Howison, Erik Knoll, Ashley S. Lee, Indra Neil Sarkar |
AMIA | 3 |
| 2013 | Building Software Environments for Research Computing Clusters
Mark Howison, Aaron Shen, Andrew Loomis |
LISA | 1 |
| 2013 | Toward a statistically explicit understanding of de novo sequence assemblyabstractMOTIVATION: Draft de novo genome assemblies are now available for many organisms. These assemblies are point estimates of the true genome sequences. Each is a specific hypothesis, drawn from among many alternative hypotheses, of the sequence of a genome. Assembly uncertainty, the inability to distinguish between multiple alternative assembly hypotheses, can be due to real variation between copies of the genome in the sample, errors and ambiguities in the sequenced data and assumptions and heuristics of the assemblers. Most assemblers select a single assembly according to ad hoc criteria, and do not yet report and quantify the uncertainty of their outputs. Those assemblers that do report uncertainty take different approaches to describing multiple assembly hypotheses and the support for each. RESULTS: Here we review and examine the problem of representing and measuring uncertainty in assemblies. A promising recent development is the implementation of assemblers that are built according to explicit statistical models. Some new assembly methods, for example, estimate and maximize assembly likelihood. These advances, combined with technical advances in the representation of alternative assembly hypotheses, will lead to a more complete and biologically relevant understanding of assembly uncertainty. This will in turn facilitate the interpretation of downstream analyses and tests of specific biological hypotheses. Mark Howison, Felipe Zapata 0001, Casey W. Dunn |
Bioinform. | 1 |
| 2013 | Agalma: an automated phylogenomics workflowabstractBACKGROUND: In the past decade, transcriptome data have become an important component of many phylogenetic studies. They are a cost-effective source of protein-coding gene sequences, and have helped projects grow from a few genes to hundreds or thousands of genes. Phylogenetic studies now regularly include genes from newly sequenced transcriptomes, as well as publicly available transcriptomes and genomes. Implementing such a phylogenomic study, however, is computationally intensive, requires the coordinated use of many complex software tools, and includes multiple steps for which no published tools exist. Phylogenomic studies have therefore been manual or semiautomated. In addition to taking considerable user time, this makes phylogenomic analyses difficult to reproduce, compare, and extend. In addition, methodological improvements made in the context of one study often cannot be easily applied and evaluated in the context of other studies. RESULTS: We present Agalma, an automated tool that constructs matrices for phylogenomic analyses. The user provides raw Illumina transcriptome data, and Agalma produces annotated assemblies, aligned gene sequence matrices, a preliminary phylogeny, and detailed diagnostics that allow the investigator to make extensive assessments of intermediate analysis steps and the final results. Sequences from other sources, such as externally assembled genomes and transcriptomes, can also be incorporated in the analyses. Agalma is built on the BioLite bioinformatics framework, which tracks provenance, profiles processor and memory use, records diagnostics, manages metadata, installs dependencies, logs version numbers and calls to external programs, and enables rich HTML reports for all stages of the analysis. Agalma includes a small test data set and a built-in test analysis of these data. In addition to describing Agalma, we here present a sample analysis of a larger seven-taxon data set. Agalma is available for download at https://bitbucket.org/caseywdunn/agalma. CONCLUSIONS: Agalma allows complex phylogenomic analyses to be implemented and described unambiguously as a series of high-level commands. This will enable phylogenomic studies to be readily reproduced, modified, and extended. Agalma also facilitates methods development by providing a complete modular workflow, bundled with test data, that will allow further optimization of each step in the context of a full phylogenomic analysis. Casey W. Dunn, Mark Howison, Felipe Zapata 0001 |
BMC Bioinform. | 2 |
| 2013 | High-Throughput Compression of FASTQ Data with SeqDBabstractCompression has become a critical step in storing next-generation sequencing (NGS) data sets because of both the increasing size and decreasing costs of such data. Recent research into efficiently compressing sequence data has focused largely on improving compression ratios. Yet, the throughputs of current methods now lag far behind the I/O bandwidths of modern storage systems. As biologists move their analyses to high-performance systems with greater I/O bandwidth, low-throughput compression becomes a limiting factor. To address this gap, we present a new storage model called SeqDB, which offers high-throughput compression of sequence data with minimal sacrifice in compression ratio. It achieves this by combining the existing multithreaded Blosc compressor with a new data-parallel byte-packing scheme, called SeqPack, which interleaves sequence data and quality scores. Mark Howison |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2012 | Parallel I/O, analysis, and visualization of a trillion particle simulationabstractPetascale plasma physics simulations have recently entered the regime of simulating trillions of particles. These unprecedented simulations generate massive amounts of data, posing significant challenges in storage, analysis, and visualization. In this paper, we present parallel I/O, analysis, and visualization results from a VPIC trillion particle simulation running on 120,000 cores, which produces ~30TB of data for a single timestep. We demonstrate the successful application of H5Part, a particle data extension of parallel HDF5, for writing the dataset at a significant fraction of system peak I/O rates. To enable efficient analysis, we develop hybrid parallel FastQuery to index and query data using multi-core CPUs on distributed memory hardware. We show good scalability results for the FastQuery implementation using up to 10,000 cores. Finally, we apply this indexing/query-driven approach to facilitate the first-ever analysis and visualization of the trillion particle dataset. Surendra Byna, Jerry Chou 0001, Oliver Rübel, Prabhat, Homa Karimabadi, William S. Daughton, Vadim Roytershteyn, E. Wes Bethel, Mark Howison, Ke-Jou Hsu, Kuan-Wu Lin, Arie Shoshani, Andrew Uselton, Kesheng Wu |
SC | 9 |
| 2012 | Hybrid Parallelism for Volume Rendering on Large-, Multi-, and Many-Core SystemsabstractWith the computing industry trending toward multi- and many-core processors, we study how a standard visualization algorithm, raycasting volume rendering, can benefit from a hybrid parallelism approach. Hybrid parallelism provides the best of both worlds: using distributed-memory parallelism across a large numbers of nodes increases available FLOPs and memory, while exploiting shared-memory parallelism among the cores within each node ensures that each node performs its portion of the larger calculation as efficiently as possible. We demonstrate results from weak and strong scaling studies, at levels of concurrency ranging up to 216,000, and with data sets as large as 12.2 trillion cells. The greatest benefit from hybrid parallelism lies in the communication portion of the algorithm, the dominant cost at higher levels of concurrency. We show that reducing the number of participants with a hybrid approach significantly improves performance. Mark Howison, E. Wes Bethel, Hank Childs |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | The mathematical imagery trainer: from embodied interaction to conceptual learningabstractWe introduce an embodied-interaction instructional design, the Mathematical Imagery Trainer (MIT), for helping young students develop grounded understanding of proportional equivalence (e.g., 2/3 = 4/6). Taking advantage of the low-cost availability of hand-motion tracking provided by the Nintendo Wii remote, the MIT applies cognitive-science findings that mathematical concepts are grounded in mental simulation of dynamic imagery, which is acquired through perceiving, planning, and performing actions with the body. We describe our rationale for and implementation of the MIT through a design-based research approach and report on clinical interviews with twenty-two 4th-6th grade students who engaged in problem-solving tasks with the MIT. Mark Howison, Dragan Trninic, Daniel L. Reinholz, Dor Abrahamson |
CHI | 1 |
| 2011 | Parallel index and query for large scale data analysisabstractModern scientific datasets present numerous data management and analysis challenges. State-of-the-art index and query technologies are critical for facilitating interactive exploration of large datasets, but numerous challenges remain in terms of designing a system for processing general scientific datasets. The system needs to be able to run on distributed multi-core platforms, efficiently utilize underlying I/O infrastructure, and scale to massive datasets. Jerry Chou 0001, Mark Howison, Brian Austin, Kesheng Wu, Ji Qiang, E. Wes Bethel, Arie Shoshani, Oliver Rübel, Prabhat, Robert D. Ryne |
SC | 2 |
| 2010 | Parallel I/O performance: From events to ensemblesabstractParallel I/O is fast becoming a bottleneck to the research agendas of many users of extreme scale parallel computers. The principle cause of this is the concurrency explosion of high-end computation, coupled with the complexity of providing parallel file systems that perform reliably at such scales. More than just being a bottleneck, parallel I/O performance at scale is notoriously variable, being influenced by numerous factors inside and outside the application, thus making it extremely difficult to isolate cause and effect for performance events. In this paper, we propose a statistical approach to understanding I/O performance that moves from the analysis of performance events to the exploration of performance ensembles. Using this methodology, we examine two I/O-intensive scientific computations from cosmology and climate science, and demonstrate that our approach can identify application and middleware performance deficiencies - resulting in more than 4× run time improvement for both examined applications. Andrew Uselton, Mark Howison, Nicholas J. Wright, David Skinner, Noel Keen, John Shalf, Karen L. Karavanic, Leonid Oliker |
IPDPS | 2 |