Marghoob Mohiyuddin

dblp:52/5405 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 47% Hardware accelerators and domain-specific architectures · 27% Energy-efficient computing · 19%
Interdisciplinary, comprehensive, and emerging computing
4 papers
Bioinformatics and computational biology · 100%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 16 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › genomics › structural variation
structural variant detection
0.422015
VarSim: a high-fidelity simulation and validation framework for high-throughput genome sequencing with cancer applications · Bioinform. 2015
MetaSV: an accurate and integrative structural-variant caller for next generation sequencing · Bioinform. 2015
Bioinformatics and computational biology
sequencing simulation
0.212016
LongISLND: in silico sequencing of lengthy and noisy datatypes · Bioinform. 2016
Energy-efficient computing
energy-efficient architecture
0.222011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011
A design methodology for domain-optimized power-efficient supercomputing · SC 2009
Bioinformatics and computational biology › sequence analysis › read mapping
short read alignment
0.112012
Fast and accurate read alignment for resequencing · Bioinform. 2012
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms
0.112009
Minimizing communication in sparse matrix solvers · SC 2009
Hardware accelerators and domain-specific architectures
domain-specific optimization
0.112009
A design methodology for domain-optimized power-efficient supercomputing · SC 2009
High-performance computing
parallel numerical algorithms
0.112009
Minimizing communication in sparse matrix solvers · SC 2009
High-performance computing
sparse linear solver
0.112009
Minimizing communication in sparse matrix solvers · SC 2009
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.112009
Minimizing communication in sparse matrix solvers · SC 2009
Bioinformatics and computational biology › genomics
variant calling
0.112016
LongISLND: in silico sequencing of lengthy and noisy datatypes · Bioinform. 2016
Electronic design automation › physical design
interconnect optimization
0.012004
Synthesizing interconnect-efficient low density parity check codes · DAC 2004
Electronic design automation
physical design
0.012004
Synthesizing interconnect-efficient low density parity check codes · DAC 2004
Coding theory
error-correcting codes
0.012004
Synthesizing interconnect-efficient low density parity check codes · DAC 2004
Coding theory › error-correcting codes
LDPC codes
0.012004
Synthesizing interconnect-efficient low density parity check codes · DAC 2004
High-performance computing
scientific computing systems
0.012011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011
High-performance computing › scientific computing systems
seismic imaging
0.012011
Hardware/software co-design for energy-efficient seismic modeling · SC 2011

Methods — techniques the papers use, named apart from their topics

consensus building · 0.2soft-clip analysis · 0.2local assembly · 0.2dynamic programming · 0.2performance and power modeling · 0.1FPGA-accelerated architectural simulation · 0.1krylov subspace method · 0.1cache blocking · 0.1auto-tuning · 0.1architecture space exploration · 0.1GMRES · 0.1heuristic synthesis · 0.1
YearPublicationVenuePosition
2016 LongISLND: in silico sequencing of lengthy and noisy datatypes
abstract
LongISLND is a software package designed to simulate sequencing data according to the characteristics of third generation, single-molecule sequencing technologies. The general software architecture is easily extendable, as demonstrated by the emulation of Pacific Biosciences (PacBio) multi-pass sequencing with P5 and P6 chemistries, producing data in FASTQ, H5, and the latest PacBio BAM format. We demonstrate its utility by downstream processing with consensus building and variant calling. AVAILABILITY AND IMPLEMENTATION: LongISLND is implemented in Java and available at http://bioinform.github.io/longislnd CONTACT: [email protected] information: Supplementary data are available at Bioinformatics online.
Bayo Lau, Marghoob Mohiyuddin, John C. Mu, Li Tai Fang, Narges Bani Asadi, Carolina Dallett, Hugo Y. K. Lam
Bioinform.2
2015 MetaSV: an accurate and integrative structural-variant caller for next generation sequencing
abstract
UNLABELLED: Structural variations (SVs) are large genomic rearrangements that vary significantly in size, making them challenging to detect with the relatively short reads from next-generation sequencing (NGS). Different SV detection methods have been developed; however, each is limited to specific kinds of SVs with varying accuracy and resolution. Previous works have attempted to combine different methods, but they still suffer from poor accuracy particularly for insertions. We propose MetaSV, an integrated SV caller which leverages multiple orthogonal SV signals for high accuracy and resolution. MetaSV proceeds by merging SVs from multiple tools for all types of SVs. It also analyzes soft-clipped reads from alignment to detect insertions accurately since existing tools underestimate insertion SVs. Local assembly in combination with dynamic programming is used to improve breakpoint resolution. Paired-end and coverage information is used to predict SV genotypes. Using simulation and experimental data, we demonstrate the effectiveness of MetaSV across various SV types and sizes. AVAILABILITY AND IMPLEMENTATION: Code in Python is at http://bioinform.github.io/metasv/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marghoob Mohiyuddin, John C. Mu, Narges Bani Asadi, Mark Gerstein, Alexej Abyzov, Wing H. Wong, Hugo Y. K. Lam
Bioinform.1
2015 VarSim: a high-fidelity simulation and validation framework for high-throughput genome sequencing with cancer applications
abstract
SUMMARY: VarSim is a framework for assessing alignment and variant calling accuracy in high-throughput genome sequencing through simulation or real data. In contrast to simulating a random mutation spectrum, it synthesizes diploid genomes with germline and somatic mutations based on a realistic model. This model leverages information such as previously reported mutations to make the synthetic genomes biologically relevant. VarSim simulates and validates a wide range of variants, including single nucleotide variants, small indels and large structural variants. It is an automated, comprehensive compute framework supporting parallel computation and multiple read simulators. Furthermore, we developed a novel map data structure to validate read alignments, a strategy to compare variants binned in size ranges and a lightweight, interactive, graphical report to visualize validation results with detailed statistics. Thus far, it is the most comprehensive validation tool for secondary analysis in next generation sequencing. AVAILABILITY AND IMPLEMENTATION: Code in Java and Python along with instructions to download the reads and variants is at http://bioinform.github.io/varsim. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
John C. Mu, Marghoob Mohiyuddin, Narges Bani Asadi, Mark Gerstein, Alexej Abyzov, Wing H. Wong, Hugo Y. K. Lam
Bioinform.2
2012 Fast and accurate read alignment for resequencing
abstract
MOTIVATION: Next-generation sequence analysis has become an important task both in laboratory and clinical settings. A key stage in the majority sequence analysis workflows, such as resequencing, is the alignment of genomic reads to a reference genome. The accurate alignment of reads with large indels is a computationally challenging task for researchers. RESULTS: We introduce SeqAlto as a new algorithm for read alignment. For reads longer than or equal to 100 bp, SeqAlto is up to 10 × faster than existing algorithms, while retaining high accuracy and the ability to align reads with large (up to 50 bp) indels. This improvement in efficiency is particularly important in the analysis of future sequencing data where the number of reads approaches many billions. Furthermore, SeqAlto uses less than 8 GB of memory to align against the human genome. SeqAlto is benchmarked against several existing tools with both real and simulated data. AVAILABILITY: Linux and Mac OS X binaries free for academic use are available at http://www.stanford.edu/group/wonglab/seqalto CONTACT: [email protected].
John C. Mu, Hui Jiang 0002, Amirhossein Kiani, Marghoob Mohiyuddin, Narges Bani Asadi, Wing Hung Wong
Bioinform.4
2011 Hardware/software co-design for energy-efficient seismic modeling
abstract
Reverse Time Migration (RTM) has become the standard for high-quality imaging in the seismic industry. RTM relies on PDE solutions using stencils that are 8th order or larger, which require large-scale HPC clusters to meet the computational demands. However, the rising power consumption of conventional cluster technology has prompted investigation of architectural alternatives that offer higher computational efficiency. In this work, we compare the performance and energy efficiency of three architectural alternatives -- the Intel Nehalem X5530 multicore processor, the NVIDIA Tesla C2050 GPU, and a general-purpose manycore chip design optimized for high-order wave equations called "Green Wave." We have developed an FPGA-accelerated architectural simulation platform to accurately model the power and performance of the Green Wave design. Results show that across highly-tuned high-order RTM stencils, the Green Wave implementation can offer up to 8x and 3.5x energy efficiency improvement per node respectively, compared with the Nehalem and GPU platforms. These results point to the enormous potential energy advantages of our hardware/software co-design methodology.
Jens Krueger, David Donofrio, John Shalf, Marghoob Mohiyuddin, Samuel Williams 0001, Leonid Oliker, Franz-Josef Pfreundt
SC4
2009 Analysis of photonic networks for a chip multiprocessor using scientific applications
abstract
As multiprocessors scale to unprecedented numbers of cores in order to sustain performance growth, it is vital that these gains are not nullified by high energy consumption from inter-core communication. With recent advances in 3D Integration CMOS technology, the possibility for realizing hybrid photonic-electronic networks-on-chip warrants investigating real application traces on functionally comparable photonic and electronic network designs. We present a comparative analysis using both synthetic benchmarks as well as real applications, run through detailed cycle accurate models implemented under the OMNeT++ discrete event simulation environment. Results show that when utilizing standard process-to-processor mapping methods, this hybrid network can achieve 75times improvement in energy efficiency for synthetic benchmarks and up to 37times improvement for real scientific applications, defined as network performance per energy spent, over an electronic mesh for large messages across a variety of communication patterns.
Gilbert Hendry, Shoaib Kamil 0001, Aleksandr Biberman, Johnnie Chan, Benjamin G. Lee, Marghoob Mohiyuddin, Keren Bergman, Luca P. Carloni, John Kubiatowicz, Leonid Oliker, John Shalf
NOCS6
2009 Minimizing communication in sparse matrix solvers
abstract
Data communication within the memory system of a single processor node and between multiple nodes in a system is the bottleneck in many iterative sparse matrix solvers like CG and GMRES. Here k iterations of a conventional implementation perform k sparse-matrix-vector-multiplications and Ω(k) vector operations like dot products, resulting in communication that grows by a factor of Ω(k) in both the memory and network. By reorganizing the sparse-matrix kernel to compute a set of matrix-vector products at once and reorganizing the rest of the algorithm accordingly, we can perform k iterations by sending O(log P) messages instead of O(k · log P) messages on a parallel machine, and reading the matrix A from DRAM to cache just once, instead of k times on a sequential machine. This reduces communication to the minimum possible. We combine these techniques to form a new variant of GMRES. Our shared-memory implementation on an 8-core Intel Clovertown gets speedups of up to 4.3x over standard GMRES, without sacrificing convergence rate or numerical stability.
Marghoob Mohiyuddin, Mark Hoemmen, James Demmel, Katherine A. Yelick
SC1
2009 A design methodology for domain-optimized power-efficient supercomputing
abstract
As power has become the pre-eminent design constraint for future HPC systems, computational efficiency is being emphasized over simply peak performance. Recently, static benchmark codes have been used to find a power efficient architecture. Unfortunately, because compilers generate sub-optimal code, benchmark performance can be a poor indicator of the performance potential of architecture design points. Therefore, we present hardware/software cotuning as a novel approach for system design, in which traditional architecture space exploration is tightly coupled with software auto-tuning for delivering substantial improvements in area and power efficiency. We demonstrate the proposed methodology by exploring the parameter space of a Tensilica-based multi-processor running three of the most heavily used kernels in scientific computing, each with widely varying micro-architectural requirements: sparse matrix vector multiplication, stencil-based computations, and general matrix-matrix multiplication. Results demonstrate that co-tuning significantly improves hardware area and energy efficiency -- a key driver for next generation of HPC system design.
Marghoob Mohiyuddin, Mark Murphy, Leonid Oliker, John Shalf, John Wawrzynek, Samuel Williams 0001
SC1
2008 Avoiding communication in sparse matrix computations
abstract
The performance of sparse iterative solvers is typically limited by sparse matrix-vector multiplication, which is itself limited by memory system and network performance. As the gap between computation and communication speed continues to widen, these traditional sparse methods will suffer. In this paper we focus on an alternative building block for sparse iterative solvers, the "matrix powers kernel" [x, Ax, A2x, ..., Akx], and show that by organizing computations around this kernel, we can achieve near-minimal communication costs. We consider communication very broadly as both network communication in parallel code and memory hierarchy access in sequential code. In particular, we introduce a parallel algorithm for which the number of messages (total latency cost) is independent of the power k, and a sequential algorithm, that reduces both the number and volume of accesses, so that it is independent of k in both latency and bandwidth costs. This is part of a larger project to develop "communication-avoiding Krylov subspace methods," which also addresses the numerical issues associated with these methods. Our algorithms work for general sparse matrices that "partition well". We introduce parallel performance models of matrices arising from 2D and 3D problems and show predicted speedups over a conventional algorithm of up to 7times on a petaflop-scale machine and up to 22times on computation across the grid. Analogous sequential performance models of the same problems predict speedups over a conventional algorithm of up to 10times on an out-of-core implementation, and up to 2.5times when we use our ideas to reduce off-chip latency and bandwidth to DRAM. Finally, we validate the model on an out-of-core sequential implementation and measured a speedup of over 3times, which is close to the predicted speedup.
James Demmel, Mark Hoemmen, Marghoob Mohiyuddin, Katherine A. Yelick
IPDPS3
2004 Synthesizing interconnect-efficient low density parity check codes
abstract
Error correcting codes are widely used in communication and storage applications. Codec complexity has usually been measured with a software implementation in mind. A recent hardware implementation of a Low Density Parity Check code (LDPC) indicates that interconnect complexity dominates the VLSI cost. We describe a heuristic interconnect-aware synthesis algorithm which generates LDPC codes that use an order of magnitude less wiring with little or no loss of coding efficiency.
Marghoob Mohiyuddin, Adnan Aziz, Marilyn Wolf
DAC1