EDBT 2026 Demo / reviewers in the wild / expert
Vipin Sachdeva
dblp:25/4260
· DBLP profile ↗
12ranked-venue papers
4as first author
3since 2021 · last 2022
0000-0001-7914-0185ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Optimized Mappings for Symmetric Range-Limited Molecular Force Calculations on FPGAsabstractIn N-body applications, the efficient evaluation of range-limited forces depends on applying certain constraints, including a cut-off radius and force symmetry (Newton's Third Law). When computing the pair-wise forces in parallel, finding the optimal mapping of particles and computations to memories and processors is surprisingly challenging, but can result in greatly reduced data movement and computation. Despite FPGAs having a distinct compute model (BRAMs/network/pipelines) from CPUs and ASICs, mappings on FPGAs have not previously been studied in depth: it was thought that the half-shell method was preferred. In this work, we find that the Manhattan method is sur-prisingly compatible with FPGA hardware. With the cache overlapping technique proposed in this paper, the ultra-fine-grained data access demanded by the Manhattan method can be satisfied, despite the fact that the memory blocks on FPGAs appear to be insufficiently fine-grained. We further demonstrate that, compared to the traditional baseline half-shell method, approximately a half of the filters (preprocessors) can be removed without performance degradation. For communication, the amount of data transferred can be reduced by 40% - 75% in the most common multi-FPGA scenarios. Moreover, data transfers are almost perfectly balanced along all directions, and the optimization requires only minimal hardware resources. The practical consequence is that nearly 2 x to 4 x the workload can be handled without upgrading the network connections between FPGAs. This is a critical finding given the relatively limited bandwidth available in many common accelerator boards and the strong-scaling applications to which FPGA clusters are being applied. Chunshu Wu, Sahan Bandara, Tong Geng, Anqi Guo, Pouya Haghi, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
FPL | 6 |
| 2021 | Particle Mesh Ewald for Molecular Dynamics in OpenCL on an FPGA ClusterabstractMolecular Dynamics (MD) simulations play a central role in physics-driven drug discovery. MD applications often use the Particle Mesh Ewald (PME) algorithm to accelerate electrostatic force computations, but efficient parallelization has proven difficult due to the high communication requirements of distributed 3D FFTs. In this paper, we present the design and implementation of a scalable PME algorithm that runs on a cluster of Intel Stratix 10 FPGAs and can handle FFT sizes appropriate to address real-world drug discovery projects (grids up to 1283). To our knowledge, this is the first work to fully integrate all aspects of the PME algorithm (charge spreading, 3D FFT/IFFT, and force interpolation) within a distributed FPGA framework. The design is fully implemented with OpenCL for flexibility and ease of development and uses 100 Gbps links for direct FPGA-to-FPGA communications without the need for host interaction. We present experimental data up to 4 FPGAs (e.g., 206 microseconds per timestep for a 65536 atom simulation and 643 3D FFT), outperforming GPUs. Additionally, we discuss design scalability on clusters with differing topologies up to 64 FPGAs (with expected performance greater than all known GPU implementations) and integration with other hardware components to form a complete molecular dynamics application. We predict best-case performance of 6.6 microseconds per timestep on 64 FPGAs. Lawrence C. Stewart, Carlo Pascoe, Emery Davis, Brian W. Sherman, Martin C. Herbordt, Vipin Sachdeva |
FCCM | 6 |
| 2021 | Upgrade of FPGA Range-Limited Molecular Dynamics to Handle Hundreds of ProcessorsabstractWith the current pandemic, the central role that Molecular Dynamics simulation (MD) plays in drug discovery makes advances in MD performance urgent. Recent work has demonstrated that among COTS devices only FPGA-centric clusters can scale beyond a few processors for relevant targets; other work has shown that single FPGA performance compares favorably to that of a GPU. In this study we demonstrate that an additional factor of 4× performance can be achieved which results in a factor of 5× speed up over a GPU. The problem addressed is that the designs of the last decade no longer scale when the number of processing pipelines grows from around ten to the hundreds. We begin by systematically evaluating existing work, exposing its flaws, and proposing a series of new design solutions. There are four major contributions. First, we address the massive routing problem by augmenting the design with three minimal networks in logic and latency. Second, we have developed a novel asynchronous out-of-order communication mechanism that removes nearly all bubbles from the routing networks. Third, we find that inverting the standard particle access algorithm results in improved locality and performance. Finally, we have created a custom numerical format that increases precision while saving space and logic. Chunshu Wu, Tong Geng, Sahan Bandara, Chen Yang 0010, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
FCCM | 5 |
| 2019 | Molecular Dynamics Range-Limited Force Evaluation Optimized for FPGAsabstractFPGA Molecular Dynamics was much studied from 2004-2010. Due to limited chip resources of that era, and the inherent variety and complexity of tasks comprising Molecular Dynamics simulations (MD), those FPGA accelerators relied on host or embedded processors to organize and pre-process input and output data. This introduced long latency for data movement between simulation iterations and, as technology advanced, drastically limited performance. Current generation FPGAs are equipped not only with abundant on-chip resources, but also have hardware support for floating point operations; these advances provide an opportunity for creating self-contained MD simulation systems on a single device. In this paper, we demonstrate such a system based on the range-limited force, which comprises 90% of the flops in a typical MD simulation. It features online particle-pair generation, hundreds of force evaluation pipelines, motion update, and particle data migration. We integrate into OpenMM and find that, for a representative dataset (liquid argon with 20K atoms), we can achieve a simulation throughput of 1.4us/day with a single FPGA, more than twice the performance of a comparable generation GPU. The bulk of the work presented here explores the design of an independent MD range-limited force evaluation system tailored for modern FPGAs without data exchange with any off-chip devices. The primary contributions are the designs of the new features, the methods for coupling those features into an integrated system, and, especially, the analysis of the most likely mappings among particles/cells, on-chip memories (BRAMs), and on-chip compute units (pipelines). Chen Yang 0010, Tong Geng, Charles Lin, Jiayi Sheng, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
ASAP | 6 |
| 2019 | Fully integrated FPGA molecular dynamics simulationsabstractThe implementation of Molecular Dynamics (MD) on FPGAs has received substantial attention. Previous work, however, has consisted of either proof-of-concept implementations of components, usually the range-limited force; full systems, but with much of the work shared by the host CPU; or prototype demonstrations, e.g., using OpenCL, that neither implement a whole system nor have competitive performance. In this paper, we present what we believe to be the first full-scale FPGA-based simulation engine, and show that its performance is competitive with a GPU (running Amber in an industrial production environment). The system features on-chip particle data storage and management, short- and long-range force evaluation, as well as bonded forces, motion update, and particle migration. Other contributions of this work include exploring numerous architectural trade-offs and analysis of various mappings schemes among particles/cells and the various on-chip compute units. The potential impact is that this system promises to be the basis for long timescale Molecular Dynamics with a commodity cluster. Chen Yang 0010, Tong Geng, Rushi Patel, Qingqing Xiong, Ahmed Sanaullah, Chunshu Wu, Jiayi Sheng, Charles Lin, Vipin Sachdeva, Woody Sherman, Martin C. Herbordt |
SC | 10 |
| 2017 | K-mer clustering algorithm using a MapReduce framework: application to the parallelization of the Inchworm module of TrinityabstractBACKGROUND: De novo transcriptome assembly is an important technique for understanding gene expression in non-model organisms. Many de novo assemblers using the de Bruijn graph of a set of the RNA sequences rely on in-memory representation of this graph. However, current methods analyse the complete set of read-derived k-mer sequence at once, resulting in the need for computer hardware with large shared memory. RESULTS: We introduce a novel approach that clusters k-mers as the first step. The clusters correspond to small sets of gene products, which can be processed quickly to give candidate transcripts. We implement the clustering step using the MapReduce approach for parallelising the analysis of large datasets, which enables the use of compute clusters. The computational task is distributed across the compute system using the industry-standard MPI protocol, and no specialised hardware is required. Using this approach, we have re-implemented the Inchworm module from the widely used Trinity pipeline, and tested the method in the context of the full Trinity pipeline. Validation tests on a range of real datasets show large reductions in the runtime and per-node memory requirements, when making use of a compute cluster. CONCLUSIONS: Our study shows that MapReduce-based clustering has great potential for distributing challenging sequencing problems, without loss of accuracy. Although we have focussed on the Trinity package, we propose that such clustering is a useful initial step for other assembly pipelines. Chang Sik Kim, Martyn D. Winn, Vipin Sachdeva, Kirk E. Jordan |
BMC Bioinform. | 3 |
| 2014 | K-mer clustering algorithm using a MapReduce approachabstractWith recent advances in high throughput sequencing platforms, it is possible to sequence RNA obtained from biological samples more cost-effectively and comprehensively. Due to the ubiquity of the technology, massive volumes of RNA sequence data are now being generated, and as a result the need for more efficient analysis software has become an urgent challenge. Chang Sik Kim, Martyn D. Winn, Vipin Sachdeva, Kirk E. Jordan |
BIBM | 3 |
| 2014 | Exploring HPC-based scientific software as a service using CometCloudabstractThe use of in-silico simulations in experimental science can greatly increase laboratory efficiency and provide additional insights into interactions not easily described by traditional methods. Such simulations require significant amounts of computational resources, accessible only via supercom Moustafa AbdelBaky, Javier Diaz Montes, Michael Johnston, Vipin Sachdeva, Richard L. Anderson 0003, Kirk E. Jordan, Manish Parashar |
CollaborateCom | 4 |
| 2010 | Evaluating Cell/B.E software cache for ClustalWabstractThis paper evaluates the performance of the bioinformatics application ClustalW developed on Cell Broadband Engine(TM) (Cell/B.E.) using a software data cache for SPEs, instead of explicit DMA transfers. The software cache of the SPEs, once it has been configured, provides the capability to access main memory with data-transfer functions that override the need for DMA commands. ClustalW exhibits high spatial locality but little temporal locality. We compare performance of ClustalW with a previous version that uses explicit DMA transfers as a means of communication with the system memory, and provide analysis and results of our comparison. Vipin Sachdeva, Michael Kistler, David A. Bader |
ISCAS | 1 |
| 2008 | Exploring the viability of the Cell Broadband Engine for bioinformatics applications
Vipin Sachdeva, Michael Kistler, William Evan Speight, Tzy-Hwa Kathy Tzeng |
Parallel Comput. | 1 |
| 2007 | Exploring the Viability of the Cell Broadband Engine for Bioinformatics ApplicationsabstractThis paper evaluates the performance of bioinformatics applications on the Cell Broadband Engine recently developed at IBM. In particular we focus on two highly popular bioinformatics applications - FASTA and ClustalW. The characteristics of these bioinformatics applications, such as small critical time-consuming code size, regular memory accesses, existing vectorized code and embarrassingly parallel computation, make them uniquely suitable for the Cell processing platform. The price and power advantages afforded by the Cell processor also make it an attractive alternative to general purpose processors. We report preliminary performance results for these applications, and contrast these results with the state-of-the-art hardware. Vipin Sachdeva, Michael Kistler, William Evan Speight, Tzy-Hwa Kathy Tzeng |
IPDPS | 1 |
| 2006 | Poster reception - Pairwise alignments on the cell processorabstractWe are evaluating the performance of bioinformatics applications on the CBE recently developed at IBM. The characteristics of pairwise alignment, critical to bioinformatics applications, makes it uniquely suitable for the Cell processor. We have implemented two different strategies for pairwise alignment on Cell. We analyze the bottlenecks for each strategy, a diagonal approach which uses the vector processing capabilities of the SPE and a row wise approach in which cells are dependant and has expensive branches. Despite the presence of branches, the row-wise approach outperforms a 2.4 Ghz Opteron by almost 5X, the diagonal approach spending a majority of time looking up the alignment matrix.Also, due to the commonality of all-to-all pairwise comparisons in bioinformatics codes, we have implemented a framework for these computations on the Cell with minimum load imbalance and minimal PPU intervention. We evaluate the performance of popular bioinformatics codes Clustalw and HMMER using such a framework, in which the workload is shared among the PPU and the SPUs. Vipin Sachdeva, Michael Kistler, William Evan Speight |
SC | 1 |