EDBT 2026 Demo / reviewers in the wild / expert
Ilya Sharapov
dblp:70/1830
· DBLP profile ↗
8ranked-venue papers
1as first author
3since 2021 · last 2024
0009-0004-8980-6170ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Breaking the Molecular Dynamics Timescale Barrier Using a Wafer-Scale SystemabstractMolecular dynamics (MD) simulations have transformed our understanding of the nanoscale, driving breakthroughs in materials science, computational chemistry, and several other fields, including biophysics and drug design. Even on exascale supercomputers, however, runtimes are excessive for systems and timescales of scientific interest. Here, we demonstrate strong scaling of MD simulations on the Cerebras Wafer-Scale Engine. By dedicating a processor core for each simulated atom, we demonstrate a 457-fold improvement in timesteps per second versus the Frontier GPU-based Exascale platform, along with a large improvement in timesteps per unit energy. Reducing every year of runtime to less than a day unlocks currently inaccessible timescales of slow microstructure transformation processes that are critical for understanding material behavior and function.Our dataflow algorithm runs Embedded Atom Method (EAM) simulations at rates over 699k timesteps per second for problems with up to 800k atoms. This demonstrated performance is unprecedented for general-purpose processing cores. Kylee Santos, Stan G. Moore, Tomas Oppelstrup, Amirali Sharifian, Ilya Sharapov, Aidan P. Thompson, Delyan Kalchev, Danny Perez, Robert Schreiber, Scott Pakin, Edgar A. León, James H. Laros III, Michael James 0002, Sivasankaran Rajamanickam |
SC | 5 |
| 2023 | Wafer-Scale Fast Fourier TransformsabstractWe have implemented fast Fourier transforms for one, two, and three-dimensional arrays on the Cerebras CS-2, a system whose memory and processing elements reside on a single silicon wafer. The wafer-scale engine (WSE) encompasses a two-dimensional mesh of roughly 850,000 processing elements (PEs) with fast local memory and equally fast nearest-neighbor interconnections. Marcelo Orenes-Vera, Ilya Sharapov, Robert S. Schreiber, Mathias Jacquelin, Philippe Vandermersch, Sharan Chetlur |
ICS | 2 |
| 2021 | ISPD 2021 Wafer-Scale Physics Modeling Contest: A New Frontier for Partitioning, Placement and RoutingabstractSolving 3-D partial differential equations in a Finite Element model is computationally intensive and requires extremely high memory and communication bandwidth. This paper describes a novel way where the Finite Element mesh points of varying resolution are mapped on a large 2-D homogenous array of processors. Cerebras developed a novel supercomputer that is powered by a 21.5cm by 21.5cm Wafer-Scale Engine (WSE) with 850,000 programmable compute cores. With 2.6 trillion transistors in a 7nm process this is by far the largest chip in the world. It is structured as a regular array of 800 by 1060 identical processing elements, each with its own local fast SRAM memory and direct high bandwidth connection to its neighboring cores. For the 2021 ISPD competition we propose a challenge to optimize placement of computational physics problems to achieve the highest possible performance on the Cerebras supercomputer. The objectives are to maximize performance and accuracy by optimizing the mapping of the problem to cores in the system. This involves partitioning and placement algorithms. Patrick Groeneveld, Michael James 0002, Vladimir Kibardin, Ilya Sharapov, Marvin Tom, Leo Wang |
ISPD | 4 |
| 2020 | Fast stencil-code computation on a wafer-scale processorabstractThe performance of CPU-based and GPU-based systems is often low for PDE codes, where large, sparse, and often structured systems of linear equations must be solved. Iterative solvers are limited by data movement, both between caches and memory and between nodes. Here we describe the solution of such systems of equations on the Cerebras Systems CS-1, a wafer-scale processor that has the memory bandwidth and communication latency to perform well. We achieve 0.86 PFLOPS on a single wafer-scale system for the solution by BiCGStab of a linear system arising from a 7-point finite difference stencil on a 600 × 595 × 1536 mesh, achieving about one third of the machine's peak performance. We explain the system, its architecture and programming, and its performance on this problem and related problems. We discuss issues of memory capacity and floating point precision. We outline plans to extend this work towards full applications. Kamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber, Michael Morrison, Vladimir Kibardin, Andrey Portnoy, Jean-Francois Dietiker, Madhava Syamlal, Michael James 0002 |
SC | 3 |
| 2019 | Online Normalization for Training Neural NetworksabstractOnline Normalization is a new technique for normalizing the hidden activations of a neural network. Like Batch Normalization, it normalizes the sample dimension. While Online Normalization does not use batches, it is as accurate as Batch Normalization. We resolve a theoretical limitation of Batch Normalization by introducing an unbiased technique for computing the gradient of normalized activations. Online Normalization works with automatic differentiation by adding statistical normalization as a primitive. This technique can be used in cases not covered by some other normalizers, such as recurrent networks, fully connected networks, and networks with activation memory requirements prohibitive for batching. We show its applications to image classification, image segmentation, and language modeling. We present formal proofs and experimental results on ImageNet, CIFAR, and PTB datasets. Vitaliy Chiley, Ilya Sharapov, Atli Kosson, Urs Köster, Ryan Reece, Sofia Samaniego de la Fuente, Vishal Subbiah, Michael James 0002 |
NeurIPS | 2 |
| 2013 | Performance Evaluation of NAS Parallel Benchmarks on Intel Xeon PhiabstractNAS parallel benchmarks (NPB) are a set of applications commonly used to evaluate parallel systems. We use the NPB-OpenMP version to examine the performance of the Intel's new Xeon Phi co-processor and focus in particular on the many core aspect of the Xeon Phi architecture. A first analysis studies the scalability up to 244 threads on 61 cores and the impact of affinity settings on scaling. It also compares performance characteristics of Xeon Phi and traditional Xeon CPUs. The application of several well-established optimization techniques allows us to identify common bottlenecks that can specifically impede performance on the Xeon Phi but are not as severe on multi-core CPUs. We also find that many of the OpenMP-parallel loops are too short (in terms of the number of loop iterations) for a balanced execution by 244 threads. New or redesigned benchmarks will be needed to accommodate the greatly increased number of cores and threads. At the end, we summarize our findings in a set recommendations for performance optimization for Xeon Phi. Arunmoezhi Ramachandran, Jérôme Vienne, Rob F. Van der Wijngaart, Lars Koesterke, Ilya Sharapov |
ICPP | 5 |
| 2007 | Characteristics of workloads used in high performance and technical computingabstractThis paper provides a systematic comparison of various characteristics of computationally-intensive workloads. Our analysis focuses on standard HPC benchmarks and representative applications. For the selected workloads we provide a wide range of characterizations based on instruction tracing and hardware counter measurements. Razvan Cheveresan, Matthew Ramsay, Chris Feucht, Ilya Sharapov |
ICS | 4 |
| 2006 | A case study in top-down performance estimation for a large-scale parallel applicationabstractThis work presents a general methodology for estimating the performance of an HPC workload when running on a future hardware architecture. Further, it demonstrates the methodology by estimating the performance of a significant scientific application -- the Gyrokinetic Toroidal Code (GTC) -- when executing on Sun's proposed next-generation petascale computer architecture.For GTC, we identify the important phases of the iteration and perform low-level analysis that includes instruction tracing and component simulations of processor and memory systems. Low-level analysis is complemented with scalability estimates based on modeling MPI, OpenMP and I/O activity in the code. The work's approach permits accurate end-to-end performance projections from the microarchitecture level to the petascale. Ilya Sharapov, Robert Kroeger, Guy Delamarter, Razvan Cheveresan, Matthew Ramsay |
PPoPP | 1 |