EDBT 2026 Demo / reviewers in the wild / expert
Simon Heybrock
dblp:32/7586
· DBLP profile ↗
1ranked-venue papers
1as first author
0since 2021 · last 2014
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 70% Memory systems · 23% Hardware accelerators and domain-specific architectures · 7% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory access optimization
data movement reduction |
0.2 | 1 | 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors · SC 2014 |
High-performance computing › scientific computing systems
lattice quantum chromodynamics |
0.2 | 1 | 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors · SC 2014 |
High-performance computing
performance optimization at scale |
0.2 | 1 | 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors · SC 2014 |
High-performance computing
scientific computing systems |
0.2 | 1 | 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors · SC 2014 |
Hardware accelerators and domain-specific architectures › many-core accelerator
many-core coprocessor |
0.1 | 1 | 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-Processors · SC 2014 |
Methods — techniques the papers use, named apart from their topics
mixed-precision arithmetic · 0.2domain decomposition · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Lattice QCD with Domain Decomposition on Intel® Xeon Phi Co-ProcessorsabstractThe gap between the cost of moving data and the cost of computing continues to grow, making it ever harder to design iterative solvers on extreme-scale architectures. This problem can be alleviated by alternative algorithms that reduce the amount of data movement. We investigate this in the context of Lattice Quantum Chromo dynamics and implement such an alternative solver algorithm, based on domain decomposition, on Intel®Xeon Phi co-processor (KNC) clusters. We demonstrate close-to-linear on-chip scaling to all 60 cores of the KNC. With a mix of single- and half-precision the domain-decomposition method sustains 400-500 Gflop/s per chip. Compared to an optimized KNC implementation of a standard solver [1], our full multi-node domain-decomposition solver strong-scales to more nodes and reduces the time-to-solution by a factor of 5. Simon Heybrock, Bálint Joó, Dhiraj D. Kalamkar, Mikhail Smelyanskiy, Karthikeyan Vaidyanathan, Tilo Wettig, Pradeep Dubey |
SC | 1 |