EDBT 2026 Demo / reviewers in the wild / expert
Jun Doi
dblp:97/3446
· DBLP profile ↗
4ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
High-performance computing · 70% Parallel and multicore computing · 13% Interconnection networks and networks-on-chip · 13% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing systems
lattice quantum chromodynamics |
0.1 | 1 | 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputer · SC 2012 |
High-performance computing
performance optimization |
0.1 | 1 | 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputer · SC 2012 |
High-performance computing
scientific computing |
0.1 | 1 | 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputer · SC 2012 |
Parallel and multicore computing › data parallelism
SIMD vectorization |
0.1 | 1 | 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputer · SC 2012 |
Interconnection networks and networks-on-chip › interprocessor communication
all-to-all communication |
0.1 | 1 | 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel Supercomputers · SC 2010 |
High-performance computing
collective communication |
0.1 | 1 | 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel Supercomputers · SC 2010 |
High-performance computing
fast fourier transform |
0.1 | 1 | 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel Supercomputers · SC 2010 |
High-performance computing › fast fourier transform
parallel FFT |
0.1 | 1 | 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel Supercomputers · SC 2010 |
Distributed systems › communication optimization
communication-computation overlap |
0.0 | 1 | 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputer · SC 2012 |
Interconnection networks and networks-on-chip › network topology
torus network |
0.0 | 1 | 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel Supercomputers · SC 2010 |
Methods — techniques the papers use, named apart from their topics
pipelining · 0.3data layout optimization · 0.1SIMD · 0.1shared memory parallel threads · 0.1overlap · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Quantum computing simulator on a heterogenous HPC systemabstractQuantum computing simulation on a classical computer is difficult due to the exponential runtime and memory overhead. Previous work addresses the difficulty by utilizing multiple Graphical Processing Units (GPUs) and multi-node computers. GPUs are efficient for handling runtime issues but have limited total accessible memory space. Meanwhile, the memory of a multi-node computer can be scaled to the petabytes order, but its bandwidth for access from host computers (CPUs) is narrow. To simultaneously accelerate simulation and enlarge the total memory space, we propose a heterogeneous parallelization approach by combining GPUs and CPUs. Our simulator allocates memory to the GPUs first, and then to the CPUs. It thus accelerates simulation by using the full capabilities of the GPUs if memory for the simulation fits in the GPUs on a cluster. Allocating memory to the CPUs reduces benefits of the GPUs but enlarges the capacity of qubits in the simulation. In such case, it can exploit the memory of the GPUs to add one more qubit in the simulation if the size of memory in a node is the power of two (such as 512GB). We show empirical performance evaluations of our simulator in a distributed environment of POWER9. Jun Doi, Hitomi Takahashi, Raymond H. Putra, Takashi Imamichi, Hiroshi Horii |
CF | 1 |
| 2012 | Peta-scale lattice quantum chromodynamics on a blue gene/Q supercomputerabstractLattice Quantum Chromodynamics (QCD) is one of the most challenging applications running on massively parallel supercomputers. To reproduce these physical phenomena on a supercomputer, a precise simulation is demanded requiring well optimized and scalable code. We have optimized lattice QCD programs on Blue Gene family supercomputers and shown the strength in lattice QCD simulation. Here we optimized on the third generation Blue Gene/Q supercomputer; i) by changing the data layout, ii) by exploiting new SIMD instruction sets, and iii) by pipelining boundary data exchange to overlap communication and calculation. The optimized lattice QCD program shows excellent weak scalability on the large scale Blue Gene/Q system, and with 16 racks we sustained 1.08 Pflop/s, 32.1% of the theoretical peak performance, including the conjugate gradient solver routines. Jun Doi |
SC | 1 |
| 2010 | Overlapping Methods of All-to-All Communication and FFT Algorithms for Torus-Connected Massively Parallel SupercomputersabstractTorus networks are commonly used for massively parallel computers, its performance often becomes the constraint on total application performance. Especially in an asymmetric torus network, network traffic along the longest axis is the performance bottleneck for all-to-all communication, so that it is important to schedule the longest-axis traffic smoothly. In this paper, we propose a new algorithm based on an indirect method for pipelining the all-to-all procedures using shared memory parallel threads, which (1) isolates the longest-axis traffic from other traffic, (2) schedules it smoothly and (3) overlaps all of the other traffic and overhead for the all-to-all communication behind the longest-axis traffic. The proposed method achieves up to 95% of the theoretical peak. We integrated the overlapped all-to-all method with parallel FFT algorithms. And local FFT calculations are also overlapped behind the longest-axis traffic. The FFT performance achieves up to 90% of the theoretical peak for the parallel 1D FFT. Jun Doi, Yasushi Negishi |
SC | 1 |
| 2004 | Efficient method of adaptive sign detection for 4 ast 4 determinants using a standard arithmetic processing unit
Toshiya Yamauchi, Norimasa Yoshida, Jun Doi, Fujio Yamaguchi |
Vis. Comput. | 3 |