EDBT 2026 Demo / reviewers in the wild / expert
Chansup Byun
dblp:116/9288
· DBLP profile ↗
5ranked-venue papers
0as first author
1since 2021 · last 2022
0009-0003-0183-914XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 60% GPUs and heterogeneous computing · 39% High-performance computing · 1% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › workload characterization
AI workload characterization |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
GPUs and heterogeneous computing › GPU computing
GPU-accelerated systems |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
Performance modeling and evaluation
workload characterization |
0.6 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 1 | 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and Implications · HPCA 2022 |
High-performance computing
domain decomposition |
0.0 | 1 | 1997 | A multi-level parallelization concept for high-fidelity multi-block solvers · SC 1997 |
Parallel and multicore computing › parallel programming models › structured parallelism
hierarchical parallelism |
0.0 | 1 | 1997 | A multi-level parallelization concept for high-fidelity multi-block solvers · SC 1997 |
Methods — techniques the papers use, named apart from their topics
user behavior analysis · 0.6job trace analysis · 0.6domain decomposition · 0.0data partitioning · 0.0data coalescing · 0.0MPI · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | AI-Enabling Workloads on Large-Scale GPU-Accelerated System: Characterization, Opportunities, and ImplicationsabstractProduction high-performance computing (HPC) systems are adopting and integrating GPUs into their design to accommodate artificial intelligence (AI), machine learning, and data visualization workloads. To aid with the design and operations of new and existing GPU-based large-scale systems, we provide a detailed characterization of system operations, job characteristics, user behavior, and trends on a contemporary GPU-accelerated production HPC system. Our insights indicate that the pre-mature phases in modern AI workflow take up significant GPU hours while underutilizing GPUs, which opens up the opportunity for a multi-tier system. Finally, we provide various potential recommendations and areas for future investment for system architects, operators, and users. Baolin Li 0001, Rohin Arora, Siddharth Samsi, Tirthak Patel, William Arcand, David Bestor, Chansup Byun, Rohan Basu Roy, Bill Bergeron, John T. Holodnak, Michael Houle 0001, Matthew Hubbell, Michael Jones 0001, Jeremy Kepner, Anna Klein, Peter Michaleas, Joseph McDonald, Lauren Milechin, Julie Mullen, Andrew Prout, Benjamin Price, Albert Reuther, Antonio Rosa, Matthew L. Weiss, Charles Yee, Daniel Edelman, Allan Vanterpool, Anson Cheng, Vijay Gadepally, Devesh Tiwari |
HPCA | 7 |
| 2018 | Scalable system scheduling for HPC and big data
Albert Reuther, Chansup Byun, William Arcand, David Bestor, Bill Bergeron, Matthew Hubbell, Michael Jones 0001, Peter Michaleas, Andrew Prout, Antonio Rosa, Jeremy Kepner |
J. Parallel Distributed Comput. | 2 |
| 2017 | Learning by doing, High Performance Computing education in the MOOC era
Julie Mullen, Chansup Byun, Vijay Gadepally, Siddharth Samsi, Albert Reuther, Jeremy Kepner |
J. Parallel Distributed Comput. | 2 |
| 2012 | Dynamic distributed dimensional data model (D4M) database and computation systemabstractA crucial element of large web companies is their ability to collect and analyze massive amounts of data. Tuple store databases are a key enabling technology employed by many of these companies (e.g., Google Big Table and Amazon Dynamo). Tuple stores are highly scalable and run on commodity clusters, but lack interfaces to support efficient development of mathematically based analytics. D4M (Dynamic Distributed Dimensional Data Model) has been developed to provide a mathematically rich interface to tuple stores (and structured query language “SQL” databases). D4M allows linear algebra to be readily applied to databases. Using D4M, it is possible to create composable analytics with significantly less effort than using traditional approaches. This work describes the D4M technology and its application and performance. Jeremy Kepner, William Arcand, Bill Bergeron, Nadya Bliss, Robert Bond, Chansup Byun, Gary Condon, Kenneth Gregson, Matthew Hubbell, Jonathan Kurz, Andrew McCabe, Peter Michaleas, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee |
ICASSP | 6 |
| 1997 | A multi-level parallelization concept for high-fidelity multi-block solversabstractThe integration of high-fidelity Computational Fluid Dynamics (CFD) analysis tools with the industrial design process benefits greatly from the robust implementations that are transportable across a wide range of computer architectures. In the present work, a hybrid domain-decomposition and parallelization concept was developed and implemented into the widely-used NASA multi-block Computational Fluid Dynamics (CFD) solvers employed in ENSAERO and OVERFLOW advanced flow analysis packages. These advanced engineering and scientific analysis packages include more than 300,000 lines of code written in FORTRAN 77 language in more than 1300 individual subprograms. The new parallel solver concept, PENS (Parallel Euler Navier-Stokes Solver), employs both fine and coarse granularity with data partitioning as well as data coalescing to obtain the desired load-balance characteristics on the available computer platforms for these legacy packages. This multi-level parallelism implementation itself introduces no changes to the numerical results, hence the original fidelity of the packages are identically preserved. The present implementation uses the Message Passing Interface (MPI) library for interprocessor message passing and memory accessing. By choosing an appropriate combination of the available partitioning and coalescing possibilities only during the execution stage, the PENS solver is used on different computer architectures from shared-memory to distributed-memory platforms with varying degrees of parallelism. Improvements in computational load-balance and speeds are extremely crucial on the realistic problems in the design of aerospace vehicles. The PENS implementation on the IBM SP2 distributed memory environment at the NASA Ames Research Center obtains 85 percent scalable parallel performance using fine-grain partitioning of single-block CFD domains using up to 128 wide computational nodes. Multi-block CFD simulations of complete aircraft geometries achieve 85 percent perfect load-balanced executions using data coalescing and the two levels of parallelism. SGI PowerChallenge, SGI Onyx2, and Cray T3E are the other platforms where the robustness, performance behavior, and the parallel scalability of the implementation are tested and fine-tuned for actual production run environments. Ferhat F. Hatay, Dennis C. Jespersen, Guru P. Guruswamy, Yehia M. Rizk, Chansup Byun, Ken Gee |
SC | 5 |