EDBT 2026 Demo / reviewers in the wild / expert
Pei-Hung Lin
dblp:62/7585
· DBLP profile ↗
13ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4977-814XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code TranslationabstractLe Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nuo Xu 0013, Winson Chen, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao |
ACL (1) | 5 |
| 2026 | Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors
Ruimin Shi, Maya B. Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Bo Peng |
Euro-Par (1) | 3 |
| 2026 | High-performance Vector-length Agnostic Quantum Circuit Simulations on ARM Processors
Ruimin Shi, Gabin Schieffer, Pei-Hung Lin, Maya B. Gokhale, Andreas Herten, Ivy Bo Peng |
IPDPS | 3 |
| 2025 | ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
Ruimin Shi, Gabin Schieffer, Maya B. Gokhale, Pei-Hung Lin, Hiren D. Patel, Ivy Bo Peng |
Euro-Par (2) | 4 |
| 2021 | Does it matter?: OMPSanitizer: an impact analyzer of reported data races in OpenMP programsabstractData races are a primary source of concurrency bugs in parallel programs. Yet, debugging data races is not easy, even with a large amount of data race detection tools. In particular, there still exists a manually-intensive and time-consuming investigation process after data races are reported by existing race detection tools. To address this issue, we present OMPSanitizer in this paper. OMPSanitizer employs a novel and semantic-aware impact analysis mechanism to assess the potential impact of detected data races so that developers can focus on data races with a high probability to produce a harmful impact. This way, OMPSanitizer can remove the heavy debugging burden of data races from developers and simultaneously enhance the debugging efficiency. We have implemented OMPSanitizer based on the widely-used dynamic binary instrumentation infrastructure, Intel Pin. Our evaluation results on a broad range of OpenMP programs from the DataRaceBench benchmark suite and an ECP Proxy application demonstrate that OMPSanitizer can precisely report the impact of data races detected by existing race detectors, e.g., Helgrind and ThreadSanitizer. We believe OMPSanitizer will provide a new perspective on automating the debugging support for data races in OpenMP programs. Wenwen Wang 0001, Pei-Hung Lin |
ICS | 2 |
| 2020 | XPlacer: Automatic Analysis of Data Access Patterns on Heterogeneous CPU/GPU SystemsabstractThis paper presents XPlacer, a framework to automatically analyze problematic data access patterns in C++ and CUDA code. XPlacer records heap memory operations in both host and device code for later analysis. To this end, XPlacer instruments read and write operations, function calls, and kernel launches. Programmers mark points in the program execution where the recorded data is analyzed and anomalies diagnosed. XPlacer reports data access anti-patterns, including alternating CPU/GPU accesses to the same memory, memory with low access density, and unnecessary data transfers. The diagnostic also produces summative information about the recorded accesses, which aids users in identifying code that could degrade performance. The paper evaluates XPlacer using LULESH, a Lawrence Livermore proxy application, Rodina benchmarks, and an implementation of the Smith-Waterman algorithm. XPlacer diagnosed several performance issues in these codes. The elimination of a performance problem in LULESH resulted in a 3x speedup on a heterogeneous platform combining Intel CPUs and Nvidia GPUs. Peter Pirkelbauer, Pei-Hung Lin, Tristan Vanderbruggen, Chunhua Liao |
IPDPS | 2 |
| 2019 | Preparation and optimization of a diverse workload for a large-scale heterogeneous systemabstractProductivity from day one on supercomputers that leverage new technologies requires significant preparation. An institution that procures a novel system architecture often lacks sufficient institutional knowledge and skills to prepare for it. Thus, the "Center of Excellence" (CoE) concept has emerged to prepare for systems such as Summit and Sierra, currently the top two systems in the Top 500. This paper documents CoE experiences that prepared a workload of diverse applications and math libraries for a heterogeneous system. We describe our approach to this preparation, including our management and execution strategies, and detail our experiences with and reasons for using different programming approaches. Our early science and performance results show that the project enabled significant early seismic science with up to a l4X throughput increase over Cori. In addition to our successes, we discuss our challenges and failures so others may benefit from our experience. Ian Karlin, Yoonho Park, Bronis R. de Supinski, Bert Still, D. A. Beckingsale, Robert Blake, Tong Chen 0001, Guojing Cong, Carlos H. A. Costa, Johann Dahm, Giacomo Domeniconi, Thomas Epperly, Aaron Fisher, Sara Kokkila Schumacher, Steve H. Langer, Hai Le, Naoya Maruyama, Xinyu Que, David F. Richards, Björn Sjögreen, Jonathan Wong, Carol S. Woodward, Ulrike Meier Yang, Bob Anderson, David Appelhans, Levi Barnes, Peter D. Barnes Jr., Sorin Bastea, David Böhme, Jamie A. Bramwell, James M. Brase, José R. Brunheroto, Barry Chen, Charway R. Cooper, Tony Degroot, Robert D. Falgout, Todd Gamblin, David J. Gardner, James N. Glosli, John A. Gunnels, Max P. Katz, Tzanio V. Kolev, I-Feng W. Kuo, Matthew P. LeGendre, Pei-Hung Lin, Shelby Lockhart, Kathleen McCandless, Claudia Misale, Jaime H. Moreno, Rob Neely, Jarom Nelson, Rao Nimmakayala, Kathryn M. O'Brien, Kevin O'Brien, Ramesh Pankajakshan, Roger A. Pearce, Slaven Peles, Phil Regier, Steven C. Rennich, Martin Schulz 0001, Howard Scott, James C. Sexton, Kathleen Shoga, Shiv Sundram, Guillaume Thomas-Collignon, Brian Van Essen, Alexey Voronin, Bob Walkup, Chris Ward, Hui-Fang Wen, Daniel A. White, Christopher Young, Cyril Zeller, Edward Zywicz |
SC | 49 |
| 2018 | Runtime and Memory Evaluation of Data Race Detection Tools
Pei-Hung Lin, Chunhua Liao, Markus Schordan, Ian Karlin |
ISoLA (2) | 1 |
| 2017 | DataRaceBench: a benchmark suite for systematic evaluation of data race detection toolsabstractData races in multi-threaded parallel applications are notoriously damaging while extremely difficult to detect. Many tools have been developed to help programmers find data races. However, there is no dedicated OpenMP benchmark suite to systematically evaluate data race detection tools for their strengths and limitations. Chunhua Liao, Pei-Hung Lin, Joshua Asplund, Markus Schordan, Ian Karlin |
SC | 2 |
| 2016 | Transforming the multifluid PPM algorithm to run on GPUs
Pei-Hung Lin, Paul R. Woodward |
J. Parallel Distributed Comput. | 1 |
| 2014 | Verification of Polyhedral Optimizations with Constant Loop Bounds in Finite State Space Computations
Markus Schordan, Pei-Hung Lin, Daniel J. Quinlan, Louis-Noël Pouchet |
ISoLA (2) | 2 |
| 2014 | Revisiting loop fusion in the polyhedral frameworkabstractLoop fusion is an important compiler optimization for improving memory hierarchy performance through enabling data reuse. Traditional compilers have approached loop fusion in a manner decoupled from other high-level loop optimizations, missing several interesting solutions. Recently, the polyhedral compiler framework with its ability to compose complex transformations, has proved to be promising in performing loop optimizations for small programs. However, our experiments with large programs using state-of-the-art polyhedral compiler frameworks reveal suboptimal fusion partitions in the transformed code. We trace the reason for this to be lack of an effective cost model to choose a good fusion partitioning among the possible choices, which increase exponentially with the number of program statements. In this paper, we propose a fusion algorithm to choose good fusion partitions with two objective functions - achieving good data reuse and preserving parallelism inherent in the source code. These objectives, although targeted by previous work in traditional compilers, pose new challenges within the polyhedral compiler framework and have thus not been addressed. In our algorithm, we propose several heuristics that work effectively within the polyhedral compiler framework and allow us to achieve the proposed objectives. Experimental results show that our fusion algorithm achieves performance comparable to the existing polyhedral compilers for small kernel programs, and significantly outperforms them for large benchmark programs such as those in the SPEC benchmark suite. Sanyam Mehta, Pei-Hung Lin, Pen-Chung Yew |
PPoPP | 2 |
| 2009 | First experience of compressible gas dynamics simulation on the Los Alamos roadrunner machineabstractAbstract We report initial experience with gas dynamics simulation on the Los Alamos Roadrunner machine. In this initial work, we have restricted our attention to flows in which the flow Mach number is less than 2. This permits us to use a simplified version of the PPM gas dynamics algorithm that has been described in detail by Woodward (2006). We follow a multifluid volume fraction using the PPB moment‐conserving advection scheme, enforcing both pressure and temperature equilibrium between two monatomic ideal gases within each grid cell. The resulting gas dynamics code has been extensively restructured for efficient multicore processing and implemented for scalable parallel execution on the Roadrunner system. The code restructuring and parallel implementation are described and performance results are discussed. For a modest grid size, sustained performance of 3.89 Gflops−1 CPU‐core−1 is delivered by this code on 36 Cell processors in 9 triblade nodes of a single rack of Roadrunner hardware. Copyright © 2009 John Wiley & Sons, Ltd. Paul R. Woodward, Jagan Jayaraj, Pei-Hung Lin, William Dai |
Concurr. Comput. Pract. Exp. | 3 |