Blaise-Pascal Tine

dblp:207/5373 · also Blaise Tine · DBLP profile ↗
← Back
11ranked-venue papers
8as first author
7since 2021 · last 2026
0000-0002-2964-8631ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 8 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Inside VOLT: Designing an Open-Source GPU Compiler (Tool)
abstract
Recent efforts in open-source GPU research are opening new avenues in a domain that has long been tightly coupled with a few commercial vendors. Emerging open GPU architectures define SIMT functionality through their own ISAs, but executing existing GPU programs and optimizing performance on these ISAs relies on a compiler framework that is technically complex and often undercounted in open-hardware development costs.
Shinnung Jeong, Chihyo Ahn, Huanzhi Pu, Jisheng Zhao, Hyesoon Kim, Blaise-Pascal Tine
CC6
2025 SoftCUDA: Running CUDA on Softcore GPU
abstract
Field-Programmable Gate Arrays (FPGAs) have been extensively employed to accelerate parallel applications by allowing designers to customize their hardware for maximum performance. However, most FPGA-based designs are constrained to specific kernels, limiting their suitability across diverse workloads. As GPU workloads grow in complexity and their requirements diverge, softcore GPU(SoftGPU) designs have emerged to exploit FPGA reconfigurability for accelerating a broader range of parallel applications. Despite their potential, these designs have seen limited adoption due to the lack of comprehensive software stack support. In a CUDA-dominated development landscape, translating CUDA source code to alternative programming models can be challenging and often lacks direct feature parity. This paper introduces SoftCUDA, a novel framework that delivers comprehensive, end-to-end CUDA support on our SoftGPU, Vortex. By fully leveraging the reconfigurable architecture of SoftGPU and maintaining a user-friendly CUDA interface, SoftCUDA enables seamless integration and execution of unmodified CUDA applications on FPGA-based platforms.
Chihyo Ahn, Ruobing Han, Udit Subramanya, Jisheng Zhao, Blaise-Pascal Tine, Hyesoon Kim
FCCM5
2025 Analysis of the RISC-V Vector Extension for Vulkan Graphics Kernels
abstract
RISC-V vector extensions have been recently officially adopted, but compiler support is still under active development. Moreover, most vectorization efforts have been concentrated on HPC and machine learning workloads. In this paper, we analyze the benefits of vectorization on RISC-V GPUs and CPUs for Vulkan 3D graphics applications.
Martin Troiber, Martin Schulz 0001, Blaise-Pascal Tine, Hyesoon Kim
ISPASS3
2024 Towards "True" GPU Performance Scaling for OpenGPU
abstract
General-Purpose Graphics Processing Units (GPGPUs) have gained significant attention for their high computational throughput and energy efficiency. Their ability to offer high-performance, general-purpose programmability makes them ideal for accelerating diverse applications in fields such as artificial intelligence, scientific computing, healthcare, financial modeling, and computer graphics. The Vortex OpenGPU introduced the first open-source full-system GPGPU, extending the RISC-V base ISA to introduce a Single-Instruction-Multiple-Threads (SIMT) execution model. In this work, we extended the OpenGPU microarchitecture into a configurable superscalar pipeline, migrating it from a CPU-oriented microarchitecture - where the performance scaling was centered on increasing the number of GPU cores - to a GPU-oriented microarchitecture where the performance scaling is centered towards scaling the parallelism inside a single core. We evaluated the design on Intel FPGA, achieving a 29% performance increase with a 4 -wide single-core 32 threads GPU versus an areaequivalent 16-core processor.
Blaise-Pascal Tine, Hyesoon Kim
HCS1
2023 Skybox: Open-Source Graphic Rendering on Programmable RISC-V GPUs
abstract
Graphics rendering remains one of the most compute intensive and memory bound applications of GPUs and has been driving their push for performance and energy efficiency since its inception. Early GPU architectures focused only on accelerating graphics rendering and implemented dedicated fixed- function rasterizer hardware to speed-up their rendering pipeline. As GPUs have become more programmable and ubiquitous in other application domains such as scientific computing, machine learning, graph analytics, and crypto-currency, generalizing GPU microarchitectures for area and power efficiency becomes necessary, especially for mobile and IoT devices. In this work, we present Skybox, a full-stack open-source GPU architecture with integrated software, compiler, hardware, and simulation environment, that enables end-to-end GPU research. Using Skybox, we explore the design space of software versus hardware graphics rendering and propose and hybrid micro-architecture that accelerates the state-of-the art Vulkan graphics API. Skybox also introduces novel compiler and system optimizations to support its unique RISC-V ISA baseline. We evaluated Skybox on high- end Altera and also Xilinx FPGAs. We were able to generate and execute a 32 cores (512 threads) Skybox graphics processor on Altera Stratix 10 FPGA, delivering a peak fill rate of 3.7 GPixels at 230 MHz. Skybox is the first open-source full-stack GPU software and hardware implementation that supports the Vulkan API
Blaise-Pascal Tine, Varun Saxena, Santosh Srivatsan, Joshua R. Simpson, Fadi Alzammar, Liam Cooper, Hyesoon Kim
ASPLOS (3)1
2022 Accelerating Graphic Rendering on Programmable RISC-V GPUs
abstract
Graphics rendering remains one of the most compute-intensive and memory-bound applications of GPUs and has been driving their push for performance and energy efficiency since its inception. Early GPU architectures focused only on accelerating graphics rendering and implemented dedicated a fixed-function rendering units. Today’s GPUs have become more programmable to address the complexity and diversity of modern graphics workloads while still accelerating several components of the graphics pipeline in fixed-function hardware.Generalizing the GPU microarchitecture and implement some of its graphics hardware blocks in software can save area that can be used to expand the generic pipeline, especially in mobile systems-on-chips environments where power and area is scarce.In this work, we propose a RISC-V-based hybrid GPU architecture that accelerates the graphics pipeline without paying the cost of a full hardware graphics pipeline. We evaluated the design on an Altera Arria 10 FPGA running at 200 MHz.
Blaise-Pascal Tine, Varun Saxena, Santosh Srivatsan, Joshua R. Simpson, Fadi Alzammar, Liam Cooper, Sam Jijina, Swetha Rajagoplan, Tejaswini Anand Kumar, Jeffrey Young 0001, Hyesoon Kim
HCS1
2021 Vortex: Extending the RISC-V ISA for GPGPU and 3D-Graphics
abstract
The importance of open-source hardware and software has been increasing. However, despite GPUs being one of the more popular accelerators across various applications, there is very little open-source GPU infrastructure in the public domain. We argue that one of the reasons for the lack of open-source infrastructure for GPUs is rooted in the complexity of their ISA and software stacks. In this work, we first propose an ISA extension to RISC-V that supports GPGPUs and graphics. The main goal of the ISA extension proposal is to minimize the ISA changes so that the corresponding changes to the open-source ecosystem are also minimal, which makes for a sustainable development ecosystem. To demonstrate the feasibility of the minimally extended RISC-V ISA, we implemented the complete software and hardware stacks of Vortex on FPGA. Vortex is a PCIe-based soft GPU that supports OpenCL and OpenGL. Vortex can be used in a variety of applications, including machine learning, graph analytics, and graphics rendering. Vortex can scale up to 32 cores on an Altera Stratix 10 FPGA, delivering a peak performance of 25.6 GFlops at 200 Mhz.
Blaise-Pascal Tine, Krishna Praveen Yalamarthy, Fares Elsabbagh, Hyesoon Kim
MICRO1
2020 Tango: An Optimizing Compiler for Just-In-Time RTL Simulation
abstract
With Moore’s law coming to an end, the advent of hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and increase design efficiency. This trend is further accentuated by rapid-pace of innovations in Machine Learning and Graph Analytic, calling for a faster product development cycle for hardware accelerators and the importance of addressing the increasing cost of hardware verification. The productivity of software-hardware co-design relies upon better integration between the software and hardware design methodologies, but more importantly in the effectiveness of the design tools and hardware simulators at reducing the development time. In this work, we developed Tango, an Optimizing compiler for Just-in-Time RTL simulation. Tango implements unique hardware-centric compiler transformations to speed up runtime code generation in a software-hardware co-design environment where hardware simulation speed is critical. Tango achieves a 6x average speedup compared to the state-of-the-art simulators.
Blaise-Pascal Tine, Sudhakar Yalamanchili, Hyesoon Kim
DATE1
2020 Cash: A Single-Source Hardware-Software Codesign Framework for Rapid Prototyping
abstract
With Moore's Law coming to an end, hardware specialization and systems on chips are providing new opportunities for continuing performance scaling while reducing the energy cost of computation. However, the current hardware design methodologies require significant engineering efforts and domain expertise, making the design process unscalable. More importantly, hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and design efficiency. In this work, we introduce Cash, a single-source hardware-software co-design framework for rapid SoC prototyping and accelerators research. Cash leverages the unique efficiency and generative attributes of Modern C++ to provide a unified development environment, aiming at closing the architecture research methodology gap. The Cash framework introduces new co-design programming abstractions that enable seamless integration with existing software from architecture research simulators to high-level synthesis.
Blaise-Pascal Tine, Fares Elsabbagh, Seyong Lee, Jeffrey S. Vetter, Hyesoon Kim
FPGA1
2020 Productive Hardware Designs using Hybrid HLS-RTL Development
abstract
Current High-Level Synthesis frameworks provide a productive hardware development methodology where hardware accelerators are generated directly from high-level languages like C/C++ or OpenCL, allowing software developers to quickly accelerate their applications. However, the hardware generated by these frameworks is sub-optimal compared to often hand-optimized RTL modules. A hybrid development approach would leverage the productive software stack and hardware board support package that HLS provides but allow for fine-grained optimization using RTL components. In this work, we introduce a new software-hardware co-design framework that integrates OpenCL/OpenACC with RTL code enabling direct execution on FPGAs as well as full emulation with a high-speed simulator to reduce the development time.
Blaise-Pascal Tine, Seyong Lee, Jeffrey S. Vetter, Hyesoon Kim
FPGA1
2019 POSTER: Tango: An Optimizing Compiler for Just-In-Time RTL Simulation
abstract
The end of Moore's law with the advent of hardware specialization presents a unique challenge for a much tighter software and hardware co-design environment to exploit domain-specific optimizations and increase design efficiency. The productivity of software-hardware codesign relies not on only in better integration between the software and hardware design methodologies but more importantly in the effectiveness of the design tools at reducing the development time. In this work, we developed Tango, an Optimizing compiler for a Just-in-Time RTL simulator. Tango implements unique hardware-centric compiler transformations to speed up runtime code generation in a software-hardware codesign environment where hardware simulation speed is critical. Tango achieves a 3x average speedup compared to the state-of-the-art RTL simulators.
Blaise-Pascal Tine, Sudhakar Yalamanchili, Hyesoon Kim, Jeffrey S. Vetter
PACT1