Utpal Bora 0001

dblp:167/1816-1 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-0076-1059ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Processor architecture and microarchitecture · 46% Parallel and multicore computing · 31% Memory systems · 20%
Software engineering, system software, and programming languages
1 paper
Concurrent programming · 50% Program analysis · 50%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › multithreading
helper thread prefetching
0.912025
Ghost Threading: Helper-Thread Prefetching for Real Systems · MICRO 2025
Processor architecture and microarchitecture › out-of-order execution
instruction window
0.912025
LoopFrog: In-Core Hint-Based Loop Parallelization · MICRO 2025
Parallel and multicore computing › loop transformation
loop parallelization
0.912025
LoopFrog: In-Core Hint-Based Loop Parallelization · MICRO 2025
Processor architecture and microarchitecture
out-of-order execution
0.912025
LoopFrog: In-Core Hint-Based Loop Parallelization · MICRO 2025
Memory systems › cache
prefetching
0.912025
Ghost Threading: Helper-Thread Prefetching for Real Systems · MICRO 2025
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.912025
LoopFrog: In-Core Hint-Based Loop Parallelization · MICRO 2025
Concurrent programming › concurrency bugs
data races
0.412020
LLOV: A Fast Static Data-Race Checker for OpenMP Programs · ACM Trans. Archit. Code Optim. 2020
Concurrent programming › concurrency bug detection
data race detection
0.412020
LLOV: A Fast Static Data-Race Checker for OpenMP Programs · ACM Trans. Archit. Code Optim. 2020
Program analysis
static analysis
0.412020
LLOV: A Fast Static Data-Race Checker for OpenMP Programs · ACM Trans. Archit. Code Optim. 2020
Program analysis › concurrent program analysis
static race detection
0.412020
LLOV: A Fast Static Data-Race Checker for OpenMP Programs · ACM Trans. Archit. Code Optim. 2020
Memory systems
cache
0.312025
Ghost Threading: Helper-Thread Prefetching for Real Systems · MICRO 2025

Methods — techniques the papers use, named apart from their topics

helper threading · 0.9compiler hints · 0.9LLVM · 0.9LLVM compiler framework · 0.9
YearPublicationVenuePosition
2025 LoopFrog: In-Core Hint-Based Loop Parallelization
abstract
To scale ILP, designers build deeper and wider out-of-order superscalar CPUs.However, this approach incurs quadratic scaling complexity, area, and energy costs with each generation.While small loops may benefit from increased instruction-window sizes and large loops may see speedups via thread-level parallelism across cores, there remains unexploited medium-granularity parallelism.We propose LoopFrog to tap into this potential by bringing thread-level speculation schemes into the modern era.LoopFrog runs multiple loop iterations from a single thread in parallel within the microarchitecture.The core can spawn future loop iterations as new microarchitectural threadlets based on compiler-inserted hints, which can leapfrog execution beyond the parent thread's instruction window, exposing a new, medium-grained parallelism, orthogonal to traditional ILP and TLP.LoopFrog monitors data dependencies between executing threadlets, forwards data for true dependencies and squashes speculative threadlets on ordering violations.Using an LLVM-based compiler to insert hints, we achieve a geometric mean loop speedup of 43%, translating to whole-program speedups of 9.2% on SPEC CPU 2006 and 9.5% on SPEC CPU 2017 benchmarks, with only modest area and power overheads.
Márton Erdos, Utpal Bora 0001, Akshay Bhosale, Bob Lytton, Ali Mustafa Zaidi, Alexandra W. Chadwick, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO2
2025 Ghost Threading: Helper-Thread Prefetching for Real Systems
Akshay Bhosale, Utpal Bora 0001, Alexandra W. Chadwick, Márton Erdos, Giacomo Gabrielli, Timothy M. Jones 0001
MICRO3
2025 LLOR: Automated Repair of OpenMP Programs
Utpal Bora 0001, Saurabh Joshi 0001, Gautam Muduganti, Ramakrishna Upadrasta
VMCAI (2)1
2020 LLOV: A Fast Static Data-Race Checker for OpenMP Programs
abstract
In the era of Exascale computing, writing efficient parallel programs is indispensable, and, at the same time, writing sound parallel programs is very difficult. Specifying parallelism with frameworks such as OpenMP is relatively easy, but data races in these programs are an important source of bugs. In this article, we propose LLOV, a fast, lightweight, language agnostic, and static data race checker for OpenMP programs based on the LLVM compiler framework. We compare LLOV with other state-of-the-art data race checkers on a variety of well-established benchmarks. We show that the precision, accuracy, and the F1 score of LLOV is comparable to other checkers while being orders of magnitude faster. To the best of our knowledge, LLOV is the only tool among the state-of-the-art data race checkers that can verify a C/C++ or FORTRAN program to be data race free.
Utpal Bora 0001, Pankaj Kukreja, Saurabh Joshi 0001, Ramakrishna Upadrasta, Sanjay V. Rajopadhye
ACM Trans. Archit. Code Optim.1