Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bozhi You

dblp:235/2493 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2025
0000-0003-4691-5874ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 61% GPUs and heterogeneous computing · 30% High-performance computing · 9%
Theoretical computer science
1 paper
Distributed computing theory · 61% Graph algorithms and graph theory · 39%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel scheduling
resource-aware scheduling
0.612022
Parla: A Python Orchestration System for Heterogeneous Architectures · SC 2022
Parallel and multicore computing
task scheduling
0.612022
Parla: A Python Orchestration System for Heterogeneous Architectures · SC 2022
Graph algorithms and graph theory › centrality
betweenness centrality
0.412019
A round-efficient distributed betweenness centrality algorithm · PPoPP 2019
Distributed computing theory › distributed graph algorithms
CONGEST model
0.412019
A round-efficient distributed betweenness centrality algorithm · PPoPP 2019
Distributed computing theory
distributed graph algorithms
0.412019
A round-efficient distributed betweenness centrality algorithm · PPoPP 2019
High-performance computing › scientific computing
scientific computing application
0.212022
Parla: A Python Orchestration System for Heterogeneous Architectures · SC 2022
Graph algorithms and graph theory › shortest path
all-pairs shortest paths
0.112019
A round-efficient distributed betweenness centrality algorithm · PPoPP 2019

Methods — techniques the papers use, named apart from their topics

task-based runtime · 0.6GPU context management · 0.6distributed-memory algorithm · 0.4
YearPublicationVenuePosition
2025 VLCs: Managing Parallelism with Virtualized Libraries
abstract
As the complexity and scale of modern parallel machines continue to grow, programmers increasingly rely on composition of software libraries to encapsulate and exploit parallelism. However, many libraries are not designed with composition in mind and assume they have exclusive access to all resources. Using such libraries concurrently can result in contention and degraded performance. Prior solutions involve modifying the libraries or the OS, which is often infeasible.
Yineng Yan, William Ruys, Ian Henriksen, Arthur Michener Peters, Sean Stephens, Bozhi You, Henrique Fingler, Martin Burtscher, Milos Gligoric 0001, Keshav Pingali, Mattan Erez, George Biros, Christopher J. Rossbach
SoCC7
2022 Parla: A Python Orchestration System for Heterogeneous Architectures
abstract
Python's ease of use and rich collection of numeric libraries make it an excellent choice for rapidly developing scientific applications. However, composing these libraries to take advantage of complex heterogeneous nodes is still difficult. To simplify writing multi-device code, we created Parla, a heterogeneous task-based programming framework that fully supports Python's scientific programming stack. Parla's API is based on Python decorators and allows users to wrap code in Parla tasks for parallel execution. Parla arrays enable automatic movement of data between devices. The Parla runtime handles resource-aware mapping, scheduling, and execution of tasks. Compared to other Python tasking systems, Parla is unique in its parallelization of tasks within a single process, its GPU context and resource-aware runtime, and its design around gradual adoption to provide easy migration of and integration into existing Python applications. We show that Parla can achieve performance competitive with hand-optimized code while improving ease of development.
William Ruys, Ian Henriksen, Arthur Michener Peters, Yineng Yan, Sean Stephens, Bozhi You, Henrique Fingler, Martin Burtscher, Milos Gligoric 0001, Karl W. Schulz, Keshav Pingali, Christopher J. Rossbach, Mattan Erez, George Biros
SC7
2019 A round-efficient distributed betweenness centrality algorithm
abstract
We present Min-Rounds BC (MRBC), a distributed-memory algorithm in the CONGEST model that computes the betweenness centrality (BC) of every vertex in a directed unweighted n-node graph in O(n) rounds. Min-Rounds BC also computes all-pairs-shortest-paths (APSP) in such graphs. It improves the number of rounds by at least a constant factor over previous results for unweighted directed APSP and for unweighted BC, both directed and undirected.
Loc Hoang, Matteo Pontecorvi, Roshan Dathathri, Gurbinder Gill, Bozhi You, Keshav Pingali, Vijaya Ramachandran
PPoPP5