Conor John Williams

dblp:371/5041 · DBLP profile ↗
← Back
1ranked-venue papers
1as first author
1since 2021 · last 2025
0000-0001-7830-7473ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models › task parallelism
fork-join parallelism
0.912025
Libfork: Portable Continuation-Stealing With Stackless Coroutines · IEEE Trans. Parallel Distributed Syst. 2025
Parallel and multicore computing
parallel programming models
0.912025
Libfork: Portable Continuation-Stealing With Stackless Coroutines · IEEE Trans. Parallel Distributed Syst. 2025
Parallel and multicore computing
parallel programming runtimes
0.912025
Libfork: Portable Continuation-Stealing With Stackless Coroutines · IEEE Trans. Parallel Distributed Syst. 2025
Parallel and multicore computing › task scheduling › dynamic scheduling
work-stealing scheduler
0.912025
Libfork: Portable Continuation-Stealing With Stackless Coroutines · IEEE Trans. Parallel Distributed Syst. 2025

Methods — techniques the papers use, named apart from their topics

stackless coroutines · 0.9segmented stacks · 0.9NUMA optimization · 0.9
YearPublicationVenuePosition
2025 Libfork: Portable Continuation-Stealing With Stackless Coroutines
abstract
Fully-strict fork-join parallelism is a powerful model for shared-memory programming due to its optimal time-scaling and strong bounds on memory scaling. The latter is rarely achieved due to the difficulty of implementing continuation-stealing in traditional High Performance Computing (HPC) languages – where it is often impossible without modifying the compiler or resorting to non-portable techniques. We demonstrate how stackless-coroutines (a new feature in C++$\bm {20}$) can enable fully-portable continuation stealing and presentlibforka wait-free fine-grained parallelism library, combining coroutines with user-space, geometric segmented-stacks. We show our approach is able to achieve optimal time/memory scaling, both theoretically and empirically, across a variety of benchmarks. Compared to openMP (libomp), libfork is on average$7.2\times$faster and consumes$10\times$less memory. Similarly, compared to Intel's TBB, libfork is on average$2.7\times$faster and consumes$6.2\times$less memory. Additionally, we introduce non-uniform memory access (NUMA) optimizations for schedulers that demonstrate performance matchingbusy-waitingschedulers.
Conor John Williams, James Elliott
IEEE Trans. Parallel Distributed Syst.1