Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

James Psota

dblp:45/4312 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
1since 2021 · last 2024
0000-0003-1371-7177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 72% High-performance computing · 23% Processor architecture and microarchitecture · 3%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel programming models
message passing
0.812024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
Parallel and multicore computing
parallel programming models
0.812024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
Parallel and multicore computing
parallel programming runtimes
0.812024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
High-performance computing › performance optimization
shared memory optimization
0.812024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
Parallel and multicore computing › load balancing › dynamic load balancing
work stealing
0.812024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
High-performance computing
scientific computing systems
0.212024
Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes · PPoPP 2024
Processor architecture and microarchitecture
instruction-level parallelism
0.012004
Evaluation of the Raw Microprocessor: An Exposed-Wire-Delay Architecture for ILP and Streams · ISCA 2004
Interconnection networks and networks-on-chip › interprocessor communication
operand transport network
0.012004
Evaluation of the Raw Microprocessor: An Exposed-Wire-Delay Architecture for ILP and Streams · ISCA 2004
Processor architecture and microarchitecture
tiled architecture
0.012004
Evaluation of the Raw Microprocessor: An Exposed-Wire-Delay Architecture for ILP and Streams · ISCA 2004
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation
0.012004
Evaluation of the Raw Microprocessor: An Exposed-Wire-Delay Architecture for ILP and Streams · ISCA 2004

Methods — techniques the papers use, named apart from their topics

work stealing · 0.8lock-free data structures · 0.8cycle-accurate simulation · 0.0
YearPublicationVenuePosition
2024 Pure: Evolving Message Passing To Better Leverage Shared Memory Within Nodes
abstract
Pure is a new programming model and runtime system explicitly designed to take advantage of shared memory within nodes in the context of a mostly message passing interface enhanced with the ability to use tasks to make use of idle cores. Pure leverages shared memory in two ways: (a) by allowing cores to steal work from each other while waiting on messages to arrive, and, (b) by leveraging efficient lock-free data structures in shared memory to achieve high-performance messaging and collective operations between the ranks within nodes. We use microbenchmarks to evaluate Pure's key messaging and collective features and also show application speedups up to 2.1× on the CoMD molecular dynamics and the miniAMR adaptive mesh refinement applications scaling up to 4,096 cores.
James Psota, Armando Solar-Lezama
PPoPP1
2010 ATAC: a 1000-core cache-coherent processor with on-chip optical network
abstract
Based on current trends, multicore processors will have 1000 cores or more within the next decade. However, their promise of increased performance will only be realized if their inherent scaling and programming challenges are overcome. Fortunately, recent advances in nanophotonic device manufacturing are making CMOS-integrated optics a reality-interconnect technology which can provide significantly more bandwidth at lower power than conventional electrical signaling. Optical interconnect has the potential to enable massive scaling and preserve familiar programming models in future multicore chips.
George Kurian, Jason E. Miller, James Psota, Jonathan Eastep, Jifeng Liu, Jürgen Michel, Lionel C. Kimerling, Anant Agarwal
PACT3
2010 ATAC: Improving performance and programmability with on-chip optical networks
abstract
Given the current trends in multicore scaling, chips with 1000 cores may exist within the next 5 to 10 years. However, their promise of increased performance will only be reached if their inherent scaling and programming challenges are overcome. Meanwhile, recent advances in nanophotonic device manufacturing are making CMOS-integrated optics a reality-interconnect technology which can provide more bandwidth at lower power than conventional electronics. Perhaps more importantly, optical interconnect also has the potential to enable new, easy-to-use programming models enabled by its inexpensive broadcast mechanism. This paper introduces ATAC, a new manycore architecture that capitalizes on the recent advances in optics to address a number of challenges that future manycore designs will face. The new constraints and opportunities of on-chip optical interconnect are presented and explored in the design of ATAC. Furthermore, this paper discusses ATAC's programming models, and introduces Consumer Tagging, a novel programming model that leverages ATAC's strengths to provide high performance and scalability.
James Psota, Jason E. Miller, George Kurian, Henry Hoffmann, Nathan Beckmann, Jonathan Eastep, Anant Agarwal
ISCAS1
2008 rMPI: Message Passing on Multicore Processors with On-Chip Interconnect
James Psota, Anant Agarwal
HiPEAC1
2004 Evaluation of the Raw Microprocessor: An Exposed-Wire-Delay Architecture for ILP and Streams
abstract
This paper evaluates the Raw microprocessor. Raw addresses the challenge of building a general-purpose architecture that performs well on a larger class of stream and embedded computing applications than existing microprocessors, while still running existing ILP-based sequential programs with reasonable performance in the face of increasing wire delays. Raw approaches this challenge by implementing plenty of on-chip resources - including logic, wires, and pins - in a tiled arrangement, and exposing them through a new ISA, so that the software can take advantage of these resources for parallel applications. Raw supports both ILP and streams by routing operands between architecturally-exposed functional units over a point-to-point scalar operand network. This network offers low latency for scalar data transport. Raw manages the effect of wire delays by exposing the interconnect and using software to orchestrate both scalar and stream data transport. We have implemented a prototype Raw microprocessor in IBM's 180 nm, 6-layer copper, CMOS 7SF standard-cell ASIC process. We have also implemented ILP and stream compilers. Our evaluation attempts to determine the extent to which Raw succeeds in meeting its goal of serving as a more versatile, general-purpose processor. Central to achieving this goal is Raw's ability to exploit all forms of parallelism, including ILP, DLP, TLP, and Stream parallelism. Specifically, we evaluate the performance of Raw on a diverse set of codes including traditional sequential programs, streaming applications, server workloads and bit-level embedded computation. Our experimental methodology makes use of a cycle-accurate simulator validated against our real hardware. Compared to a 180nm Pentium-III, using commodity PC memory system components, Raw performs within a factor of 2/spl times/ for sequential applications with a very low degree of ILP, about 2/spl times/ to 9/spl times/ better for higher levels of ILP, and 10/spl times/-100/spl times/ better when highly parallel applications are coded in a stream language or optimized by hand. The paper also proposes a new versatility metric and uses it to discuss the generality of Raw.
Michael B. Taylor, Walter Lee, Jason E. Miller, David Wentzlaff, Ian Bratt, Ben Greenwald, Henry Hoffmann, Paul R. Johnson, Jason Sungtae Kim, James Psota, Arvind Saraf, Nathan Shnidman, Volker Strumpen, Matthew I. Frank, Saman P. Amarasinghe, Anant Agarwal
ISCA10