James A. Ross

dblp:56/9913 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 72% GPUs and heterogeneous computing · 12% Hardware accelerators and domain-specific architectures · 12%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms › neuromorphic computing
brain-inspired computing
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Emerging computing paradigms
neuromorphic computing
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
Emerging computing paradigms
neuromorphic hardware
0.212016
Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications · SC 2016
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing
0.112011
Hybrid Core Acceleration of UWB SIRE Radar Signal Processing · IEEE Trans. Parallel Distributed Syst. 2011
Hardware accelerators and domain-specific architectures › signal processing accelerator
radar signal processing accelerator
0.112011
Hybrid Core Acceleration of UWB SIRE Radar Signal Processing · IEEE Trans. Parallel Distributed Syst. 2011
Embedded and real-time systems
real-time signal processing
0.012011
Hybrid Core Acceleration of UWB SIRE Radar Signal Processing · IEEE Trans. Parallel Distributed Syst. 2011

Methods — techniques the papers use, named apart from their topics

software ecosystem · 0.2scalable systems · 0.2brook+ · 0.1CUDA · 0.1
YearPublicationVenuePosition
2016 Truenorth ecosystem for brain-inspired computing: scalable systems, software, and applications
abstract
Abstract not provided
Jun Sawada, Filipp Akopyan, Andrew S. Cassidy, Brian Taba, Michael DeBole, Pallab Datta, Rodrigo Alvarez-Icaza, Arnon Amir, John V. Arthur, Alexander Andreopoulos, Rathinakumar Appuswamy, Heinz Baier, Davis Barch, David J. Berg, Carmelo di Nolfo, Steven K. Esser, Myron Flickner, Thomas A. Horvath, Bryan L. Jackson, Jeffrey A. Kusnitz, Scott Lekuch, Michael Mastro, Timothy Melano, Paul Merolla, Steven E. Millman, Tapan K. Nayak, Norm Pass, Hartmut Penner, William P. Risk, Kai Schleupen, Ben Shaw 0001, Hayley Wu, Brian Giera, Adam Moody, T. Nathan Mundhenk, Brian Van Essen, Eric X. Wang, David P. Widemann, William E. Murphy, Jamie K. Infantolino, James A. Ross, Dale R. Shires, Manuel M. Vindiola, Raju Namburu, Dharmendra S. Modha
SC42
2014 Cycle-accurate 8080 emulation using an ARM11 processor with dynamic binary translation
abstract
We describe the investigation of methods for cycle-accurate emulation of an Intel 8080 CPU using a low-power ARM11 CPU. Source-level optimizations and several threaded dispatch designs are explored, as well as an extreme optimization technique using direct register mapping and dynamic binary translation between the two Instruction Set Architectures. The cycle efficiency of the fastest design is nearly as fast as the original 8080 processor.
David A. Richie, James A. Ross
MEMOCODE2
2014 A Class-Structured Approach to Couple Application and Hybrid Core Parallelism
abstract
This paper presents an application that performs multi-objective geospatial optimizations for tactical mission planning in an urban environment. Utilizing a XML-driven C++ framework developed for hybrid platforms, the application distributes computational tasks to a collection of heterogeneous OpenCL devices and achieves efficient task parallel scaling. The framework abstracts the details of the mission scenario setup, compute device architecture details, and task scheduling to enable robust extensibility in software. Performance scalability and multi-objective parameter visualization for a prototypical mission scenario are documented in this study.
James A. Ross, David A. Richie, Song Jun Park, Dale R. Shires, Brian J. Henz
PDP1
2011 Hybrid Core Acceleration of UWB SIRE Radar Signal Processing
abstract
To move High-Performance Computing (HPC) closer to forward operating environments and missions, the Army Research Laboratory is developing approaches using hybrid, asymmetric core computing. By blending capabilities found in Graphics Processing Units (GPUs) and traditional von Neumann multicore Central Processing Units (CPUs), approaches are being developed and optimized to provide at or near real-time processing speeds for research project applications. Algorithms are designed to partition work to resources best designed to handle the processing load. The use of commodity resources allows the design to be flexible throughout the life cycle without the costly and time-consuming delays associated with Application-Specific Integrated Circuit (ASIC) development. This paradigm allows for rapid technology transfer to end users. In this paper, we describe a synchronous impulse reconstruction radar imaging algorithm that has been designed for hybrid CPU-GPU processing. We discuss various optimizations such as asynchronous task partitioning between the CPU and GPU as well as data movement reduction. We also discuss analysis and design of the algorithms within the context of two programming models: NVIDIA's CUDA and AMD's ATI Brook+. Finally, we report on the speedup achieved by this approach that allowed us to take a code once restricted to postprocessing and transform it into one that exceeds real-time performance requirements.
Song Jun Park, James A. Ross, Dale R. Shires, David A. Richie, Brian J. Henz, Lam H. Nguyen
IEEE Trans. Parallel Distributed Syst.2