Michael A. Parker

dblp:34/9601 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
0since 2021 · last 2012
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 64% Embedded and real-time systems · 24% Parallel and multicore computing · 10%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 77% Geometric modeling and processing · 23%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory management
address remapping
0.122006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Impulse: Building a Smarter Memory Controller · HPCA 1999
Memory systems › shared memory
distributed shared memory
0.112006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Memory systems
memory controller
0.122001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Impulse: Building a Smarter Memory Controller · HPCA 1999
Embedded and real-time systems › real-time scheduling
hierarchical scheduling
0.012003
Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003
Embedded and real-time systems
real-time scheduling
0.012003
Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003
Embedded and real-time systems › real-time scheduling › schedulability analysis
response time analysis
0.012003
Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003
Memory systems › memory access
indirect addressing
0.012001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Visualization and visual analytics
volume visualization
0.011999
Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999
Parallel and multicore computing › parallel computing
parallel rendering
0.011999
Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999
Memory systems › cache
prefetching
0.011999
Impulse: Building a Smarter Memory Controller · HPCA 1999
Memory systems
cache coherence
0.012006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Parallel and multicore computing
multiprocessor system
0.012006
Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006
Memory systems
memory-bound computation
0.022001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Impulse: Building a Smarter Memory Controller · HPCA 1999
Parallel and multicore computing › parallel computing › parallel program analysis › concurrency bug detection
race detection
0.012003
Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003
Memory systems
DRAM
0.012001
The Impulse Memory Controller · IEEE Trans. Computers 2001
Geometric modeling and processing
isosurface extraction
0.011999
Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999
High-performance computing
scientific computing
0.011999
Impulse: Building a Smarter Memory Controller · HPCA 1999

Methods — techniques the papers use, named apart from their topics

simulation · 0.1volume bricking · 0.0shallow data hierarchy · 0.0
YearPublicationVenuePosition
2012 Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang
J. Supercomput.6
2007 Active memory operations
abstract
The performance of modern microprocessors is increasingly limited by their inability to hide main memory latency. The problem is worse in large-scale shared memory systems, where remote memory latencies are hundreds, and soon thousands, of processor cycles. To mitigate this problem, we propose the use of Active Memory Operations (AMOs), in which select operations can be sent to and executed on the home memory controller of data. AMOs can eliminate significant number of coherence messages, minimize intranode and internode memory traffic, and create opportunities for parallelism. Our implementation of AMOs is cache-coherent and requires no changes to the processor core or DRAM chips.
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Ali Ibrahim, Michael A. Parker
ICS5
2006 The design and utility of the ML-RSIM system simulator
Lambert Schaelicke, Michael A. Parker
J. Syst. Archit.2
2006 Efficient address remapping in distributed shared-memory systems
abstract
As processor performance continues to improve at a rate much higher than DRAM and network performance, we are approaching a time when large-scale distributed shared memory systems will have remote memory latencies measured in tens of thousands of processor cycles. The Impulse memory system architecture adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to control how data is accessed and cached and thereby improve cache and bus utilization and reduce the number of memory accesses required. Previous Impulse work focuses on uniprocessor systems and relies on software to flush processor caches when necessary to ensure data coherence. In this paper, we investigate an extension of Impulse to multiprocessor systems that extends the coherence protocol to maintain data coherence without requiring software-directed cache flushing. Specifically, the multiprocessor Impulse controller can gather/scatter data across the network while its coherence protocol guarantees that each gather request gets coherent data and each scatter request updates every coherent replica in the system. Our simulation results demonstrate that the proposed system can significantly outperform conventional systems, achieving an average speedup of 9X on four memory-bound benchmarks on a 32-processor system.
Lixin Zhang 0002, Michael A. Parker, John B. Carter
ACM Trans. Archit. Code Optim.2
2005 Fast synchronization on shared-memory multiprocessors: An architectural approach
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Liqun Cheng, Michael A. Parker
J. Parallel Distributed Comput.5
2004 A low power architecture for embedded perception
abstract
Recognizing speech, gestures, and visual features are important interface capabilities for future embedded mobile systems. Unfortunately, the real-time performance requirements of complex perception applications cannot be met by current embedded processors and often even exceed the performance of high performance microprocessors whose energy consumption far exceeds embedded energy budgets. Though custom ASICs provide a solution to this problem, they incur expensive and lengthy design cycles and are inflexible. This paper introduces a VLIW perception processor which uses a combination of clustered function units, compiler controlled dataflow and compiler controlled clock-gating in conjunction with a scratch-pad memory system to achieve high performance for perceptual algorithms at low energy consumption. The architecture is evaluated using ten benchmark applications taken from complex speech and visual feature recognition, security, and signal processing domains. The energy-delay product of a 0.13μ implementation of this architecture is compared against ASICs and general purpose processors. Using a combination of Spice simulations and real processor power measurements, we show that the cluster running at 1 GHz clock frequency outperforms a 2.4 GHz Pentium 4 by a factor of 1.75 while simultaneously achieving 159 times better energy delay product than a low power Intel XScale embedded processor.
Binu K. Mathew, Al Davis, Michael A. Parker
CASES3
2004 Energy efficient cluster co-processors [3G wireless applications]
abstract
New 3G wireless algorithms require more performance than can be currently provided by embedded processors. ASICs provide the necessary performance but are costly to design and sacrifice generality. This paper introduces a clustered VLIW coprocessor approach that organizes the execution and storage resources differently than a traditional general-purpose processor or DSP. The execution units of the coprocessor are clustered and embedded in a rich set of communication resources. Fine grain control of these resources is imposed by a wide-word horizontal micro-code program. The advantages of this approach are quantified on a suite of six algorithms that are taken from both traditional DSP applications and from the new 3G cellular telephony domain. The result is surprising. The execution clusters retain much of the generality of a conventional processor while simultaneously improving performance by one to two orders of magnitude and by reducing energy-delay by three to four orders of magnitude when compared to a conventional embedded processor such as the Intel XScale.
Ali Ibrahim, Michael A. Parker, Al Davis
ICASSP (5)2
2003 Evolving real-time systems using hierarchical scheduling and concurrency analysis
abstract
We have developed a new way to look at real-time and embedded software: as a collection of execution environments created by a hierarchy of schedulers. Common schedulers include those than run interrupts, bottom-half handlers, threads, and events. We have created algorithms for deriving response times, scheduling overheads, and blocking terms for tasks in systems containing multiple execution environments. We have also created task scheduler logic, a formalism that permits checking systems for race conditions and other errors. Concurrency analysis of low-level software is challenging because there are typically several kinds of locks, such as thread mutexes and disabling interrupts, and groups of cooperating tasks may need to acquire some, all or none of the available types of locks to create correct software. Our high-level goal is to create systems that are evolvable: they are easier to modify in response to changing requirements than are systems created using traditional techniques. We have applied our approach to two case studies in evolving software for networked sensor nodes.
John Regehr, Alastair Reid 0001, Kirk Webb, Michael A. Parker, Jay Lepreau
RTSS4
2001 The Impulse Memory Controller
abstract
Impulse is a memory system architecture that adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to remap their data structures in memory. As a result, they can control how their data is accessed and cached, which can improve cache and bus utilization. The Impulse design does not require any modification to processor, cache, or bus designs since all the functionality resides at the memory controller. As a result, Impulse can be adopted in conventional systems without major system changes. We describe the design of the Impulse architecture and how an Impulse memory system can be used in a variety of ways to improve the performance of memory-bound applications. Impulse can be used to dynamically create superpages cheaply, to dynamically recolor physical pages, to perform strided fetches, and to perform gathers and scatters through indirection vectors. Our performance results demonstrate the effectiveness of these optimizations in a variety of scenarios. Using Impulse can speed up a range of applications from 20 percent to over a factor of 5. Alternatively, Impulse can be used by the OS for dynamic superpage creation; the best policy for creating superpages using Impulse outperforms previously known superpage creation policies.
Lixin Zhang 0002, Zhen Fang 0002, Michael A. Parker, Binu K. Mathew, Lambert Schaelicke, John B. Carter, Wilson C. Hsieh, Sally A. McKee
IEEE Trans. Computers3
1999 Impulse: Building a Smarter Memory Controller
abstract
Impulse is a new memory system architecture that adds two important features to a traditional memory controller. First, Impulse supports application-specific optimizations through configurable physical address remapping. By remapping physical addresses, applications control how their data is accessed and cached, improving their cache and bus utilization. Second, Impulse supports prefetching at the memory controller, which can hide much of the latency of DRAM accesses. In this paper we describe the design of the Impulse architecture, and show how an Impulse memory system can be used to improve the performance of memory-bound programs. For the NAS conjugate gradient benchmark, Impulse improves performance by 67%. Because it requires no modification to processor, cache, or bus designs, Impulse can be adopted in conventional systems. In addition to scientific applications, we expect that Impulse will benefit regularly strided memory-bound applications of commercial importance, such as database and multimedia programs.
John B. Carter, Wilson C. Hsieh, Leigh Stoller, Mark R. Swanson, Lixin Zhang 0002, Erik Brunvand, Al Davis, Chen-Chi Kuo, Ravindra Kuramkote, Michael A. Parker, Lambert Schaelicke, Terry Tateyama
HPCA10
1999 Interactive Ray Tracing for Volume Visualization
abstract
Presents a brute-force ray-tracing system for interactive volume visualization. The system runs on a conventional (distributed) shared-memory multiprocessor machine. For each pixel, we trace a ray through a volume to compute the color for that pixel. Although this method has a high intrinsic computational cost, its simplicity and scalability make it ideal for large data sets on current high-end parallel systems. To gain efficiency, several optimizations are used, including a volume bricking scheme and a shallow data hierarchy. These optimizations are used in three separate visualization algorithms: isosurfacing of rectilinear data, isosurfacing of unstructured data, and maximum-intensity projection on rectilinear data. The system runs interactively (i.e. at several frames per second) on an SGI Reality Monster. The graphics capabilities of the Reality Monster are used only for display of the final color image.
Steven G. Parker, Michael A. Parker, Yarden Livnat, Peter-Pike J. Sloan, Charles D. Hansen, Peter Shirley
IEEE Trans. Vis. Comput. Graph.2