EDBT 2026 Demo / reviewers in the wild / expert
Michael A. Parker
dblp:34/9601
· DBLP profile ↗
11ranked-venue papers
0as first author
0since 2021 · last 2012
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8Graphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Memory systems · 64% Embedded and real-time systems · 24% Parallel and multicore computing · 10% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 77% Geometric modeling and processing · 23% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › memory management
address remapping |
0.1 | 2 | 2006 | Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006 Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Memory systems › shared memory
distributed shared memory |
0.1 | 1 | 2006 | Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006 |
Memory systems
memory controller |
0.1 | 2 | 2001 | The Impulse Memory Controller · IEEE Trans. Computers 2001 Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Embedded and real-time systems › real-time scheduling
hierarchical scheduling |
0.0 | 1 | 2003 | Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003 |
Embedded and real-time systems
real-time scheduling |
0.0 | 1 | 2003 | Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003 |
Embedded and real-time systems › real-time scheduling › schedulability analysis
response time analysis |
0.0 | 1 | 2003 | Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003 |
Memory systems › memory access
indirect addressing |
0.0 | 1 | 2001 | The Impulse Memory Controller · IEEE Trans. Computers 2001 |
Visualization and visual analytics
volume visualization |
0.0 | 1 | 1999 | Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999 |
Parallel and multicore computing › parallel computing
parallel rendering |
0.0 | 1 | 1999 | Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999 |
Memory systems › cache
prefetching |
0.0 | 1 | 1999 | Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Memory systems
cache coherence |
0.0 | 1 | 2006 | Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006 |
Parallel and multicore computing
multiprocessor system |
0.0 | 1 | 2006 | Efficient address remapping in distributed shared-memory systems · ACM Trans. Archit. Code Optim. 2006 |
Memory systems
memory-bound computation |
0.0 | 2 | 2001 | The Impulse Memory Controller · IEEE Trans. Computers 2001 Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Parallel and multicore computing › parallel computing › parallel program analysis › concurrency bug detection
race detection |
0.0 | 1 | 2003 | Evolving real-time systems using hierarchical scheduling and concurrency analysis · RTSS 2003 |
Memory systems
DRAM |
0.0 | 1 | 2001 | The Impulse Memory Controller · IEEE Trans. Computers 2001 |
Geometric modeling and processing
isosurface extraction |
0.0 | 1 | 1999 | Interactive Ray Tracing for Volume Visualization · IEEE Trans. Vis. Comput. Graph. 1999 |
High-performance computing
scientific computing |
0.0 | 1 | 1999 | Impulse: Building a Smarter Memory Controller · HPCA 1999 |
Methods — techniques the papers use, named apart from their topics
simulation · 0.1volume bricking · 0.0shallow data hierarchy · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2012 | Active memory controller
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Sally A. McKee, Ali Ibrahim, Michael A. Parker, Xiaowei Jiang |
J. Supercomput. | 6 |
| 2007 | Active memory operationsabstractThe performance of modern microprocessors is increasingly limited by their inability to hide main memory latency. The problem is worse in large-scale shared memory systems, where remote memory latencies are hundreds, and soon thousands, of processor cycles. To mitigate this problem, we propose the use of Active Memory Operations (AMOs), in which select operations can be sent to and executed on the home memory controller of data. AMOs can eliminate significant number of coherence messages, minimize intranode and internode memory traffic, and create opportunities for parallelism. Our implementation of AMOs is cache-coherent and requires no changes to the processor core or DRAM chips. Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Ali Ibrahim, Michael A. Parker |
ICS | 5 |
| 2006 | The design and utility of the ML-RSIM system simulator
Lambert Schaelicke, Michael A. Parker |
J. Syst. Archit. | 2 |
| 2006 | Efficient address remapping in distributed shared-memory systemsabstractAs processor performance continues to improve at a rate much higher than DRAM and network performance, we are approaching a time when large-scale distributed shared memory systems will have remote memory latencies measured in tens of thousands of processor cycles. The Impulse memory system architecture adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to control how data is accessed and cached and thereby improve cache and bus utilization and reduce the number of memory accesses required. Previous Impulse work focuses on uniprocessor systems and relies on software to flush processor caches when necessary to ensure data coherence. In this paper, we investigate an extension of Impulse to multiprocessor systems that extends the coherence protocol to maintain data coherence without requiring software-directed cache flushing. Specifically, the multiprocessor Impulse controller can gather/scatter data across the network while its coherence protocol guarantees that each gather request gets coherent data and each scatter request updates every coherent replica in the system. Our simulation results demonstrate that the proposed system can significantly outperform conventional systems, achieving an average speedup of 9X on four memory-bound benchmarks on a 32-processor system. Lixin Zhang 0002, Michael A. Parker, John B. Carter |
ACM Trans. Archit. Code Optim. | 2 |
| 2005 | Fast synchronization on shared-memory multiprocessors: An architectural approach
Zhen Fang 0002, Lixin Zhang 0002, John B. Carter, Liqun Cheng, Michael A. Parker |
J. Parallel Distributed Comput. | 5 |
| 2004 | A low power architecture for embedded perceptionabstractRecognizing speech, gestures, and visual features are important interface capabilities for future embedded mobile systems. Unfortunately, the real-time performance requirements of complex perception applications cannot be met by current embedded processors and often even exceed the performance of high performance microprocessors whose energy consumption far exceeds embedded energy budgets. Though custom ASICs provide a solution to this problem, they incur expensive and lengthy design cycles and are inflexible. This paper introduces a VLIW perception processor which uses a combination of clustered function units, compiler controlled dataflow and compiler controlled clock-gating in conjunction with a scratch-pad memory system to achieve high performance for perceptual algorithms at low energy consumption. The architecture is evaluated using ten benchmark applications taken from complex speech and visual feature recognition, security, and signal processing domains. The energy-delay product of a 0.13μ implementation of this architecture is compared against ASICs and general purpose processors. Using a combination of Spice simulations and real processor power measurements, we show that the cluster running at 1 GHz clock frequency outperforms a 2.4 GHz Pentium 4 by a factor of 1.75 while simultaneously achieving 159 times better energy delay product than a low power Intel XScale embedded processor. Binu K. Mathew, Al Davis, Michael A. Parker |
CASES | 3 |
| 2004 | Energy efficient cluster co-processors [3G wireless applications]abstractNew 3G wireless algorithms require more performance than can be currently provided by embedded processors. ASICs provide the necessary performance but are costly to design and sacrifice generality. This paper introduces a clustered VLIW coprocessor approach that organizes the execution and storage resources differently than a traditional general-purpose processor or DSP. The execution units of the coprocessor are clustered and embedded in a rich set of communication resources. Fine grain control of these resources is imposed by a wide-word horizontal micro-code program. The advantages of this approach are quantified on a suite of six algorithms that are taken from both traditional DSP applications and from the new 3G cellular telephony domain. The result is surprising. The execution clusters retain much of the generality of a conventional processor while simultaneously improving performance by one to two orders of magnitude and by reducing energy-delay by three to four orders of magnitude when compared to a conventional embedded processor such as the Intel XScale. Ali Ibrahim, Michael A. Parker, Al Davis |
ICASSP (5) | 2 |
| 2003 | Evolving real-time systems using hierarchical scheduling and concurrency analysisabstractWe have developed a new way to look at real-time and embedded software: as a collection of execution environments created by a hierarchy of schedulers. Common schedulers include those than run interrupts, bottom-half handlers, threads, and events. We have created algorithms for deriving response times, scheduling overheads, and blocking terms for tasks in systems containing multiple execution environments. We have also created task scheduler logic, a formalism that permits checking systems for race conditions and other errors. Concurrency analysis of low-level software is challenging because there are typically several kinds of locks, such as thread mutexes and disabling interrupts, and groups of cooperating tasks may need to acquire some, all or none of the available types of locks to create correct software. Our high-level goal is to create systems that are evolvable: they are easier to modify in response to changing requirements than are systems created using traditional techniques. We have applied our approach to two case studies in evolving software for networked sensor nodes. John Regehr, Alastair Reid 0001, Kirk Webb, Michael A. Parker, Jay Lepreau |
RTSS | 4 |
| 2001 | The Impulse Memory ControllerabstractImpulse is a memory system architecture that adds an optional level of address indirection at the memory controller. Applications can use this level of indirection to remap their data structures in memory. As a result, they can control how their data is accessed and cached, which can improve cache and bus utilization. The Impulse design does not require any modification to processor, cache, or bus designs since all the functionality resides at the memory controller. As a result, Impulse can be adopted in conventional systems without major system changes. We describe the design of the Impulse architecture and how an Impulse memory system can be used in a variety of ways to improve the performance of memory-bound applications. Impulse can be used to dynamically create superpages cheaply, to dynamically recolor physical pages, to perform strided fetches, and to perform gathers and scatters through indirection vectors. Our performance results demonstrate the effectiveness of these optimizations in a variety of scenarios. Using Impulse can speed up a range of applications from 20 percent to over a factor of 5. Alternatively, Impulse can be used by the OS for dynamic superpage creation; the best policy for creating superpages using Impulse outperforms previously known superpage creation policies. Lixin Zhang 0002, Zhen Fang 0002, Michael A. Parker, Binu K. Mathew, Lambert Schaelicke, John B. Carter, Wilson C. Hsieh, Sally A. McKee |
IEEE Trans. Computers | 3 |
| 1999 | Impulse: Building a Smarter Memory ControllerabstractImpulse is a new memory system architecture that adds two important features to a traditional memory controller. First, Impulse supports application-specific optimizations through configurable physical address remapping. By remapping physical addresses, applications control how their data is accessed and cached, improving their cache and bus utilization. Second, Impulse supports prefetching at the memory controller, which can hide much of the latency of DRAM accesses. In this paper we describe the design of the Impulse architecture, and show how an Impulse memory system can be used to improve the performance of memory-bound programs. For the NAS conjugate gradient benchmark, Impulse improves performance by 67%. Because it requires no modification to processor, cache, or bus designs, Impulse can be adopted in conventional systems. In addition to scientific applications, we expect that Impulse will benefit regularly strided memory-bound applications of commercial importance, such as database and multimedia programs. John B. Carter, Wilson C. Hsieh, Leigh Stoller, Mark R. Swanson, Lixin Zhang 0002, Erik Brunvand, Al Davis, Chen-Chi Kuo, Ravindra Kuramkote, Michael A. Parker, Lambert Schaelicke, Terry Tateyama |
HPCA | 10 |
| 1999 | Interactive Ray Tracing for Volume VisualizationabstractPresents a brute-force ray-tracing system for interactive volume visualization. The system runs on a conventional (distributed) shared-memory multiprocessor machine. For each pixel, we trace a ray through a volume to compute the color for that pixel. Although this method has a high intrinsic computational cost, its simplicity and scalability make it ideal for large data sets on current high-end parallel systems. To gain efficiency, several optimizations are used, including a volume bricking scheme and a shallow data hierarchy. These optimizations are used in three separate visualization algorithms: isosurfacing of rectilinear data, isosurfacing of unstructured data, and maximum-intensity projection on rectilinear data. The system runs interactively (i.e. at several frames per second) on an SGI Reality Monster. The graphics capabilities of the Reality Monster are used only for display of the final color image. Steven G. Parker, Michael A. Parker, Yarden Livnat, Peter-Pike J. Sloan, Charles D. Hansen, Peter Shirley |
IEEE Trans. Vis. Comput. Graph. | 2 |