Anthony-Trung Nguyen

dblp:94/6023 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2000
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Memory systems · 93% Interconnection networks and networks-on-chip · 4% Processor architecture and microarchitecture · 4%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache coherence
0.122000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
High-Throughput Coherence Controllers · HPCA 2000
Memory systems › cache coherence
cache-coherent shared memory
0.012000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
Memory systems › cache coherence
coherence controller
0.012000
High-Throughput Coherence Controllers · HPCA 2000
Memory systems › cache coherence
directory-based coherence
0.012000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
Memory systems › shared memory
distributed shared memory
0.012000
High-Throughput Coherence Controllers · HPCA 2000
Memory systems
processing-in-memory
0.012000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
Memory systems › processing-in-memory
processor-in-memory architecture
0.012000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
Processor architecture and microarchitecture
multiprocessor architecture
0.012000
Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration · HPCA 2000
Interconnection networks and networks-on-chip
network interface
0.012000
High-Throughput Coherence Controllers · HPCA 2000

Methods — techniques the papers use, named apart from their topics

simulation · 0.1
YearPublicationVenuePosition
2000 High-Throughput Coherence Controllers
abstract
Recent research shows that the occupancy of the coherence controllers is a major performance bottleneck for distributed cache coherent shared memory multiprocessors. In this paper we study three approaches to alleviating this problem in hardwired coherence controllers, namely, multiple protocol engines, pipelined protocol engines, and split request-response streams. Split request-response streams is an innovative contribution of this paper. The performance of pipelining in the context of coherence controllers has not been presented in the literature. Multiple protocol engines has not been studied in the context of hardwired controllers except for a study of ours and only to a limited extent. Using both commercial and scientific benchmarks on detailed simulation models, we present experimental results that show that each mechanism is highly effective at reducing controller occupancy by as much as 66% and improving execution time by as much as 51%, for applications with high communication bandwidth requirement. A combination of mechanisms further reduces controller occupancy and execution time by as much as 78% and 61%, respectively. Our results show that applying any of the parallel mechanisms in the coherence controllers allows integrating four times as many processors per coherence controller, thus reducing system cost, while maintaining or even exceeding the performance of systems with larger number of coherence controllers.
Ashwini K. Nanda, Anthony-Trung Nguyen, Maged M. Michael, Douglas J. Joseph
HPCA2
2000 Toward a Cost-Effective DSM Organization That Exploits Processor-Memory Integration
abstract
Dramatic increases in the number of transistors that can be integrated on a VLSI chip will soon allow commodity microprocessors to include both processor and a sizable fraction of main memory on chip. Distributed Shared-Memory (DSM) multiprocessors typically use the latest off-the-shelf microprocessors and thus will be affected by the upcoming processor-memory integration. In this paper, we explore how a cache-coherent DSM machine built around Processor-In-Memory (PIM) chips might be cost-effectively organized. To take advantage of the close coupling between processor and memory, we propose tagging the memory and organizing it as a cache. Furthermore, commercial considerations dictate the use of off-the-shelf hardware largely designed for uniprocessors. Consequently, we keep the directory control off-chip. To keep the multiprocessor cheap and simple, and to allow for reconfigurability, directory control is performed by chips that are identical to the ones used as compute nodes. As a result, the machine hardware can be easily reconfigured for computing or coherence-handling depending on the needs of the application. We also propose a cache coherence protocol that is tailored to our architecture: it uses the memory very efficiently while exploiting the large caching space available. Overall, the resulting machine is simple and inexpensive, and delivers performance that is comparable to, and higher than, the more expensive traditional COMA and CC-NUMA organizations, respectively.
Josep Torrellas, Liuxi Yang, Anthony-Trung Nguyen
HPCA3
1996 The Augmint multiprocessor simulation toolkit for Intel x86 architectures
abstract
Most publicly available simulation tools only simulate RISC architectures. These tools cannot capture the instruction mix and memory reference patterns of CISC architectures. We present an overview of Augmint, an execution driven multiprocessor simulation toolkit that fills this gap by supporting Intel x86 architectures. Augmint also supports trace driven simulation for uniprocessors as well as multiprocessors, with minor effort on the part of simulator developers. Augmint runs m4 macro extended C and C++ applications such as those in the SPLASH and SPLASH-2 benchmark suites. Augmint supports a thread based programming model with shared global address space and private stack space. Augmint supports a simulator interface compatible with that of the MINT simulation toolkit for MIPS architectures, thus allowing the reuse of most architecture simulators written for MINT. Augmint simulations run on x8d based uniprocessor systems under Unix or Windows NT. The source code of Augmint is publicly available from http://www.csrd.uiuc.edu/iacoma/augmint.
Anthony-Trung Nguyen, Maged M. Michael, Josep Torrellas
ICCD1