Marie Nguyen

dblp:159/9467 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-7226-1598ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Systematic CXL Memory Characterization and Performance Analysis at Scale
abstract
Compute Express Link (CXL) has emerged as a pivotal interconnect for memory expansion. Despite its potential, the performance implications of CXL across devices, latency regimes, processors, and workloads remain underexplored. We present Melody, a framework for systematic characterization and analysis of CXL memory performance. Melody builds on an extensive evaluation spanning 265 workloads, 4 real CXL devices, 7 latency levels, and 5 CPU platforms. Melody yields many insights: workload sensitivity to sub-μs CXL latencies (140-410ns), the first disclosure of CXL tail latencies, CPU tolerance to CXL latencies, a novel approach (SPA) for pinpointing CXL bottlenecks, and CPU prefetcher inefficiencies under CXL.
Jinshu Liu, Hamid Hadian, Yuyue Wang 0001, Daniel S. Berger, Marie Nguyen, Xun Jian 0002, Sam H. Noh, Huaicheng Li
ASPLOS (2)5
2025 Can a Client-Server Cache Tango Accelerate Disaggregated Storage?
abstract
Disaggregated storage architectures have become a critical component in modern data centers, offering independent scaling of compute and storage. However, disaggregation introduces performance challenges, particularly due to the overhead of remote storage access. We first conduct a detailed end-to-end analysis of existing caching strategies and identify the critical issues, such as high server resource consumption, data duplication, lack of fair server resources scheduling, inefficient eviction, and prefetching policies. To address these limitations, we present the preliminary design of OrcaCache, an orchestrated, unified caching framework that coordinates between clients and storage servers. OrcaCache aims to carefully shift cache indexing to clients by exposing a global cache view, with an aim to reduce server CPU usage and duplication of data across caches. OrcaCache also aims to improve cache efficiency, adaptiveness, and fairness across servers and clients.
Linjie Ma, Marie Nguyen, Sudarsun Kannan
HotStorage3
2024 OmniCache: Collaborative Caching for Near-storage Accelerators
Yujie Ren, Marie Nguyen, Changwoo Min, Sudarsun Kannan
FAST3
2024 Context-aware Prefetching for Near-Storage Accelerators
abstract
We present ContextPrefetcher, a host-guided high-performant prefetching framework for near-storage accelerators that prefetches data blocks from storage (e.g., NAND) to device-level RAM. Efficiently prefetching data blocks to device-level RAM reduces storage access costs and improves I/O performance. We introduce a novel abstraction, Cross-layered Context (CLC), a virtual entity that spans across the host and the device and is used for identifying, managing, and tracking active and inactive data such as files, objects (within object stores), or a range of blocks. To support efficient prefetching of actively used CLCs to device memory without incurring near-device resource (memory and compute) bottlenecks, ContextPrefetcher delegates prefetching management to the host, guiding near-device compute to prefetch blocks of active CLC. Finally, ContextPrefetcher facilitates the swift reclamation of blocks associated with inactive CLC. Preliminary evaluation against state-of-the-art near-storage accelerator designs demonstrates performance gains of up to 1.34X.
Marie Nguyen, Sanidhya Kashyap, Sudarsun Kannan
HotStorage2
2020 Partial Reconfiguration for Design Optimization
abstract
FPGA designers have traditionally shared a similar design methodology with ASIC designers. Most notably, at design time, FPGA designers commit to a fixed allocation of logic resources to modules in a design. At runtime, some of the occupied resources could be left under-utilized due to hard-to-avoid sources of inefficiencies (e.g., operation dependencies, unbalanced pipelines). With partial reconfiguration (PR), FPGA resources can be re-allocated over time. Therefore, using PR, a designer can attempt to reduce under-utilization with better area-time scheduling. In this paper, we offer definitions, insights, and equations to explain when, how, and why PR-style designs can improve over the performance-area Pareto front of ASIC-style designs (without PR). We first introduce the concept of area-time volume to explain why PR-style designs can improve upon ASIC-style designs. We identify resource under-utilization as an opportunity that can be exploited by PR-style designs. We then present a first-order analytical model to help a designer decide if a PR-style design can be beneficial. When it is the case, the model points to the most suitable PR execution strategy and provides an estimate of the improvement. The model is validated in a case study.
Marie Nguyen, Nathan Serafin, James C. Hoe
FPL1
2019 Quantifying the Benefits of Dynamic Partial Reconfiguration for Embedded Vision Applications
abstract
Dynamic partial reconfiguration (DPR) allows parts of an FPGA to be reprogrammed at runtime (i.e., repurposed). Though DPR has been supported by commercial devices and tools for more than a decade, it has been underutilized, perhaps, due to a shortage of demonstrated use-cases and quantified benefits over static FPGA mapping (without DPR). In this paper, we quantify the benefits of dynamic FPGA mapping (with DPR) over traditional static FPGA mapping for two vision applications deployed on systems with area/device cost, power or energy constraints (i.e., smart car and smart robot). In both applications, the FPGA needs to accelerate multiple tasks at 60 fps. However, all tasks are not required at the same time. In this work, instead of mapping all tasks statically on a large FPGA, the set of tasks needed at a given time is (1) repurposed on a smaller FPGA and (2) still meets the functional and performance requirements (i.e., 60 fps). In the two application examples, we show that dynamic mapping on smaller FPGAs reduces logic resource utilization by up to 3.2x, device cost by up to 10x, and power and energy consumption by up to 30% in comparison with static mapping on larger FPGAs. These benefits are crucial for applications deployed on systems where reducing area/device cost, power and energy is as important as meeting performance requirement.
Marie Nguyen, Robert Tamburo, Srinivasa G. Narasimhan, James C. Hoe
FPL1
2018 Time-Shared Execution of Realtime Computer Vision Pipelines by Dynamic Partial Reconfiguration
abstract
This paper presents an FPGA runtime framework that demonstrates the feasibility of using dynamic partial reconfiguration (DPR) for time-sharing an FPGA by multiple realtime computer vision pipelines. The presented time-sharing runtime framework manages an FPGA fabric that can be round-robin time-shared by different pipelines at the time scale of individual frames. In this new use-case, the challenge is to achieve useful performance despite high reconfiguration time. The paper describes the basic runtime support as well as four optimizations necessary to achieve realtime performance given the limitations of DPR on today's FPGAs. The paper provides a characterization of a working runtime framework prototype on a Xilinx ZC706 development board. The paper also reports the performance of streaming vision pipelines when time-shared.
Marie Nguyen, James C. Hoe
FPL1
2014 GraphGen: An FPGA Framework for Vertex-Centric Graph Computation
abstract
Vertex-centric graph computations are widely used in many machine learning and data mining applications that operate on graph data structures. This paper presents GraphGen, a vertex-centric framework that targets FPGA for hardware acceleration of graph computations. GraphGen accepts a vertex-centric graph specification and automatically compiles it onto an application-specific synthesized graph processor and memory system for the target FPGA platform. We report design case studies using GraphGen to implement stereo matching and handwriting recognition graph applications on Terasic DE4 and Xilinx ML605 FPGA boards. Results show up to 14.6× and 2.9× speedups over software on Intel Core i7 CPU for the two applications, respectively.
Eriko Nurvitadhi, Gabriel Weisz, Yu Wang 0110, Skand Hurkat, Marie Nguyen, James C. Hoe, José F. Martínez, Carlos Guestrin
FCCM5