Chengyi Juan

dblp:430/6862 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0007-3800-2481ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 61% Storage systems · 30% Cloud and datacenter computing · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory disaggregation
CXL memory
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Storage systems
file systems
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Memory systems
shared memory
1.012026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.312026
No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

vector similarity search · 1.0retrieval-augmented generation · 1.0
YearPublicationVenuePosition
2026 No Atomics, No Problem. Developing a RAG Pipeline for Shared CXL Memory
abstract
We share early experiments with software development for shared CXL memory on the H3 Falcon C5022 CXL switch. We describe how the different stages in the development of a commercial RAG pipeline were impacted by the presence of CXL memory. All of Wikipedia is split into 46 million text passages which are embedded into a high-dimensional embedding space and subjected to heavy load in single-and multi-host experiments. We observe significant performance advantages available to DuckDB and Faiss without changing the software; we measure up to 12× latency reduction with real workloads and 40× reduction with synthetic workloads on CXL, compared to the same queries run with demand paging on NVMe. We explain the current challenges of working with two disjoint cache coherency domains and explain how the Fabric-Attached Memory File System (famfs) and famfs producer-consumer queues provide synchronization patterns without atomic operations on current ×86 CPUs. We then discuss open challenges remaining for high performance atomics and locking mechanisms in shared memory. Lastly we show how famfs page-level interleaving enables near-linear throughput scaling when a second host serves queries from the same Faiss index on shared fabric-attached memory and how shared memory allocations can be orchestrated by Kubernetes in a full vertical RAG deployment.
Alfred Bratterud, Gisle Dankel, Amin Farajianzadeh, John Groves, Chengyi Juan, Joshua Suetterlein, Andrés Márquez 0001, Petter Gustad, Arnt Emil Ingulstad
IEEE Trans. Computers5