Albert Cho

dblp:83/10341 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Power Sloshing in Compound Servers for Large-Scale AI Inference Workloads
Albert Cho, Jovan Stojkovic, Leonardo Piga, Abhishek Dhanotia, Sultan Mahmud Sajal, Gefei Zuo, Krishna T. Malladi, Devon Akers, Kalyan Subramanian, Shobhit O. Kanaujia, Alexandros Daglis
ISCA1
2024 StarNUMA: Mitigating NUMA Challenges with Memory Pooling
abstract
Large multi-socket machines are mission-critical high-performance systems for workloads requiring massive memory shared by hundreds of processors. Beyond eight sockets, such systems typically feature multi-hop inter-socket networks, exacerbating the Non-Uniform Memory Access (NUMA) challenge. NUMA effects stem from major disparity in latency and bandwidth characteristics of local and remote memory, often in the 4–10× range. While judicious data placement across the distributed memory's fragments can ameliorate NUMA effects, we observe that in challenging workloads with irregular access patterns, a large fraction of accessed pages are “vagabond”: being actively shared by multiple sockets, they lack a fitting home socket location. On 16-socket systems, such pages incur up to 75% remote memory accesses, which encounter significant latency overheads and bandwidth bottlenecks. StarNUMA introduces a new architectural block for multi-socket architectures to ameliorate the challenge posed by vagabond pages. By leveraging the capabilities of the emerging CXL interconnect, StarNUMA augments a typical NUMA architecture with a memory pool that is directly accessible by every socket in a single high-bandwidth interconnect hop. We show that placement of vagabond pages in StarNUMA's memory pool effectively curbs the latency overheads and queuing delays of the bandwidth-constrained multi-hop inter-socket network, reducing the average memory access time of 16-socket systems by 48%. In turn, faster memory access yields performance improvements of 1.54× on average, and up to 2.17×.
Albert Cho, Alexandros Daglis
MICRO1
2024 COAXIAL: A CXL-Centric Memory System for Scalable Servers
abstract
The memory system is a major performance determinant for server processors. Ever-growing core counts and datasets demand higher memory bandwidth and capacity. DDR—the dominant processor interface to memory—requires a large number of on-chip pins, which is a scarce resource, thus limiting the processor’s memory bandwidth. With limited bandwidth, multiple concurrent memory requests experience significant queuing delays that often overshadow DRAM’s service time and degrade performance. We present CoaXial, a memory system design for throughput-oriented manycore servers that replaces all of the processor’s DDR interfaces with the pin-efficient CXL interface, which offers $4 \times$ higher bandwidth per pin. While such replacement incurs a considerable latency overhead, we demonstrate that, for many workloads, and with careful integration, CXL’s higher bandwidth more than offsets its latency premium. Our evaluation shows that CoaXial improves the performance of manycore throughput-oriented servers by $1.39 \times$ on average and by up to $3 \times$.
Albert Cho, Anish Saxena, Moinuddin K. Qureshi, Alexandros Daglis
SC1
2022 Patching up Network Data Leaks with Sweeper
abstract
Datacenters have witnessed a staggering evolution in networking technologies, driven by insatiable application demands for larger datasets and inter-server data transfers. Modern NICs can already handle 100s of Gbps of traffic, a bandwidth capability equivalent to several memory channels. Direct Cache Access mechanisms like DDIO that contain network traffic inside the CPU’s caches are therefore essential to effectively handle growing network traffic rates. However, a growing body of work reveals instances of a critical DDIO weakness known as “leaky DMA”, occurring when a significant fraction of network traffic leaks from the CPU’s caches to memory. We find that such network data leaks cap the network bandwidth a server can effectively utilize. We identify that a major culprit for such network data leaks are evictions of already consumed dirty network buffers. Our key insight is that buffers already consumed by the application typically need not be written back to memory, as their next reuse will be a full overwrite with new network data by the NIC. We introduce Sweeper, a hardware extension and API that allows applications to mark such consumed network buffers. Hardware then skips writing marked buffers back to memory, drastically reducing memory bandwidth consumption and mitigating the performance penalty of network data leaks. Sweeper boosts a 24-core server’s peak sustainable network bandwidth by up to $2. 6 \times $ as compared to DDIO-based configurations.
Marina Vemmou, Albert Cho, Alexandros Daglis
MICRO2
2011 Gopher: Global observation of Planetary Health and Ecosystem Resources
abstract
GOPHER (Global Observation of Planetary Health and Ecosystem Resources) is a collection of data mining algorithms for detecting global land use and land cover change that builds on a decade of research on spatio temporal data mining at the University of Minnesota; the GOPHER approach to analyzing remote sensing imagery provides such a solution by providing rapid, inexpensive, robust, scalable, and precise detection of land use change. GOPHER algorithms are able to detect historical changes with high precision, and more recent changes with reasonable precision in as little as 8 weeks after changes occur (this is due to a combination of remote sensing data availability and modeling needs). This paper outlines the GOPHER approach and provides some illustrative results to demonstrate its utility.
Ashish Garg 0001, Varun Mithal, Yashu Chamber, Ivan Brugere, Vijay Chaudhari, Marc Dunham, Vikrant Krishna, Sairam Krishnamurthy, Sruthi Vangala, Shyam Boriah, Michael S. Steinbach, Vipin Kumar 0001, Albert Cho, J. D. Stanley, Teji Abraham, Juan Carlos Castilla-Rubio, Christopher Potter, Steven A. Klooster
IGARSS13