Yunhua Fang

dblp:402/3905 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
0009-0009-4718-8825ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Hardware accelerators and domain-specific architectures · 57% Memory systems · 43%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory compression
cache compression
1.012026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026
Memory systems › memory disaggregation
CXL memory
1.012026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator
KV cache compression
1.012026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures › machine learning accelerator › transformer accelerator
LLM inference accelerator
1.012026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.012026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026
Memory systems
DRAM
0.312026
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling · IEEE Trans. Computers 2026

Methods — techniques the papers use, named apart from their topics

precision scaling · 1.0lossless compression · 1.0bit-plane layout · 1.0
YearPublicationVenuePosition
2026 TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
abstract
LLM inference is increasingly limited by memory bandwidth, and the bottleneck worsens at long context as the KV cache grows. CXL memory adds capacity to offload weights and KV, but its link and device-side DDR bandwidth are far below HBM, so decoding stalls once traffic shifts to the CXL tier. Many CXL controllers are starting to add genericlosslesscompression, yet applying commodity codecs directly to standard word-major LLM tensors is largely ineffective, especially for token-major KV streams. We propose TRACE (Traffic-Reduced Architecture for Compression and Elasticity), which preserves the unmodified CXL.mem interface but changes the device-internal representation. It stores tensors in a channel-major, disaggregated bit-plane layout, and applies a KV-specific transform before compression, converting mixed-field words into low-entropy plane streams that commodity codecs can compress. The same substrate enables precision-proportional fetch by reading only the required bit-planes. Across public LLMs, TRACE reduces BF16 weight footprint by 25.2% and BF16 KV footprint by 46.9% losslessly, with per-layer KV ratios peaking at 2.69×. In tracedriven system modeling, once KV spills to CXL, GPT-OSS-120B-MXFP4 improves throughput at 128k tokens from 16.28 to 68.99 tok/s (4.24×). DRAMSim3 shows up to 40.3% lower DRAM access energy under plane-aligned fetch. A 7nm SystemVerilog implementation sustains 256 GB/s device bandwidth. Relative to a CXL controller with generic inline lossless compression, TRACE only adds 7.2% area, 4.7% power, and 6.0% load-to-use latency at 2 GHz and 0.7V.
Rui Xie 0006, Asad Ul Haq, Yunhua Fang, Linsen Ma, Zirak Burzin Engineer, Liu Liu 0017, Tong Zhang 0002
IEEE Trans. Computers3