Junhui Luo

dblp:99/4983 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 31% GPUs and heterogeneous computing · 31% Performance modeling and evaluation · 19%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 100%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache coherence
cache-coherent interconnect
1.012026
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation · HPCA 2026
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation
0.312026
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation · HPCA 2026
Distributed systems
remote procedure call
0.312026
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation · HPCA 2026
Distributed systems › remote procedure call
RPC offloading
0.312026
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation · HPCA 2026
Performance modeling and evaluation
simulation
0.312026
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation · HPCA 2026
Programming languages and type systems
language implementation
0.011993
On Implementing a Language for Specifying Active Database Execution Models · VLDB 1993

Methods — techniques the papers use, named apart from their topics

remote atomic operations · 1.0hardware-calibrated simulation · 1.0cache-coherent interconnect · 1.0
YearPublicationVenuePosition
2026 Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
abstract
Conventional heterogeneous computing systems built on PCIe interconnects suffer from inefficient fine-grained host-device interactions and complex programming models. In recent years, many proprietary and open cache-coherent interconnect standards have emerged, among which compute express link (CXL) prevails in the open-standard domain after acquiring several competing solutions. Although CXL-based coherent heterogeneous computing holds the potential to fundamentally transform the collaborative computing mode of CPUs and XPUs, research in this direction remains hampered by the scarcity of available CXL-supported platforms, immature software/hardware ecosystems, and unclear application prospects. This paper presents Cohet, the first CXL-driven coherent heterogeneous computing framework. Cohet decouples the compute and memory resources to form unbiased CPU and XPU pools which share a single unified and coherent memory pool. It exposes a standard malloc/mmap interface to both CPU and XPU compute threads, which share a single per-process page table for user applications, leaving the OS dealing with smart memory allocation, page auto-migration, and management of heterogeneous resources. This design significantly simplifies heterogeneous parallel programming to a level comparable to homogeneous programming. To facilitate Cohet research, we also present a fullsystem cycle-level simulator named SimCXL, which is capable of modeling all CXL sub-protocols and device types. SimCXL has been rigorously calibrated against a real CXL testbed with various CXL memory and accelerators, showing an average simulation error of 3 %. Our evaluation reveals that CXL.cache reduces latency by 68 % and increases bandwidth by$14.4 \times$compared to DMA transfers at cacheline granularity. Building upon these insights, we demonstrate the benefits of Cohet with two killer apps, which are remote atomic operation (RAO) and remote procedure call (RPC). Compared to PCIe-NIC design, CXL-NIC achieves a 5.5 to$40.2 \times$speedup for RAO offloading and an average speedup of$\mathbf{1. 8 6} \times$for$\mathbf{R P C}$(de)serialization offloading.
Yanjing Wang 0007, Lizhou Wu, Sunfeng Gao, Yibo Tang, Junhui Luo, Zicong Wang, Dezun Dong, Nong Xiao 0001
HPCA5
2024 Multiview-Ensemble-Learning-Based Robust Graph Convolutional Networks Against Adversarial Attacks
abstract
Graph neural networks (GNNs) have been widely applied in the Internet of Things (IoT) for the intelligent analysis of data collected by sensors, particularly complex relationships and dependent information between IoT devices. However, recent studies have shown that GNNs are vulnerable to adversarial attacks, which significantly limits their application in safety-critical IoT systems such as smart health monitoring, traffic monitoring, and autonomous driving. To address this issue, in addition to the low feature similarity, this study examines the vulnerability of GNNs empirically and reveals that adversarial perturbations against GNNs tend to have low structural proximity in local neighborhoods. Thus, a natural approach for defending GNNs against adversarial attacks is to utilize the related high-order robust information of the perturbed graphs. In this study, we construct auxiliary views with high-order structure and feature similarity from a perturbed graph and propose a multi-view ensemble learning-based robust graph convolutional network (MV-RGCN). Each base model in the MV-RGCN aggregates the adversarial perturbed graph and the constructed view through an adaptive aggregation mechanism, thereby eliminating the impact of adversarial perturbations. Robust representations of the base models are then integrated using an adaptive ensemble mechanism to generate predictions. Extensive experiments under adversarial attack scenarios demonstrate that the MV-RGCN outperforms state-of-the-art methods and can achieve satisfactory performance without affecting its accuracy on the original graph data. This code is available at https://github.com/thomaslok0516/MVRGCN.
Tao Wu 0003, Junhui Luo, Shaojie Qiao, Chao Wang 0025, Lin Yuan 0002, Xiao Pu 0002, Xingping Xian
IEEE Internet Things J.2
1993 On Implementing a Language for Specifying Active Database Execution Models
Shahram Ghandeharizadeh, Richard Hull 0001, Dean Jacobs, Jaime Castillo, Martha Escobar-Molano, Shih-Hui Lu, Junhui Luo, Chiu Tsang
VLDB7