EDBT 2026 Demo / reviewers in the wild / expert
Jiahao Shan
dblp:342/8022
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-7941-9933ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 55% Hardware accelerators and domain-specific architectures · 13% Reconfigurable computing and FPGAs · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing systems
molecular dynamics simulation |
1.7 | 2 | 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential · SC 2025 TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion Atoms · PPoPP 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential · SC 2025 |
Reconfigurable computing and FPGAs
molecular dynamics acceleration |
0.9 | 1 | 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential · SC 2025 |
High-performance computing
scientific computing systems |
0.9 | 1 | 2025 | TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion Atoms · PPoPP 2025 |
Processor architecture and microarchitecture
SIMD |
0.7 | 1 | 2023 | Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores · ASPLOS (3) 2023 |
High-performance computing
large-scale simulation |
0.3 | 1 | 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential · SC 2025 |
Compilers and program optimization
vectorization |
0.2 | 1 | 2023 | Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU Cores · ASPLOS (3) 2023 |
Methods — techniques the papers use, named apart from their topics
machine-learning interatomic potential · 1.7phase behavior analysis · 1.3dynamic lane partitioning · 1.3kernel optimization · 0.9ab initio accuracy · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open Set RF Fingerprinting Identification: A Joint Prediction and Siamese Comparison FrameworkabstractRadio Frequency Fingerprinting Identification (RFFI) is a lightweight physical layer identity authentication technique. It identifies the radio frequency device by analyzing the signal feature differences caused by the inevitable minor hardware impairments. However, existing RFFI methods based on closed set recognition struggle to detect unknown unauthorized devices in open environments. Moreover, the feature interference among legitimate devices can further compromise identification accuracy. In this paper, we propose a joint radio frequency fingerprint prediction and siamese comparison (JRFFP-SC) framework for open set recognition. Specifically, we first employ a radio frequency fingerprint prediction network to predict the most probable category result. Then a detailed comparison among the test sample's features with registered samples is performed in a siamese network. The proposed JRFFP-SC framework eliminates inter-class interference and effectively addresses the challenges associated with open set identification. The simulation results show that our proposed JRFFP-SC framework can achieve excellent rogue device detection and generalization capability for classifying devices. Donghong Cai, Jiahao Shan, Ning Gao 0001, Bingtao He, Yingyang Chen, Shi Jin 0002, Pingzhi Fan |
ICC | 2 |
| 2025 | Unsupervised Learning for Solving the Graph Edit Distance
Jiahao Shan |
ICIC (21) | 1 |
| 2025 | TensorMD: Molecular Dynamics Simulation with Ab Initio Accuracy of 50 Billion AtomsabstractMolecular dynamics simulation emerges as an important area that HPC+AI helps to investigate the physical properties, with machine-learning interatomic potentials (MLIPs) being used. General-purpose machine-learning (ML) tools have been leveraged in MLIPs, but they are not perfectly matched with each other, since many optimization opportunities in MLIPs have been missed by ML tools. This inefficiency arises from the fact that HPC+AI applications work with far more computational complexity compared with pure AI scenarios. This paper has developed an MLIP, named TensorMD, independently from any ML tool. TensorMD has been evaluated on two supercomputers and scaled to 51.8 billion atoms, i.e., ~ 3× compared with state-of-the-art. Yucheng Ouyang, Ying Liu 0055, Honghui Shang, Zhenchuan Chen, Jiahao Shan, Huimin Cui, Xiaobing Feng 0002, Xingyu Gao 0003, Haifeng Song 0003, Xin Chen 0023, Rongfen Lin |
PPoPP | 5 |
| 2025 | TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic PotentialabstractAI has been integrated into HPC across various scientific fields, significantly enhancing performance. In molecular dynamics simulations, HPC+AI facilitates the investigation of atomic-scale physical properties using machine-learning interatomic potentials (MLIPs). However, general-purpose ML tools (e.g., TensorFlow) used in MLIPs are not optimally matched, leading to missed optimization opportunities due to the higher computational complexity and greater diversity of HPC+AI applications compared to pure AI scenarios. To address this, we introduce TensorMD, an MLIP independent of existing ML tools, enabling flexible optimizations that standard ML frameworks cannot support. TensorMD outperforms a state-of-the-art MLIP—winner of the 2020 Gordon Bell Prize and built on an ML tool—by 1.88 × on NVIDIA A100 GPU. Additionally, TensorMD was evaluated on two supercomputers with different architectures, achieving significantly reduced time-to-solution and supporting molecular dynamics simulations at scales beyond 50 billion atoms. Yucheng Ouyang, Ying Liu 0055, Xin Chen 0023, Honghui Shang, Zhenchuan Chen, Rongfen Lin, Xingyu Gao 0003, Jiahao Shan, Haifeng Song 0003, Huimin Cui, Xiaobing Feng 0002, Jingling Xue |
SC | 11 |
| 2023 | Occamy: Elastically Sharing a SIMD Co-processor across Multiple CPU CoresabstractSIMD extensions are widely adopted in multi-core processors to exploit data-level parallelism. However, when co-running workloads on different cores, compute-intensive workloads cannot take advantage of the underutilized SIMD lanes allocated to memoryintensive workloads, reducing the overall performance. This paper proposes Occamy, a SIMD co-processor that can be shared by multiple CPU cores, so that their co-running workloads can spatially share its SIMD lanes. The key idea is to enable elastic spatial sharing by dynamically partitioning all the SIMD lanes across different workloads based on their phase behaviors, so that each workload may execute in variable-length SIMD mode. We also introduce an Occamy compiler to support such variable-length vectorization by analyzing such phase behaviors and generating the vectorized code that works with varying vector lengths. We demonstrate that Occamy can improve SIMD utilization, and consequently, performance over three representative SIMD architectures, with negligible chip area cost. Zhongcheng Zhang, Yan Ou, Ying Liu 0055, Chenxi Wang 0005, Yongbin Zhou, Yucheng Ouyang, Jiahao Shan, Ying Wang 0001, Jingling Xue, Huimin Cui, Xiaobing Feng 0002 |
ASPLOS (3) | 9 |
| 2023 | Hierarchical Sparse Estimation of Non-Stationary Channel for Uplink Massive MIMO SystemsabstractThis paper proposes a hierarchical sparse estimation of spatial non-stationarity channel for uplink massive multiple-input multiple-output (MIMO) systems without prior information. Especially, the non-zero rows of non-stationarity channel matrix are estimated according to the in-row correlation in the first layer; while the non-zero elements of the estimated non-zero rows are further refined in the second layer. A row-wise sparse adaptive matching pursuit (SAMP) is used to find the non-zero rows in the first layer of the proposed algorithms, and multiple non-zero rows can be estimated in one iteration, which has higher precision and lower complexity, compared to the conventional SAMP. Different from the existing two-layer iteration algorithms, a threshold is designed to estimate the non-zero elements replacing the iterative algorithm in the second layer. Further, the computation complexity is analyzed and compared. The simulation results demonstrate that the proposed threshold-enhanced hierarchical spatial non-stationary channel estimation algorithms achieve better performance compared to various state-of-the-art baselines in terms of channel coefficient estimation, and computational efficiency. Chongyang Tan, Donghong Cai, Fang Fang 0005, Jiahao Shan, Yanqing Xu 0003, Zhiguo Ding 0001, Pingzhi Fan |
GLOBECOM | 4 |