Jingheng Xu

dblp:180/7981 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 48% GPUs and heterogeneous computing · 24% Performance modeling and evaluation · 12%
Artificial intelligence
1 paper
3D vision · 87% Segmentation and scene understanding · 13%
Network and information security
1 paper
Malware analysis · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth estimation
1.012026
Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance · IEEE Trans. Multim. 2026
Computer vision › 3D vision › depth estimation
panoramic depth estimation
1.012026
Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance · IEEE Trans. Multim. 2026
Malware analysis › ransomware
ransomware detection
0.812024
CanCal: Towards Real-time and Lightweight Ransomware Detection and Response in Industrial Environments · CCS 2024
High-performance computing
performance optimization at scale
0.622019
Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor · ACM Trans. Archit. Code Optim. 2019
Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer · SC 2016
GPUs and heterogeneous computing
GPU computing
0.412019
Optimizing Finite Volume Method Solvers on Nvidia GPUs · IEEE Trans. Parallel Distributed Syst. 2019
GPUs and heterogeneous computing › GPU memory management
GPU memory optimization
0.412019
Optimizing Finite Volume Method Solvers on Nvidia GPUs · IEEE Trans. Parallel Distributed Syst. 2019
Hardware accelerators and domain-specific architectures
hardware-aware optimization
0.412019
Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor · ACM Trans. Archit. Code Optim. 2019
Performance modeling and evaluation
performance tuning
0.412019
Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor · ACM Trans. Archit. Code Optim. 2019
High-performance computing
stencil computation
0.412019
Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor · ACM Trans. Archit. Code Optim. 2019
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312026
Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance · IEEE Trans. Multim. 2026
High-performance computing › scientific computing systems
climate modeling
0.212016
Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer · SC 2016
High-performance computing
scientific computing systems
0.212016
Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer · SC 2016
Compilers and program optimization › program transformation
source-to-source transformation
0.112016
Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer · SC 2016

Methods — techniques the papers use, named apart from their topics

semantic assistance · 1.0attention module · 1.0thread rescheduling · 0.8process filtering · 0.8memory access optimization · 0.8behavioral analysis · 0.8source-to-source translation · 0.5on-chip buffering · 0.5OpenACC · 0.5performance tuning · 0.4algorithm modification · 0.4
YearPublicationVenuePosition
2026 Useg-PanoDepth:Unified $360^{\circ }$ Depth Estimation for Indoor and Outdoor Scenes With Semantic Assistance
abstract
In complex$360^{\circ }$scenes, depth estimation is challenging for small objects and the depth of object boundaries, which cannot be effectively solved with existing works.$360^{\circ }$depth estimation is unable to produce uniform depth estimate findings in both indoor and outdoor settings due to the datasets. In this paper, the Useg-PanoDepth and PanoDepth dataset is proposed to improve the above problems effectively. The Diagonal-aware Attention Module (DAM) effectively estimates small objects in complex scenes. Enhanced Boundary Module (EBM), for enhancing boundary information,can also effectively solve the problem of depth unification of indoor and outdoor scenes. Extensive experiments on our constructed PanoDepth dataset, Useg-PanoDepth achieves SOTA results. The Relative accuracy (deltahttps://github.com/xjh6/Useg-PanoDepth.
Qingling Chang, Jingheng Xu, Yan Cui 0011, Yikui Zhai, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti
IEEE Trans. Multim.2
2024 CanCal: Towards Real-time and Lightweight Ransomware Detection and Response in Industrial Environments
abstract
Ransomware attacks have emerged as one of the most significant cybersecurity threats. Despite numerous methods proposed for detecting and defending against ransomware, existing approaches face two fundamental limitations in large-scale industrial applications: (1) Behavior-based detection engines suffer from the enormous overhead of monitoring all processes and resource constraints for model inference, failing to meet the requirements for real-time detection; (2) Decoy-based detection engines generate an overwhelming number of false positives in large-scale industrial clusters, leading to intolerable disruptions to critical processes and excessive inspection efforts from security analysts. To address these challenges, we propose CanCal, a real-time and lightweight ransomware detection system. Specifically, instead of indiscriminately analyzing all processes, CanCal selectively filters suspicious processes by the monitoring layers and then performs in-depth behavioral analysis to isolate ransomware activities from benign operations, minimizing alert fatigue while ensuring lightweight computational and storage overhead. The experimental results on a large-scale industrial environment (1,761 ransomware, ~ 3 million events, continuous test over 5 months) indicate that CanCal achieves a remarkable 99.65% true positive rate on 555,678 unknown ransomware behavior events, with near-zero false positives. CanCal is as effective as state-of-the-art techniques while enabling rapid inference within 30ms and real-time response within a maximum of 3 seconds. CanCal dramatically reduces average CPU utilization by 91.04% (from 6.7% to 0.6%) and peak CPU utilization by 76.69% (from 26.6% to 6.2%), while avoiding 76.50% (from 3,192 to 750) of the inspection efforts from security analysts. By the time of this writing, CanCal has been integrated into a commercial product and successfully deployed on 3.32 million endpoints for over a year. From March 2023 to April 2024, CanCal successfully detected and thwarted 61 ransomware attacks. A detailed manual forensic analysis of 27 ransomware attacks from March to June 2023 (including 13 n-day exploits and 5 high-risk zero-day attacks) demonstrates the effectiveness of CanCal in combating sophisticated and unknown ransomware threats in real-world scenarios.
Shenao Wang 0001, Feng Dong 0008, Hangfeng Yang, Jingheng Xu, Haoyu Wang 0001
CCS4
2024 PanoDthNet: Depth Estimation Based on Indoor and Outdoor Panoramic Images
Jieyuan Cai, Jingheng Xu, Qingling Chang
PRCV (3)2
2021 Highly scalable parallel genetic algorithm on Sunway many-core processors
Zhiyong Xiao 0001, Jingheng Xu, Qingxiao Sun, Lin Gan 0001
Future Gener. Comput. Syst.3
2019 Million-Core-Scalable Simulation of the Elastic Migration Algorithm on Sunway TaihuLight Supercomputer
abstract
Migration algorithm is one of the most essential methods in seismic application to image the underground geology, and to help scientists and researchers in geophysics exploration better understand the earth system. However, due to the desire in migration algorithm for covering lager region and acquiring better resolution, many tough challenges have to be tackled for current state-of-the-art computing systems. This work optimized and scaled the elastic migration algorithm onto the Sunway TaihuLight supercomputer, one of the most powerful systems of the world. Targeting at the major process, the reverse time migration (RTM) algorithm, a set of algorithmic, process-level, and thread-level optimizations is proposed, to significantly improve the performance (up to 163× speedup in time-to-solution) on Sunway CPU. Our design is successfully scaled to over two million cores (2,662,400 cores in total) on the Sunway TaihuLight supercomputer, with nearly ideal weak-scaling efficiency. The largest run is able to achieve a sustainable performance of processing over 859 billion cells per second.
Lin Gan 0001, Jingheng Xu, Xin Wang 0233, Sihai Wu, Xiaohui Duan, Haohuan Fu, Guangwen Yang 0002
CCGRID2
2019 SunwayLB: Enabling Extreme-Scale Lattice Boltzmann Method Based Computing Fluid Dynamics Simulations on Sunway TaihuLight
abstract
The Lattice Boltzmann Method (LBM) is a relatively new class of Computational Fluid Dynamics methods. In this paper, we report our work on SunwayLB, which enables LBM based solutions aiming for industrial applications. We propose several techniques to boost the simulation speed and improve the scalability of SunwayLB, including a customized multi-level domain decomposition and data sharing scheme, a carefully orchestrated strategy to fuse kernels with different performance constraints for a more balanced workload, and optimization strategies for assembly code, which bring up to 137x speedup. Based on these optimization schemes, we manage to perform the largest direct numerical simulation which involves up to 5.6 trillion lattice cells, achieving 11,245 billion cell updates per second (GLUPS), 77% memory bandwidth utilization and a sustained performance of 4.7 PFlops. We also demonstrate a series of computational experiments for extreme-large scale fluid flow, as examples of real-world applications, to check the validity and performance of our work. The results show that SunwayLB is competent for a practical solution for industrial applications.
Xuesen Chu, Xiaojing Lv, Hongsong Meng, Shupeng Shi, Wenji Han, Jingheng Xu, Haohuan Fu, Guangwen Yang 0002
IPDPS7
2019 Performance Tuning and Analysis for Stencil-Based Applications on POWER8 Processor
abstract
This article demonstrates an approach for combining general tuning techniques with the POWER8 hardware architecture through optimizing three representative stencil benchmarks. Two typical real-world applications, with kernels similar to those of the winning programs of the Gordon Bell Prize 2016 and 2017, are employed to illustrate algorithm modifications and a combination of hardware-oriented tuning strategies with the application algorithms. This work fills the gap between hardware capability and software performance of the POWER8 processor, and provides useful guidance for optimizing stencil-based scientific applications on POWER systems.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Wayne Luk, Guangwen Yang 0002
ACM Trans. Archit. Code Optim.1
2019 Optimizing Finite Volume Method Solvers on Nvidia GPUs
abstract
As scientific applications are increasingly ported to GPUs to benefit from both the powerful computing capacity and high throughput, accelerating explicit solvers for GPU-based finite volume methods is gaining more and more attention. In this paper, based on the detailed analysis of the FVM algorithm, we present a set of novel optimization methods, including the explicit data cache mechanism, optimal global memory loading strategy, as well as the inner-thread rescheduling method, which derives a suitable mapping from the solver algorithm to the underlying GPU hardware architecture, so as to remarkably improve the solving performance of structured mesh based FVM. We demonstrate the impact of our tuning techniques on two widely-used atmospheric dynamic kernels (3-D Euler and 2-D SWE) on five kinds of mainstream GPU platforms, and make a detailed analysis of the different tuning methodologies so as to demonstrate how to select the proper tuning strategy to different applications on various GPU platforms. Specifically, 93.9x speedup is achieved for the 3D Euler solver on Nvidia V100 over one 12-core Intel E5-2697 (v2) CPU, which is a 77 percent improvement compared with the original speedup without adopting the tuning techniques presented in this work.
Jingheng Xu, Guangwen Yang 0002, Haohuan Fu, Wayne Luk, Lin Gan 0001, Wei Xue 0003, Chao Yang 0002, Yong Jiang 0001, Conghui He
IEEE Trans. Parallel Distributed Syst.1
2018 PLZMA: A Parallel Data Compression Method for Cloud Computing
Xin Wang 0233, Lin Gan 0001, Jingheng Xu, Jinzhe Yang, Maocai Xia, Haohuan Fu, Xiaomeng Huang, Guangwen Yang 0002
ICA3PP (3)3
2016 Unleashing the performance potential of CPU-GPU platforms for the 3D atmospheric Euler solver
abstract
As a traditional application on various supercomputers, atmospheric modeling has long been suffering from the low performance efficiency. In this paper, we pick the 3D Euler equation solver (the most essential dynamic component for a non-hydrostatic atmospheric model) as the target application, and explore the maximum performance efficiency that can be achieved on CPU-GPU hybrid architectures. Besides presenting the suitable hybrid domain decomposition methodology and taking proper usage of tuning techniques for both the CPU and GPU parts, we further propose a novel GPU tuning technique, namely the customizable data caching mechanism with thread warp rescheduling scheme, which is specifically designed for the Euler solver. Combining all the optimizing approaches together, remarkable performance boost has been achieved on mainstream GPU architectures including Tesla Fermi C2050, K20×, K40 and K80. Especially, on the latest Tesla K80, we demonstrate a 31.64× speedup over the performance of 12-core E5-2697 CPU. In addition, based on a hybrid CPU-GPU node with two 12-core E5-2697 CPUs and two Tesla K80 GPUs, a sustained double-precision performance of 1.04 Tflops (16% of the peak) is achieved, which is remarkably higher than the efficiency of similar optimizing tasks based on heterogeneous platforms (strictly less than 10%, as demonstrated in the related work). In addition, a nearly linear weak scaling efficiency is achieved which demonstrate the effectiveness of our domain decomposition method.
Haohuan Fu, Jingheng Xu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Wenlai Zhao, Guangwen Yang 0002
ASAP2
2016 Performance optimization of Jacobi stencil algorithms based on POWER8 architecture
abstract
In this paper we choose the widely used Jacobi stencil algorithm as our target program to evaluate the effectiveness of tuning techniques based on the latest POWER8 processor, thus to provide optimization guidelines to similar stencil based algorithms.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Hongbo Peng, Guangwen Yang 0002
ASAP1
2016 Generalized GPU Acceleration for Applications Employing Finite-Volume Methods
abstract
Scientific HPC applications are increasingly ported to GPUs to benefit from both the high throughput and the powerful computing capacity. Many of these applications, such as atmospheric modeling and hydraulic erosion simulation, are adopting the finite volume method (FVM) as the solver algorithm. However, the communication components inside these applications generally lead to a low flop-to-byte ratio and an inefficient utilization of GPU resources. This paper aims at optimizing FVM solver based on the structured mesh. Besides a high-level overview of the finite-volume method as well as its basic optimizations on modern GPU platforms, we further present two generalized tuning techniques including an explicit cache mechanism as well as an inner-thread rescheduling method that tries to achieve a suitable mapping between the algorithm feature and the platform architecture. To the end, we demonstrate the impact of our generalized optimization methods in two typical atmospheric dynamic kernels (Euler and SWE) based on four mainstream GPU platforms. According to the experimental results of Tesla K80, speedups of 24.4x for SWE and 31.5x for Euler could be achieved over a 12-core Intel E5-2697 CPU, which is a great promotion compared with its original speedup (18x and 15.47x) without applying these two methods.
Jingheng Xu, Haohuan Fu, Lin Gan 0001, Chao Yang 0002, Wei Xue 0003, Shizhen Xu, Wenlai Zhao, Bingwei Chen, Guangwen Yang 0002
CCGrid1
2016 Refactoring and optimizing the community atmosphere model (CAM) on the sunway taihulight supercomputer
abstract
This paper reports our efforts on refactoring and optimizing the Community Atmosphere Model (CAM) on the Sunway TaihuLight supercomputer, which uses a many-core processor that consists of management processing elements (MPEs) and clusters of computing processing elements (CPEs). To map the large code base of CAM to the millions of cores on the Sunway system, we take OpenACC-based refactoring as the major approach, and apply source-to-source translator tools to exploit the most suitable parallelism for the CPE cluster, and to fit the intermediate variable into the limited on-chip fast buffer. For individual kernels, when comparing the original ported version using only MPEs and the refactored version using both the MPE and CPE clusters, we achieve up to 22× speedup for the compute-intensive kernels. For the 25km resolution CAM global model, we manage to scale to 24,000 MPEs, and 1,536,000 CPEs, and achieve a simulation speed of 2.81 model years per day.
Haohuan Fu, Junfeng Liao, Wei Xue 0003, Lanning Wang, Dexun Chen, Long Gu, Jinxiu Xu 0001, Nan Ding 0006, Conghui He, Shizhen Xu, Yishuang Liang, Jiarui Fang, Yuanchao Xu 0001, Weijie Zheng 0001, Jingheng Xu, Zhen Zheng, Wanjing Wei, Bingwei Chen, Xiaomeng Huang, Guangwen Yang 0002
SC16