Hailong Jiang

dblp:137/6173 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CarbonEdge: Carbon-Aware Deep Learning Inference Framework for Sustainable Edge Computing
Guilin Zhang, Wulan Guo, Chuanyi Sun, Hailong Jiang
IWCMC5
2025 Serverless GPU Architecture for Enterprise HR Analytics: A Production-Scale BDaaS Implementation
Guilin Zhang, Wulan Guo, Srinivas Vippagunta, Suchitra Raman, Shreeshankar Chatterjee, Ju Lin, Mary Schladenhauffen, Jeffrey Luo, Hailong Jiang
IEEE Big Data11
2025 Can Large Language Models Understand Intermediate Representations in Compilers?
abstract
Intermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by *Large Language Models* (LLMs) remains underexplored. In this paper, we present an explorative empirical study evaluating the capabilities of six state-of-the-art LLMs—GPT-4, GPT-3, DeepSeek, Gemma 2, Llama 3, and Code Llama—in understanding IRs. Specifically, we assess model performance across four core tasks: *control flow graph reconstruction*, *decompilation*, *code summarization*, and *execution reasoning*. While LLMs exhibit competence in parsing IR syntax and identifying high-level structures, they consistently struggle with instruction-level reasoning, especially in control flow reasoning, loop handling, and dynamic execution. Common failure modes include misinterpreting branching instructions, omitting critical operations, and relying on heuristic reasoning rather than on precise instruction-level logic. Our findings highlight the need for IR-specific enhancements in LLM design. We recommend fine-tuning on structured IR datasets and integrating control-flow-sensitive architectures to improve the models’ effectiveness on IR-related tasks. All the experimental data and source code are publicly available at [https://github.com/hjiang13/LLM4IR](https://github.com/hjiang13/LLM4IR).
Hailong Jiang, Yao Wan 0001, Bo Fang 0002, Hongyu Zhang 0002, Ruoming Jin, Qiang Guan
ICML1
2025 KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-Scaling
abstract
Autoscaling GPU inference workloads in Kubernetes remains challenging due to the reactive and threshold-based nature of default mechanisms such as the Horizontal Pod Autoscaler (HPA), which struggle under dynamic and bursty traffic patterns and lack integration with GPU-level metrics. We present KIS-S, a unified framework that combines KISim, a GPU-aware Kubernetes Inference Simulator, with KIScaler, a Proximal Policy Optimization (PPO)-based autoscaler. KISim enables safe, high-fidelity scheduling emulation with real GPU hardware and Prometheus integration, while KIScaler learns latency-aware and resource-efficient scaling policies entirely in simulation. KIScaler observes system metrics via Prometheus and adjusts replica counts via the Kubernetes API. We evaluate KIS-S across four synthetic traffic patterns-ramp, periodic, random, and spike-and compare it against conventional baselines including HPA and fixed-resource deployments. Despite training with synthetic feedback due to single-GPU hardware constraints, KIScaler's moving average reward improves from 1.05 to 1.84 (a 75.2 % increase) over 100 training episodes, reduces P95 latency by up to$6.7 \times$over CPU-only baselines, and generalizes across all traffic patterns without retraining. These results highlight the value of combining simulation and learning, bridging the gap between reactive autoscaling and intelligent orchestration for scalable, GPU-accelerated Kubernetes environments.
Guilin Zhang, Wulan Guo, Qiang Guan, Hailong Jiang
IPCCC5
2024 HAppA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded
abstract
High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method - HAppA-LSTM - achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of key words representing the source code. A comprehensive importance analysis of these key words further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.
Hailong Jiang, Bo Fang 0002, Kevin J. Barker, Ruoming Jin, Qiang Guan
SRDS1
2023 Visilience: An Interactive Visualization Framework for Resilience Analysis using Control-Flow Graph
abstract
Soft errors have become one of the main concerns for the resilience of HPC applications, as these errors can cause HPC applications to generate serious outcomes such as silent data corruption (SDC). Many approaches have been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to perform resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of detail, which can create obstacles to compare and explore the resilience of the HPC program execution. To this end, we present Visilience, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, Visilience leverages an effective visualization approach, Control Flow Graph (CFG) to present a function execution. Furthermore, three widely used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly integrated into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework Visilience.
Hailong Jiang, Shaolun Ruan, Bo Fang 0002, Yong Wang 0021, Qiang Guan
PRDC1
2022 BatchLens: A Visualization Approach for Analyzing Batch Jobs in Cloud Systems
abstract
Cloud systems are becoming increasingly powerful and complex. It is highly challenging to identify anomalous execution behaviors and pinpoint problems by examining the overwhelming intermediate results/states in complex application workflows. Domain scientists urgently need a friendly and functional interface to understand the quality of the computing services and the performance of their applications in real time. To meet these needs, we explore data generated by job schedulers and investigate general performance metrics (e.g., utilization of CPU, memory and disk I/O). Specifically, we propose an interactive visual analytics approach, BatchLens, to provide both providers and users of cloud service with an intuitive and effective way to explore the status of system batch jobs and help them conduct root-cause analysis of anomalous behaviors in batch jobs. We demonstrate the effectiveness of BatchLens through a case study on the public Alibaba bench workload trace datasets.
Shaolun Ruan, Yong Wang 0021, Hailong Jiang, Weijia Xu, Qiang Guan
DATE3
2021 Genome-wide variant-based study of genetic effects with the largest neuroanatomic coverage
abstract
BACKGROUND: Brain image genetics provides enormous opportunities for examining the effects of genetic variations on the brain. Many studies have shown that the structure, function, and abnormality (e.g., those related to Alzheimer's disease) of the brain are heritable. However, which genetic variations contribute to these phenotypic changes is not completely clear. Advances in neuroimaging and genetics have led us to obtain detailed brain anatomy and genome-wide information. These data offer us new opportunities to identify genetic variations such as single nucleotide polymorphisms (SNPs) that affect brain structure. In this paper, we perform a genome-wide variant-based study, and aim to identify top SNPs or SNP sets which have genetic effects with the largest neuroanotomic coverage at both voxel and region-of-interest (ROI) levels. Based on the voxelwise genome-wide association study (GWAS) results, we used the exhaustive search to find the top SNPs or SNP sets that have the largest voxel-based or ROI-based neuroanatomic coverage. For SNP sets with >2 SNPs, we proposed an efficient genetic algorithm to identify top SNP sets that can cover all ROIs or a specific ROI. RESULTS: We identified an ensemble of top SNPs, SNP-pairs and SNP-sets, whose effects have the largest neuroanatomic coverage. Experimental results on real imaging genetics data show that the proposed genetic algorithm is superior to the exhaustive search in terms of computational time for identifying top SNP-sets. CONCLUSIONS: We proposed and applied an informatics strategy to identify top SNPs, SNP-pairs and SNP-sets that have genetic effects with the largest neuroanatomic coverage. The proposed genetic algorithm offers an efficient solution to accomplish the task, especially for identifying top SNP-sets.
Peihua Bao, Yanzhao Li, Hailong Jiang, Shiaofen Fang
BMC Bioinform.8
2020 Chaser: An Enhanced Fault Injection Tool for Tracing Soft Errors in MPI Applications
abstract
Resilient computation has been an emerging topic in the field of high-performance computing (HPC). In particular, studies show that tolerating faults on leadership-class supercomputers (such as exascale supercomputers) is expected to be one of the main challenges. In this paper, we utilize dynamic binary instrumentation and virtual machine based fault injection to emulate soft errors and study the soft errors' impact on the behavior of applications. We propose Chaser, a fine-grained, accountable, flexible, and efficient fault injection framework built on top of QEMU. Chaser offers just-in-time fault injection, the ability to trace fault propagation, and flexible and programable interfaces. In the case study, we demonstrate the usage of Chaser on Matvec and a real DOE mini MPI application
Qiang Guan, Xunchao Hu, Terence Grove, Bo Fang 0002, Hailong Jiang, Heng Yin 0001, Nathan DeBardeleben
DSN5
2020 Semi-Supervised Subspace Learning for Pattern Classification via Robust Low Rank Constraint
Ao Li 0002, Ruoqi An, Guanglu Sun, Xin Liu 0085, Qidi Wu, Hailong Jiang
Mob. Networks Appl.7
2013 IM-Torch: Interference Mitigation via Traffic Offloading in Macro/Femtocell+WiFi HetNets
abstract
Interference management is a hot issue in Heterogeneous Networks (HetNets), which is very crucial for the performance promotion in heterogeneous cellular networks with full frequency reuse. Focusing on mitigating the interference between Macrocell and Femtocell, we propose the IM-Torch (Interference Mitigation via Traffic Offloading in Macro/Femtocell + WiFi Heterogeneous Networks) scheme to handle this problem via traffic offloading. We formulate it as a Mixed Integer Nonlinear Program which is hard to solve, and design a two-step heuristic algorithm to solve this problem. In the first self-scheduling step, Femtocell tries to reallocate the power and PRBs (Physical Resource Blocks) for HUEs (Home User Equipment) to mitigate the interference. After the failure of the first step, Femtocell offloads some necessary HUEs with data services to WiFi and re-adjusts the resources for HUEs to alleviate the interference. The analysis and simulations validate that IM-Torch scheme can greatly alleviate the interference in HetNets and thus improve the system total throughput of Femtocell while guarantee the QoS of Macrocell and HUEs with real time applications.
Liang Wang 0014, Min Sheng, Yan Zhang 0006, Hailong Jiang
PIMRC4