James Lin 0001

dblp:54/5881-1 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0003-4404-5027ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 5 since 2021Security and privacy · 2 · 2 since 2021Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
Yunzhe Li 0001, Hongzi Zhu, James Lin 0001, Shan Chang, Minyi Guo
NDSS4
2026 On the Availability Risks of Production LLM Services Under Unbounded Inference
abstract
Large Language Models (LLMs) have become foundational components in a wide range of applications, including natural language understanding and generation, embodied intelligence, and scientific discovery. As their computational requirements continue to grow, these models are increasingly deployed as cloud-based services, allowing users to access powerful LLMs via the Internet. However, this deployment model introduces a new class of threat: denial-of-service (DoS) attacks via unbounded reasoning, where adversaries craft specially designed inputs that cause the model to enter excessively long or infinite generation loops. These attacks can exhaust backend compute resources, degrading or denying service to legitimate users. To mitigate such risks, many LLM providers adopt a closed-source, black box setting to obscure model internals. In this paper, we propose ThinkTrap, a novel input-space optimization framework for DoS attacks against LLM services even in black-box environments. The core idea of ThinkTrap is to first map discrete tokens into a continuous embedding space, then undertake efficient black-box optimization in a low-dimensional subspace exploiting input sparsity. The goal of this optimization is to identify adversarial prompts that induce extended or non-terminating generation across several state-of-the-art LLMs, achieving DoS with minimal token overhead. We evaluate ThinkTrap across multiple commercial, closed-source LLM services and observe that it can consistently induce abnormally long outputs and noticeable response-side degradation under black-box access. To further quantify the system-level impact of the attack, we conduct controlled experiments on private LLM deployments, where ThinkTrap reduces service throughput to as low as 1% of its original capacity and, in extreme cases, induces complete service failure due to resource exhaustion.
Yunzhe Li 0001, Hongzi Zhu, James Lin 0001, Shan Chang, Minyi Guo
IEEE Trans. Dependable Secur. Comput.4
2025 Design and Operation of Elastic GPU-Pooling on Campus
Kaicheng Guo, Yun Wang 0039, Semakin Anton, Tovmachenko Dmitry, Jiajie Sheng, Jianwen Wei, James Lin 0001, Zhengwei Qi, Haibing Guan
Euro-Par (1)8
2024 TacVar: Tackling Variability in Short-Interval Timing Measurements on X86 Processors
abstract
When timing short-interval loops, different timing methods may produce different results caused by timing fluctuations. These unstable results prevent us from drawing reliable conclusions in processor performance benchmark. To tackle this issue, we proposed the TacVar framework including three components: 1) we developed the TVkern benchmark to highlight misleading tail-end timing results caused by timing fluctuations; 2) we built the TVconv convolution model to describe the mechanism of coupling between timing fluctuations and performance variability; 3) we designed the TVfilt algorithm to filter out timing fluctuations from measurement results using in-situ timing fluctuation samples. We evaluated TacVar with three TVkern and two real-world stencil kernels on two mainstream X86 processors. The results showed TacVar can reduce measurement deviations caused by timing fluctuations by up to 99.0%.
Qiucheng Liao, James Lin 0001
CCGrid2
2023 Characterizing Performance Impacts of Subnormal Numbers on Vector Instructions and Transcendental Functions
abstract
We investigated the performance impact of IEEE-754 double-precision floating-point subnormal numbers, focusing on vector arithmetic and transcendental functions across Intel, AMD, and HiSilicon CPUs. We developed a benchmark tool, SNBench, and uncovered that subnormal numbers can extend instruction latency and reciprocal throughput by up to 162 cycles on Intel, 13 cycles on AMD, and 2 cycles on HiSilicon processors. The impact is more pronounced in transcendental functions, with performance losses reaching up to 562 cycles on Intel, 404 cycles on AMD, and 62 cycles on HiSilicon CPUs.
James Lin 0001, Qiucheng Liao
ICPADS2
2022 SchedP: I/O-aware Job Scheduling in Large-Scale Production HPC Systems
Kaiyue Wu, Jianwen Wei, James Lin 0001
NPC3
2021 tcFFT: A Fast Half-Precision FFT Library for NVIDIA Tensor Cores
abstract
Mixed-precision computing becomes an inevitable trend for HPC and AI applications due to the increasing using mixed-precision units such as NVIDIA Tensor Cores. Fast Fourier transform (FFT) is one of the most widely-used scientific kernels and hence mixed-precision FFT is highly demanded. However, few existing FFT libraries (or algorithms) can support universal size of FFTs on Tensor Cores. Therefore, we proposed tcFFT, a fast half-precision FFT library on Tensor Cores that can support universal size of 1D and 2D FFTs. Our work consists of two parts: framework design and performance optimizations. We designed the tcFFT library framework to support all power-of-two size and multi-dimension of FFTs; we applied two performance optimizations, one to use Tensor Cores efficiently and the other to ease GPU memory bottlenecks. We evaluated tcFFT with a wide range size of 1D and 2D FFTs on NVIDIA V100 and A100 GPUs. The results show that tcFFT can outperform 1.29X-3.24X and 1.10X-3.03X higher on average than NVIDIA cuFFT v11.0 in FP16 on V100 and A100, respectively.
Bin-Rui Li, Shenggan Cheng, James Lin 0001
CLUSTER3
2020 CUBE - Towards an Optimal Scaling of Cosmological N-body Simulations
abstract
N-body simulations are essential tools in physical cosmology to understand the large-scale structure (LSS) formation of the universe. Large-scale simulations with high resolution are important for exploring the substructure of universe and for determining fundamental physical parameters like neutrino mass. However, traditional particle-mesh (PM) based algorithms use considerable amounts of memory, which limits the scalability of simulations. Therefore, we designed a two-level PM algorithm CUBE towards optimal performance in memory consumption reduction. By using the fixed-point compression technique, CUBE reduces the memory consumption per N-body particle to only 6 bytes, an order of magnitude lower than the traditional PM-based algorithms. We scaled CUBE to 512 nodes (20,480 cores) on an Intel Cascade Lake based supercomputer with ≃95% weak-scaling efficiency. This scaling test was performed in Cosmo-π - a cosmological LSS simulation using ≃4.4 trillion particles, tracing the evolution of the universe over ≃13.7 billion years. To our best knowledge, Cosmo-π is the largest completed cosmological N-body simulation. We believe CUBE has a huge potential to scale on exascale supercomputers for larger simulations.
Shenggan Cheng, Hao-Ran Yu, Derek Inman, Qiucheng Liao, Qiaoya Wu, James Lin 0001
CCGRID6
2020 NeoMPX: Characterizing and Improving Estimation of Multiplexing Hardware Counters for PAPI
abstract
Modern processors provide hundreds of low-level hardware events (such as cache miss rate), but offer only a small number (usually 6-12) of hardware counters to collect these events due to limited register resource. Multiplexing (MPX) is an estimation-based technique designed to collect hardware events simultaneously with few hardware counters. However, the low-accuracy of existing MPX methods prevents this technique from wide usage in real conditions. To obtain accurate and reliable hardware counter values, we conducted this work in three steps: 1) to explore the root cause of inaccuracy, we characterized the estimation errors of MPX, and found that estimation errors arise from the outliers in PAPI; 2) to eliminate these outliers and improve MPX accuracy, we proposed two non-linear growth rate gradient estimation methods: divided curved-area method (DCAM) and curved-area method (CAM); 3) based on these two methods, we developed a new MPX library for PAPI, NeoMPX. We evaluated NeoMPX with six Rodinia benchmarks on four mainstream x86 and ARM server processors, and compared the results with PAPI default MPX and two other state-of-art MPX methods, DIRA, and TAM. Evaluations show that, for collecting 16 evaluated hardware events, our methods can improve up to 59% accuracy than PAPI default MPX, and achieve 36% and 5% higher accuracy than DIRA and TAM, respectively. We have open-sourced NeoMPX and expect it to enable PAPI MPX for practical usage.
Yichao Wang 0001, Jin-Kun Chen, Sicheng Zuo, Xiaoming Su, James Lin 0001
CLUSTER6
2020 FTL: A Universal Framework for Training Low-Bit DNNs via Feature Transfer
Kunyuan Du, Ya Zhang 0002, Haibing Guan, Qi Tian 0001, Yanfeng Wang 0001, Shenggan Cheng, James Lin 0001
ECCV (25)7
2019 An Empirical Study of HPC Workloads on Huawei Kunpeng 916 Processor
abstract
The ARM-based server processors have been gaining momentum in high performance computing (HPC). While not designed specifically for HPC, Huawei Kunpeng 916 processor has 32 ARMv8 cores and is tempting for HPC workloads. However, its potential remains unknown. To throughly understand the potential, we conducted a systematic evaluation in three steps by using: 1) three well-known benchmarks (HPL, STREAM, and LMbench); 2) three typical scientific kernels (SpMV, N-body, and GEMM); 3) three widely used mini-apps (TeaLeaf, Neutral, and SNAP) and a real-world application GTC-P. We compared the performance results of Kunpeng 916 with that of Intel Xeon E5-2680v3/4 (Haswell/Broadwell). The evaluation results show that Kunpeng 916 has higher memory bandwidth than the two Intel processors, thus it can achieve compelling performance for running memory bound HPC applications.
Yichao Wang 0001, Jin-Kun Chen, Bin-Rui Li, Sicheng Zuo, William Tang 0002, Bei Wang 0002, Qiucheng Liao, James Lin 0001
ICPADS9
2018 Optimizing Preconditioned Conjugate Gradient on TaihuLight for OpenFOAM
abstract
Porting the domain-specific software OpenFOAM onto the TaihuLight supercomputer is a challenging task, due to the highly memory-bound nature of both the supercomputer's processor (SW26010) and the software's liner solvers. Our study tackles this technical challenge, in three steps, by optimizing the linear solvers, such as Preconditioned Conjugate Gradient (PCG), on the SW26010. First, in order to minimize the all_reduce communication cost of PCG, we developed a new algorithm RNPCG, a non-blocking PCG leveraging the on-chip register communication. Second, we optimized three key kernels of the PCG, including proposing a localized version of the Diagonal-based Incomplete Cholesky (LDIC) preconditioner. Third, to scale the RNPCG on TaihuLight, we designed the three-level non-blocking all_reduce operations. With these three steps, we implemented the RNPCG in OpenFOAM. The experimental results on TaihuLight show that 1) compared with the default implementations of OpenFOAM, the RNPCG and the LDIC on a single-core group of SW26010 can achieve a maximum speedup of 8.9X and 3.1X, respectively; 2) the scalable RNPCG can outperform the standard PCG both in the strong and the weak scaling up to 66,560 cores.
James Lin 0001, Minhua Wen, Delong Meng, Xin Liu 0020, Akira Nukada, Satoshi Matsuoka
CCGrid1
2018 OpenACC vs the Native Programming on Sunway TaihuLight: A Case Study with GTC-P
abstract
Sunway TaihuLight is China's recent top-ranked supercomputer worldwide that was the first to be built entirely with home-grown processors. This supercomputer can be programmed with two approaches: directive-based OpenACC and native programming. These approaches are studied here using GTC-P, a particle-in-cell code for investigating micro-turbulence in magnetic fusion plasmas. We have compared the performance and programming efforts between the OpenACC and the native version of GTC-P. Associated results show that in the OpenACC version, the kernel with irregular memory access becomes the main performance bottleneck due to poor data locality. To address this issue, we have applied two optimizations on the native version: (1) register level communication (RLC); and (2) an "asynchronization" strategy. With these two optimizations, the native version can achieve up to 2.5X speedup for the memory-bound kernel compared with the OpenACC version. In addition, we have now scaled GTC-P on 4,259,840 cores of TaihuLight and demonstrate performance comparisons with several world-leading supercomputers.
Linjin Cai, Yichao Wang 0001, William Tang 0002, Bei Wang 0002, Stéphane Ethier, James Lin 0001
CLUSTER7
2018 Optimizing Deep Learning Frameworks Incrementally to Get Linear Speedup: A Comparison Between IPoIB and RDMA Verbs
abstract
Deeper models and larger datasets are two major ingredients for applying deep learning (DL) on real-world problems, which inevitably shifts model training from on a single GPU card to on a GPU clusters due to limited GPU memory and time-to-solution requirements. High-speed low-latency RDMA-capable network fabrics like Infiniband and RoCE play an important role on coping with enoumous data exchanged during training. DL frameworks are built upon these fabrics with various APIs including IPoIB, MPI and RDMA Verbs. Tradeoffs are made between performance and usability when adapting DL frameworks onto RDMA-capable networks, which may result in high-performance yet hard-to-maintain and hard-to-merge code if improper design choices are made. This paper presents our approach to adapt MXNet, a modular versatile DL framework onto RDMA-capable networks. Dividing the training process on MXN et into P2P communication and A11Reduce commnunication, we add incremental optimizations on its message passing code. Experiments show that our approach exhibits near-linear speedups, whose parallel efficiency reaches 96% compared to 53% of the original IPoIB version when scaling to 100 GPU cards. In contrast to other MPI-based porting approach, our modifications are limited within MXNet's Parameter Server module, which is transparent for upper-layer operations, thus making no sacrifice on features like auto recovery and user-controlled consistency view.
Jianwen Wei, Yichao Wang 0001, Minhua Wen, Simon See, James Lin 0001
ICPADS6
2018 Evaluating the SW26010 many-core processor with a micro-benchmark suite for performance optimizations
James Lin 0001, Zhigeng Xu, Linjin Cai, Akira Nukada, Satoshi Matsuoka
Parallel Comput.1
2017 Optimizations of Two Compute-Bound Scientific Kernels on the SW26010 Many-Core Processor
abstract
The home-grown SW26010 many-core processor enabled the production of China's first independently developed number-one ranked supercomputer - the Sunway TaihuLight. The design of the limited off-chip memory bandwidth, however, renders the SW26010 a highly memory-bound processor. To compensate for this limitation, the processor was designed with a unique hardware feature, "Register Level Communication" (RLC), to share register data among its 8 × 8 computing processing elements (CPEs) via a 2D onchip network. Such a radical architecture has sparked global researchers' concerns regarding the programming challenges this may cause. To address these concerns, we adopted two compute-bound scientific kernels as benchmarks to identify the potential programming challenges. The first kernel is doubleprecision general matrix-multiplication (DGEMM). An RLCfriendly algorithm was designed for this kernel to reuse the data that already reside in the registers of 64 CPEs. This novel optimization enables the kernel to achieve up to 88.7% efficiency in one core group of the SW26010. This paper reveals, for the first time, the details of how the highly efficient DGEMM is implemented on the home-grown processor. The second kernel that we used is N-body. Due to the inefficient hardware support for transcendental operations on the SW26010, we replaced the reciprocal square root (rsqrt) instruction of N-body with a software routine to tackle the problem. Based on the programming challenges identified through these two optimized kernels, we proposed a three-level programming guideline for the SW26010. The paper concludes with our crucial finding that the critical step towards bridging the ninja performance gap on the SW26010 is to design an RLC-friendly algorithm to increase arithmetic intensity.
James Lin 0001, Zhigeng Xu, Akira Nukada, Naoya Maruyama, Satoshi Matsuoka
ICPP1
2016 Performance and Portability Studies with OpenACC Accelerated Version of GTC-P
abstract
Accelerator-based heterogeneous computing is of paramount importance to High Performance Computing. The increasing complexity of the cluster architectures requires more generic, high-level programming models. OpenACC is a directive-based parallel programming model, which provides performance on and portability across a wide variety of platforms, including GPU, multicore CPU, and many-core processors. GTC-P is a discovery-science-capable real-world application code based on the Particle-In-Cell (PIC) algorithm that is well-established in the HPC area. Several native versions of GTC-P have been developed for supercomputers on TOP500 with different architectures, including Titan, Mira, etc. Motivated by the state-of-art portability, we implemented the first OpenACC version of GTC-P and evaluated its performance portability across NVIDIA GPUs, Intel x86 and OpenPOWER CPUs. In this paper, we also proposed two key optimization methods for OpenACC implementation of PIC algorithm on multicore CPU and GPU including removing atomic operation and taking advantage of shared memory. OpenACC shows both impressive productivity and performance in a perspective of portability and scalability. The OpenACC version achieves more than 90% performance compared with the native versions with only about 300 LOC.
Yueming Wei, Yichao Wang 0001, Linjin Cai, William Tang 0002, Bei Wang 0002, Stéphane Ethier, Simon See, James Lin 0001
PDCAT8
2015 Modeling Gather and Scatter with Hardware Performance Counters for Xeon Phi
abstract
Intel Initial Many-Core Instructions (IMCI) for Xeon Phi introduces hardware-implemented Gather and Scatter (G/S) load/store contents of SIMD registers from/to non-contiguous memory locations. However, they can be one of key performance bottlenecks for Xeon Phi. Modelling G/S can provide insights to the performance on Xeon Phi, however, the existing solution needs a hand-written assembly implementation. Therefore, we modeled G/S with hardware performance counters which can be profiled by the tools like PAPI. We profiled Address Generation Interlock (AGI) events as the number of G/S, estimated the average latency of G/S with VPU_DATA_READ, and combined them to model the total latencies of G/S. We applied our model to the 3D 7-point stencil and the result showed G/S spent nearly 40% of total kernel time. We also validated the model by implementing a G/S- free version with intrinsics. The contribution of the work is a performance model for G/S built with hardware counters. We believe the model can be generally applicable to CPU as well.
James Lin 0001, Akira Nukada, Satoshi Matsuoka
CCGRID1