Jingchang Qin

dblp:352/6520 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0009-0009-0090-5030ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2025 rInfer: A Generic and High-Performance Framework for Remote Inference with Heterogeneous Accelerators
abstract
Inference applications leverage various heterogeneous accelerators, including GPUs, TPUs, and FPGAs, to achieve remarkable performance. Remote inference, which enables applications and accelerators to reside on different nodes, can effectively address low hardware resource utilization and enhance flexibility in resource allocation. However, existing APIremoting solutions face compatibility issues, as they demand extensive development efforts to support diverse heterogeneous accelerators and adapt to rapidly evolving device runtimes. On the other hand, current inference service systems demand cumbersome cross-server configuration and suffer from suboptimal data transmission performance, overlooking data copy overhead in the network stack and not leveraging high-performance RDMA. In this paper, we present rInfer, a generic and highperformance framework for remote inference with heterogeneous accelerators. The key idea of rInfer is to abstract the entire remote inference process and optimize network transmission. On the client node, rInfer offers users an inference-oriented rInfer Device and an ease-of-use rInfer programming model, shielding users from the complexities of server configuration and network communication. On the server node, rInfer integrates with diverse inference frameworks, such as TensorRT and Torch, ensuring high compatibility with various accelerators. Furthermore, we optimize network data transmission by eliminating the need for serialization and deserialization in Protobuf to minimize data copying and adopt RDMA to enhance data transfer efficiency. Experimental results demonstrate that rInfer achieves performance nearly equivalent to local execution for the BERT model, with only a minor difference of 1.5 %. Compared to TorchServe, rInfer-TCP achieves a 51.6 % reduction in execution time for the ResNet models and consumes 37.3% less CPU and 61.2% less memory bandwidth on the server node.
Zhen Jin 0008, Yiquan Chen, Yin Du, Keyao Zhang, Jiexiong Xu, Wenhai Lin, Jingchang Qin, Kanghua Fang, Wenzhi Chen
CCGrid8
2024 CINDA: Don't Ignore Instructions When Cloning Memory Access Behavior
abstract
Existing workload cloning methods suffer from low accuracy as they primarily focus on data access patterns and ignore instruction access. This limitation reduces the accuracy of shared L2 cache design exploration and impedes processor designers from optimizing Icache and ITLB designs. In this paper, we propose CINDA, a novel workload cloning technique that can Capture both INstruction and DAta access patterns of applications. In particular, CINDA separates the instruction and data traces of applications to generate proxy instruction and proxy data traces, subsequently merging them. The results show that CINDA can accurately replicate memory access behavior with 99.1%, 99.9%, and 96.2% accuracy in replicating L1 Icache, ITLB and L2 cache performance, respectively. Furthermore, CINDA outperforms the state-of-the-art methods by reducing 7.7% L2 cache miss error.
Wenhai Lin, Yiquan Chen, Jiexiong Xu, Zhen Jin 0008, Peiyu Liu 0003, Shishun Cai, Yuzhong Zhang, Jingchang Qin, Yiquan Lin, Wenzhi Chen
CCGrid8
2024 BlueJay: A Platform to Quantifying the Impact of Memory Latency on Datacenter Application Performance
abstract
Understanding the impact of memory latency on datacenter application performance can provide decision support to memory subsystem designers. Currently, various methods are available to quantify this impact, including cycle-accurate simulators, memory-level parallelism models, and software delay injection techniques. However, these methods suffer from several limitations, such as slow simulation speed, inaccuracy, and insufficient compatibility that requires application modification.This paper proposes BlueJay, a novel platform to quantify the impact of memory latency on the end-to-end performance of datacenter applications, avoiding slow simulation and providing high accuracy and compatibility. The key idea of BlueJay is to control the consumed memory bandwidth and read/write ratio, thereby manipulating memory latency to achieve quantification. Experiment shows that BlueJay provides accurate quantification with an average error of 3.04%. In addition, we built regression models for five applications deployed at scale in Alibaba data centers. The results reveal that a 10 ns increase in memory latency results in a performance decrease of 2.61%-3.31% for enterprise Java applications and databases, while the elastic block storage service experiences a more modest performance decrease of 0.73%-0.92%.
Jingchang Qin, Yiquan Chen, Shishun Cai, Wenhai Lin, Jiexiong Xu, Zhen Jin 0008, Lifa Cao, Yuzhong Zhang, Wenzhi Chen
CCGrid1
2024 PARS: A Pattern-Aware Spatial Data Prefetcher Supporting Multiple Region Sizes
abstract
Hardware data prefetching is a well-studied technique to bridge the processor-memory performance gap. Bit-pattern-based prefetchers are one of the most promising spatial data prefetchers that achieve substantial performance gains. In bit-pattern-based prefetchers, the region size is a crucial parameter, which denotes the memory size that can be recorded by a pattern or prefetched by a prediction. However, existing bit-pattern-based prefetchers only support one fixed region size. Our experiment shows that the fixed region size cannot meet the requirements for numerous applications and leads to suboptimal performance and high hardware overhead. In this article, we propose PARS, a pattern-aware spatial data prefetcher supporting multiple region sizes. The key idea of PARS is that it supports multiple region sizes, enabling it to simultaneously enhance application performance while reducing the hardware overhead. Moreover, PARS supports dynamically switching appropriate region sizes for different patterns through an adaptive RS-switching mechanism. We evaluated PARS on numerous workloads and results show that PARS provides an average performance improvement of 40.6% over a baseline with no data prefetchers and outperforms the two state-of-the-art prefetchers Bingo by 2.1% (up to 24.4%) and Pythia by 3.9% (up to 111.2%) in the single-core system. In the four-core system, PARS outperforms Bingo by 5.0% (up to 66.0%) and Pythia by 5.4% (up to 177.9%).
Yiquan Lin, Wenhai Lin, Jiexiong Xu, Yiquan Chen, Zhen Jin 0008, Jingchang Qin, Shishun Cai, Yuzhong Zhang, Zonghui Wang, Wenzhi Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2023 JACO: JAva Code Layout Optimizer Enabling Continuous Optimization without Pausing Application Services
abstract
Many Java applications in data centers suffer from severe processor pipeline frontend bottlenecks, which can be mitigated by profile-guided code layout optimizations (PGCLO). To maximize optimization opportunities, state-of-the-art PGCLO solutions adopt continuous optimization to ensure that the code layout consistently matches ever-changing application control flow characteristics. However, existing continuous optimizations inevitably pause the application to execute the new code completely, which leads to high response latency and significantly deteriorates user experience.In this paper, we propose JACO, a novel profile-guided Java code layout optimizer, enabling continuous optimization without pausing application services. The key idea of JACO is to enable the execution of both the old and new code simultaneously rather than completely switching to the new code. In particular, JACO is composed of three components: (1) A lightweight profiler captures the control flow information of the application and then generates an optimized function order. (2) A control flow switcher generates new code based on optimized function order and switches the application to execute the new code without pausing the application services. (3) A selective code reclaimer only frees the memory occupied by the inactive old code. We evaluated JACO on both open-source applications and real-world applications from a world-leading company. JACO achieved up to a 16.36% performance improvement for real-world applications. The state-of-the-art approach introduces up to 37.93x latency overhead that will interrupt application services, while JACO only introduces a negligible 7% latency overhead.
Wenhai Lin, Jingchang Qin, Yiquan Chen, Zhen Jin 0008, Jiexiong Xu, Yuzhong Zhang, Shishun Cai, Lirong Fu, Wenzhi Chen
CLUSTER2
2023 How IoT Re-using Threatens Your Sensitive Data: Exploring the User-Data Disposal in Used IoT Devices
abstract
With the rapid technology evolution of the Internet of Things (IoT) and increasing user needs, IoT device re-using becomes more and more common nowadays. For instance, more than 300,000 used IoT devices are selling on Craigslist. During IoT re-using, sensitive data such as credentials and biometrics residing in these devices may face the risk of leakage if a user fails properly dispose of the data. Thus, a critical security concern is raised: do (or can) users properly dispose of the sensitive data in used IoT? To the best of our knowledge, it is still an unexplored problem that desires a systematic study.In this paper, we perform the first in-depth investigation on the user-data disposal of used IoT devices. Our investigation integrates multiple research methods to explore the status quo and the root causes of the user-data leakages with used IoT devices. First, we conduct a user study to investigate the user awareness and understanding of data disposal. Then, we conduct a large-scale analysis on 4,749 IoT firmware images to investigate user-data collection. Finally, we conduct a comprehensive empirical evaluation on 33 IoT devices to investigate the effectiveness of existing data disposal methods.Through the systematical investigation, we discover that IoT devices collect more sensitive data than users expect. Specifically, we detect 121,984 sensitive data collections in the tested firmware. Moreover, users usually do not or even cannot properly dispose of the sensitive data. Worse, due to the inherent characteristics of storage chips, 13.2% of the investigated firmware perform "shallow" deletion, which may allow adversaries to obtain sensitive data after data disposal. Given the large-scale IoT re-using, such leakage would cause a broad impact. We have reported our findings to world-leading companies. We hope our findings raise awareness of the failures of user-data disposal with IoT devices and promote the protection of users’ sensitive data in IoT devices.
Peiyu Liu 0003, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang 0002, Jingchang Qin, Wenhai Wang, Wenzhi Chen
SP6