Kaicheng Guo

dblp:308/5921 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0003-0493-1147ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FractalGPU: Fair and Elastic GPU Sharing for General-Purpose Computing
Kaicheng Guo, Lingyun Yang, Wenda Tang, Pengwei Du, Qian Da, Zhengwei Qi
ICDCS2
2026 gPooling: An Elastic GPU Resource Management Framework for On-Demand Virtualization in Shared Accelerator Clusters
Kaicheng Guo, Chen Chen 0067, Yun Wang 0039, Pengwei Du, Zhengwei Qi, Haibing Guan
IEEE Trans. Parallel Distributed Syst.1
2025 Design and Operation of Elastic GPU-Pooling on Campus
Kaicheng Guo, Yun Wang 0039, Semakin Anton, Tovmachenko Dmitry, Jiajie Sheng, Jianwen Wei, James Lin 0001, Zhengwei Qi, Haibing Guan
Euro-Par (1)1
2025 Exploring Efficient Hardware Accelerator for Learning-Based Image Compression
abstract
Recently, learning-based image compression (LIC) methods have surpassed manually designed approaches in both compression quality and bitrate. However, increasing computational demands and insufficient optimizations in codec performance have hindered the advancement of LIC acceleration. Most researches focus on optimizing specific components, often neglecting the sources of underutilization during the execution of LIC models. Generally, efficient LIC acceleration encounters three primary challenges: 1) extra overheads introduced by individual optimizations; 2) load and computation imbalances in small kernels; and 3) mismatches between hardware configurations and the LIC models. To address these challenges, we propose a framework named extensive accelerator for LIC (X-LIC) for efficiently exploring the design space under constrained resources. First, we quantitatively characterize a representative LIC model, including its latency, computation size, and temporal utilization across various accelerators. We design a hardware-optimized quantization method to compensate for the lack of LIC-oriented research, particularly regarding data precision, distortion, and resource consumption. Additionally, we propose a parameterized LIC accelerator architecture that integrates seamlessly with existing loop optimization models and supports various LIC operators. Two optimization schemes are proposed for redundant computation in transposed convolution and load and computation imbalance in small kernels. Experimental results show that our framework demonstrates significant flexibility across a broad design space, achieving an average of 78%–95% of the theoretical peak performance and up to 688.2/759.1 GOP/s en/de-coder performance with INT8 precision. As a result, the en/de-coder performance can reach up to 33/36 FPS in 720P resolution. An FPGA demo of X-LIC is available athttps://github.com/sjtu-tcloud/X-LIC.
Chen Chen 0067, Kaicheng Guo, Xingzi Yu, Weidong Qiu, Zhengwei Qi, Haibing Guan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2023 Optimum: Runtime optimization for multiple mixed model deployment deep learning inference
Kaicheng Guo, Yixiao Xu, Zhengwei Qi, Haibing Guan
J. Syst. Archit.1
2023 DVHN: A Deep Hashing Framework for Large-Scale Vehicle Re-Identification
abstract
Vehicle re-identification is a pervasive technology in real-world intelligence transportation systems. Conventional methods generally perform re-identification tasks by representing vehicle images as real-valued feature vectors and then ranking the gallery images by computing the corresponding Euclidean distances. Despite achieving remarkable retrieval accuracy, these high-dimensional real-valued feature vectors are not tailored for fast indexing and matching and require tremendous memory and computation when the gallery set is large, making them inapplicable in a large-scale real-world retrieval setting. In light of this limitation, in this paper, we make the very first attempt to develop an efficient vehicle re-identification system (DVHN) for real-world large-scale retrieval tasks with deep hashing learning. It could substantially reduce memory usage and enhances retrieval efficiency while maintaining retrieval accuracy. Concretely, DVHN directly learns discrete compact binary hashing codes for each image by jointly optimizing the feature learning network and the hash code generating module. Specifically, we directly constrain the output from the convolutional neural network to be discrete binary codes and ensure the learned binary codes are optimal for classification. To optimize the deep discrete hashing framework, we further propose an alternating minimization method for learning binary similarity-preserved hashing codes. Extensive experiments on two widely-studied vehicle re-identification datasets- VehicleID and VeRi- have demonstrated the superiority of our method against the state-of-the-art deep hash methods. DVHN of 2048 bits can achieve 13.94% and 10.21% accuracy improvement in terms of mAP and Rank@1 for VehicleID (800) dataset. For VeRi, we achieve 35.45% and 32.72% performance gains for Rank@1 and mAP, respectively.
Yongbiao Chen, Fangxin Liu, Kaicheng Guo, Zhengwei Qi
IEEE Trans. Intell. Transp. Syst.5
2022 Supervised Contrastive Vehicle Quantization for Efficient Vehicle Retrieval
abstract
This paper considers large-scale efficient vehicle re-identification (Vehicle ReID). Existing works adopting deep hashing techniques function by projecting vehicle images into compact binary codes in the Hamming space. Since Hamming distance is less distinct, a considerable amount of discriminative information will be lost, leading to degraded retrieval performances. Inspired by the recent advancements in contrastive learning, we put forward the very first product quantization based framework for large-scale efficient vehicle re-identification: Supervised Contrastive Vehicle Quantization (SCVQ). Specifically, we integrate the product quantization process into deep supervised learning by designing a differentiable quantization network. In addition, we propose a novel supervised cross-quantized contrastive quantization (SCQC) loss for similarity-preserving learning, which is tailored for the asymmetric retrieval in the product quantization process. Comprehensive experiments on two public benchmarks have evidenced the superiority of our framework against the state-of-the-arts. Our work is open-sourced at https://github.com/chrisbyd/ContrastiveVehicleQuant
Yongbiao Chen, Kaicheng Guo, Fangxin Liu, Zhengwei Qi
ICMR2