Jiajian Zhang

dblp:227/0259 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-2903-9411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LTA-Gait: In-Network Temporal Activation of Large Vision Models for Gait Recognition
abstract
While gait recognition based on Large Vision Models (LVMs) benefits from powerful spatial representations, existing paradigms often reduce video sequences to unordered image sets, inherently neglecting the inter-frame temporal causality. To bridge the modal gap between static image perception and dynamic video understanding, we propose the Layer-wise Spatio-Temporal Activation (LTA) framework and the LTA-Gait model. Leveraging the linear complexity of State Space Models (SSMs), our method alleviates the efficiency challenge of introducing temporal modeling into deep LVMs, enabling full-level dense temporal adaptation in a practical manner. Addressing the challenges of recursive noise sensitivity and spatio-temporal granularity misalignment specific to gait video, we design an Anatomy-Aware Temporal Adapter (AATA). By introducing an input-output synergistic dual-gating mechanism, we apply a binary mask-based input constraint to suppress background noise before temporal recursion, while using adaptive soft gating at the output to stabilize feature injection during fine-tuning. Furthermore, we construct a progressive temporal receptive field strategy to better match temporal modeling granularity with the hierarchical spatial features of the LVM. Extensive experiments demonstrate that LTA-Gait achieves State-of-the-Art (SOTA) performance on benchmarks such as CCPG, CCGR, and CASIA-B*. Notably, significant performance gains in challenging clothing-change scenarios validate the core value of robust ordered modeling in extracting intrinsic gait dynamics. All implementation code is fully available at our project repository: https://github.com/zjucgl/LTA-Gait.
Genlang Chen, Chengcheng Jia, Jiajian Zhang
ICMR4
2025 WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU Communications
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Chaoyi Pang
USENIX ATC1
2025 An attention-based framework for integrating WSI and genomic data in cancer survival prediction
Genlang Chen, Sixuan Sui, Jiajian Zhang, Ping Cai
J. Biomed. Informatics3
2025 AlignMalloc: Warp-Aware Memory Rearrangement Aligned With UVM Prefetching for Large-Scale GPU Dynamic Allocations
abstract
As parallel computing tasks rapidly expand in both complexity and scale, the need for efficient GPU dynamic memory allocation becomes increasingly important. While progress has been made in developing dynamic allocators for substantial applications, their real-world applicability is still limited due to inefficient memory access behaviors. This paper introduces AlignMalloc, a novel memory management system that aligns with the Unified Virtual Memory (UVM) prefetching strategy, significantly enhancing both memory allocation and access performance in large-scale dynamic allocation scenarios. We analyze the fundamental inefficiencies in UVM access and first reveal the mismatch between memory access and UVM prefetching methods. To resolve this issue, AlignMalloc implements a warp-aware memory rearrangement strategy that exploits the regularity of warps to align with the UVM's static prefetching setup. Additionally, AlignMalloc introduces an OR tree-based structure within a host-co-managed framework to further optimize dynamic allocation. Comprehensive experiments demonstrate that AlignMalloc substantially outperforms current state-of-the-art systems, achieving up to$2.7 \times$improvement in dynamic allocation and$2.3 \times$in memory access. Additionally, eight real-world applications with diverse memory access patterns exhibit consistent performance enhancements, with average speedups$1.5 \times$.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Eng Gee Lim, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.1
2024 SyncMalloc: A Synchronized Host-Device Co-Management System for GPU Dynamic Memory Allocation across All Scales
abstract
Dynamic memory allocation on GPUs, increasingly crucial for applications with dynamic computational patterns, encounters significant challenges due to the complex calculations with intricate branches and substantial memory resources consumed by metadata from massive thread allocations. Despite the current research, there is a lack of a scalable and flexible solution that effectively manages dynamic memory allocation while minimizing memory usage on GPUs. This paper introduces SyncMalloc, a synchronized Host-Device Co-Management system that is specifically designed to adeptly handle dynamic memory allocations of diverse magnitudes. Through the integration of pipelining and producer-consumer mechanisms, SyncMalloc effectively reduces communication overhead and resolves architectural mismatches, further enhancing its capability through synergistic integration with CUDA’s unified memory to facilitate oversubscription. Moreover, SyncMalloc advances slab-based memory management to enhance the efficiency of small allocations, reducing conflict probabilities and overhead in high-activity scenarios. Finally, we present a comprehensive performance evaluation, expanding benchmarks and measurement dimensions to reflect the performance of real-world applications more accurately. The experimental results demonstrate the effectiveness of SyncMalloc in supporting dynamic GPU allocations scaled from 4B to 200GB from multiple perspectives. Our source code is available at https://github.com/jjZhang94/SyncMalloc.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Genlang Chen, Qiufeng Wang 0001
ICPP1
2023 A deep image segmentation-based method for stitching ancient-book images without an overlapping region
abstract
Abstract With continuous advancements in ancient‐book digitization and preservation research, the problems with the stitching of ancient‐book images have become increasingly prominent, as traditional feature‐mapping‐based methods cannot satisfactorily stitch non‐overlapping images. To realize the accurate stitching of the left and right pages of ancient‐book images, this paper proposes a method for ancient‐book image stitching to meet the requirements of their digitization in back‐wrapped binding and other binding forms. First, a dataset of the black text frames from ancient‐book images was established and then used to train a VGG16‐UNet network for the extraction of black text frames. Then, the Douglas–Peucker algorithm was used to fit the black text frames and filter outliers. Finally, a sliding matching algorithm based on the position information of black text frames was proposed for the rectification of misalignments. The results showed that the method achieved a satisfying stitching effect and had good robustness.
Genlang Chen, Guanghui Song, Jiajian Zhang
IET Image Process.5
2022 CRAC: An automatic assistant compiler of checkpoint/restart for OpenCL program
abstract
Summary Nowadays, people use multiple devices to meet the growing requirement for computing. With the application of multicard computing, fault tolerance, load balance, and resource sharing have been the hot issues and the checkpoint/restart (CPR) mechanism is critical in a preemptive system. This article proposes a CPR framework including the automatic compiler (CRAC) to achieve a feasible CPR system, especially for graphics processing unit applications on heterogeneous devices in OpenCL programs. By offering the positions of the CPR in source code, CRAC inserts primitives into programs and invokes the runtime support modules for final results. A comprehensive example and experiments have demonstrated the feasibility and effectiveness of proposed framework.
Genlang Chen, Jiajian Zhang, Zufang Zhu, Hai Jiang 0003, Chaoyi Pang
Concurr. Comput. Pract. Exp.2
2021 Video Text Tracking With a Spatio-Temporal Complementary Model
abstract
Text tracking is to track multiple texts in a video, and construct a trajectory for each text. Existing methods tackle this task by utilizing the tracking-by-detection framework, i.e., detecting the text instances in each frame and associating the corresponding text instances in consecutive frames. We argue that the tracking accuracy of this paradigm is severely limited in more complex scenarios, e.g., owing to motion blur, etc., the missed detection of text instances causes the break of the text trajectory. In addition, different text instances with similar appearance are easily confused, leading to the incorrect association of the text instances. To this end, a novel spatio-temporal complementary text tracking model is proposed in this paper. We leverage a Siamese Complementary Module to fully exploit the continuity characteristic of the text instances in the temporal dimension, which effectively alleviates the missed detection of the text instances, and hence ensures the completeness of each text trajectory. We further integrate the semantic cues and the visual cues of the text instance into a unified representation via a text similarity learning network, which supplies a high discriminative power in the presence of text instances with similar appearance, and thus avoids the mis-association between them. Our method achieves state-of-the-art performance on several public benchmarks. The source code is available at https://github.com/lsabrinax/VideoTextSCM.
Yuzhe Gao, Jiajian Zhang, Yu Zhou 0016, Jing Wang 0221, Shenggao Zhu, Xiang Bai
IEEE Trans. Image Process.3
2021 CRState: checkpoint/restart of OpenCL program for in-kernel applications
Genlang Chen, Jiajian Zhang, Zufang Zhu, Qiangqiang Jiang, Hai Jiang 0003, Chaoyi Pang
J. Supercomput.2
2019 CRState: In-Kernel Checkpoint/Restart of OpenCL Program Execution on GPU
abstract
Checkpoint/restart is an important mechanism to achieve fault tolerance, load balancing and resources sharing in a preemptive system. As Graphics Processing Unit (GPU) becomes quite popular in high performance computing as well as OpenCL programs are portable across various CPUs and GPUs, checkpoint/restart of OpenCL programs on GPUs is in demand. However, due to the intricacy of computation states inside GPUs, there is no effective checkpoint/restart scheme for heterogeneous devices now. This paper proposes a feasible system, CRState, to achieve checkpoint/restart in GPU kernels. With the assistant of a pre-compiler, the primitives are inserted into programs. In run-time, the computation state existing in the underlying hardware is concretized and reconstructed at application level and is ported to heterogeneous devices. Comprehensive experiments have been conducted to demonstrate CRState's feasibility and effectiveness. The experimental results also indicate that CRState has the potential to reschedule resources and balance workload across heterogeneous devices.
Genlang Chen, Jiajian Zhang, Qiuru Lin, Hai Jiang 0003, Chaoyi Pang
ICPADS2
2019 Application of deep learning fusion algorithm in natural language processing in emotional semantic analysis
abstract
Summary With the development of network technology, people are facing more and more massive information. How to extract emotional information in massive information rapidly has received more and more attention from people. This paper introduces the principle and structure of the traditional emotional model. Different personality, emotional states, and external stimuli will have different effects on emotional semantic analysis. In addition, this paper has proposed emotional semantic analysis method based on wake‐sleep and SVM method. The model starts from the description and calculation of the dynamic characteristics of emotions and more fully predicts the process characteristics that describe the evolution of emotions. Search and category browsing allows users to quickly access these information points. In addition, this paper provides a deep learning fusion algorithm in emotional semantic analysis, introduces its reference implementation and related key technologies, and supports business intelligence to a certain extent, and it has a strong application prospect on the network data information.
Yunlu Gong, Nannan Lu, Jiajian Zhang
Concurr. Comput. Pract. Exp.3
2019 Generic attribute revocation systems for attribute-based encryption in cloud storage
abstract
Attribute-based encryption (ABE) has been a preferred encryption technology to solve the problems of data protection and access control, especially when the cloud storage is provided by third-party service providers. ABE can put data access under control at each data item level. However, ABE schemes have practical limitations on dynamic attribute revocation. We propose a generic attribute revocation system for ABE with user privacy protection. The attribute revocation ABE (AR-ABE) system can work with any type of ABE scheme to dynamically revoke any number of attributes.
Genlang Chen, Zhiqian Xu 0001, Jiajian Zhang, Guojun Wang 0001, Hai Jiang 0003, Miaoqing Huang
Frontiers Inf. Technol. Electron. Eng.3