Jin Zhang 0003

dblp:43/6657-3 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-9086-1178ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 JSQKV: Joint Sparsification and Quantization for KV-Cache Compression and Decode Acceleration
Xiaoli Gong, Huayou Su, Qingxia Chen, Jin Zhang 0003
APPT6
2026 Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Xiaoli Gong, Jin Zhang 0003
Signal Process.6
2026 Fine-Grained Audio-Visual Event Localization
abstract
Audio-visual event localization (AVEL) aims to recognize events in videos by associating audio-visual information. However, events involved in existing AVEL tasks are usually coarse-grained events. Actually, finer-grained events are sometimes necessary to be distinguished, especially in certain expert-level applications or rich-content-generation studies. However, this is challenging because they are more difficult to detect or distinguish compared with coarse-grained events. To better address this problem, we discuss a new setting of fine-grained AVEL from dataset to method. First, we constructed the first fine-grained audio-visual event dataset, which is called IT-AVE, relying on videos of playing musical instruments, containing 13k video clips and over 52k audio-visual events. All events are labeled from professional music practitioners, and the event categories are all derived from playing techniques, which are fine-grained with little interclass variation. Next, we designed a new fine-grained event localization method, spatial-temporal video event detector (SVED), which focuses on the challenges that fine-grained events are more imperceptible and prone to be disturbed. Finally, we conduct extensive experiments based on the proposed IT-AVE dataset versus fine-grained versions of two existing related datasets, including UnAV-22 derived from UnAV-100 and FineAction-AV derived from FineAction. Experimental results demonstrate the effectiveness of our method. We hope that this work will contribute to the exploration of an integrated understanding of audio-visual videos.
Baoyu Fan, Lu Liu 0009, Xiaochuan Li 0001, Jin Zhang 0003
IEEE Trans. Neural Networks Learn. Syst.6
2025 So Far Yet So Near: Time Series Data Augmentation with Exploring non-Semantic Boundaries based on Reinforcement Learning
abstract
Data augmentation effectively expands feature distribution in time series classification, enhancing downstream task performance. However, existing techniques often fail to maintain semantic consistency between augmented and original time series data, causing label noise and thereby degrading downstream task performance. We argue that data augmentation should preserve time series semantic consistency and expand the non-semantic information space. In this paper, we reformulate data augmentation as a semantic path planning problem between original data and augmented data, modeled as a Markov Decision Process (MDP). We propose a reinforcement learning-based algorithm (RL) named FreqSYN, where the action space is defined by a set of learnable Gaussian kernels that perturbs the frequency domain of the original data to generate augmented samples. The confidence coefficients of augmented data in semantically relevant classification tasks are used as a reward to iteratively refine the FreqSYN. Our method is validated across four datasets, achieving state-of-the-art performance, with a 2% improvement in F1 score over the SimPSI method. The code and models are available at https://github.com/NKU-EmbeddedSystem/FreqSYN.
Haoran Li 0014, Jiarong Kang, Xun Jiang 0001, Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Andrzej Cichocki
ICASSP6
2025 Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model
abstract
Electroencephalography (EEG) is a time-series signal containing semantic information that can be used to determine human brain activities. Artifacts within EEG data can interfere with the intrinsic distribution of this semantic information, so removing artifacts is crucial for improving EEG analysis performance on downstream tasks. In this paper, we redefine the efficacy of the artifact removal model by evaluating the performance of the noisy EEG data in downstream tasks before and after artifact removal. Currently, most artifact removal models fail to ensure semantic consistency, rendering them ineffective. To solve it, we propose an artifact removal model based on the 1-dimensional diffusion model utilizing the U-Net, referred to as Essentia. Moreover, we find that the skip-connection layer in U-Net contains mid-to-high-frequency information that interferes with the semantic representation. We introduce a semantic guidance module (SGM) that leverages contrastive learning to generate semantic distribution weights, boosting semantic representation. We evaluate Essentia on three datasets with six solutions. The accuracy of downstream tasks from the denoised EEG data increased by 4% compared with the DeepSeparetor. The code and models are available at https://github.com/NKU-EmbeddedSystem/Essentia.
Haoran Li 0014, Xiaoli Gong, Jin Zhang 0003, Tingjuan Lu, Zhe Sun 0009, Andrzej Cichocki
ICASSP5
2024 Prism: Decomposing Program Semantics for Code Clone Detection through Compilation
abstract
Code clone detection (CCD) is of critical importance in software engineering, while semantic similarity is a key evaluation factor for CCD. The embedding technique, which represents an object using a numerical vector, is utilized to generate code representations, where code snippets with similar semantics (clone pairs) should have similar vectors. However, due to the diversity and flexibility of high-level program languages, the code representation of clone pairs may be inconsistent. Assembly code provides the program execution trace and can normalize the diversity of high-level languages in terms of the program behavior semantics. After revisiting the assembly language, we find that different assembly codes can align with the computational logic and memory access patterns of cloned pairs. Therefore, the use of multiple assembly languages can capture the behavior semantics to enhance the understanding of programs. Thus, we propose Prism, a new method for code clone detection fusing behavior semantics from multiple architecture assembly code, which directly captures multilingual domains' syntax and semantic information. Additionally, we introduce a multi-feature fusion strategy that leverages global information interaction to expand the representation space. This fusion process allows us to capture the complementary information from each feature and leverage the relationships between them to create a more expressive representation of the code. After testing the OJClone dataset, the Prism model exhibited exceptional performance with precision and recall scores of 0.999 and 0.999, respectively.
Haoran Li 0014, Siqian Wang, Weihong Quan, Xiaoli Gong, Huayou Su, Jin Zhang 0003
ICSE6
2024 SyncIntellects: Orchestrating LLM Inference with Progressive Prediction and QoS-Friendly Control
abstract
Large Language Models (LLMs) have shown impressive capabilities, especially in the realm of Human-Machine Chat Systems. Nevertheless, these models entail significant computational expenses, particularly when generating tokens. As a remedy to enhance system throughput and hardware utilization, batch scheduling is commonly adopted. This method involves initiating a batch of inference requests concurrently and then waiting for their completion. A significant challenge encountered with task-batching is the need to group requests with similar response lengths. However, accurately predicting response length proves to be a daunting task, and the inherent variability in response length leads to suboptimal resource utilization.In this paper, we introduce SyncIntellects, a framework designed to orchestrate Large Language Model (LLM) Inference with fine-grained response length prediction and Quality of Service (QoS)-Friendly length control. Specifically, SyncIntellects enhances response length prediction by leveraging embedding information during token generation through a transformer-based model. Subsequently, a dynamic response length controller based on Prompt Engineering techniques is employed to ensure alignment of response lengths without compromising the QoS of the responses. We have implemented SyncIntellects and seamlessly integrated it with a chatbot engine based on the llama2 7B model. We conduct comprehensive experiments on an NVIDIA A100-based testbed, and the results demonstrate a significant reduction in latency by 17.76% on average, along with an increase in throughput by 9.34%.
Xue Lin 0006, Peining Yue, Haoran Li 0014, Jin Zhang 0003, Baoyu Fan, Huayou Su, Xiaoli Gong
IWQoS5
2024 OneGraph: a cross-architecture framework for large-scale graph computing on GPUs based on oneAPI
Jiaxun Han, Xiaoli Gong, Gang Wang 0001, Jin Zhang 0003, Xuqiang Wang
CCF Trans. High Perform. Comput.8
2024 JiuJITsu: Removing Gadgets with Safe Register Allocation for JIT Code Generation
abstract
Code-reuse attacks have the capability to craft malicious instructions from small code fragments, commonly referred to as “gadgets.” These gadgets are generated by JIT (Just-In-Time) engines as integral components of native instructions, with the flexibility to be embedded in various fields, including Displacement . In this article, we introduce a novel approach for potential gadget insertion, achieved through the manipulation of ModR/M and SIB bytes via JavaScript code. This manipulation influences a JIT engine’s register allocation and code generation algorithms. These newly generated gadgets do not rely on constants and thus evade existing constant blinding schemes. Furthermore, they can be combined with 1-byte constants, a combination that proves to be challenging to defend against using conventional constant blinding techniques. To showcase the feasibility of our approach, we provide proof-of-concept (POC) code for three distinct types of gadgets. Our research underscores the potential for attackers to exploit ModR/M and SIB bytes within JIT-generated native instructions. In response, we propose a practical defense mechanism to mitigate such attacks. We introduce JiuJITsu , a security-enhanced register allocation scheme designed to prevent harmful register assignments during the JIT code generation phase, thereby thwarting the generation of these malicious gadgets. We conduct a comprehensive analysis of JiuJITsu ’s effectiveness in defending against code-reuse attacks. Our findings demonstrate that it incurs a runtime overhead of under 1% when evaluated using JetStream2 benchmarks and real-world websites.
Zhang Jiang, Ying Chen 0034, Xiaoli Gong, Jin Zhang 0003, Wenwen Wang 0001, Pen-Chung Yew
ACM Trans. Archit. Code Optim.4
2024 Hybrid-Memcached: A Novel Approach for Memcached Persistence Optimization With Hybrid Memory
abstract
Memcached is a widely adopted, high-performance, in-memory key-value object caching system utilized in data centers. Nonetheless, its data is stored in volatile DRAM, making the cached data susceptible to loss during system shutdowns. Consequently, cold restarts experience significant delays. Persistent memory is a byte-addressable, large-capacity, and non-volatility storage media, which can be employed to avoid the cold restart problem. However, deploying Memcached on persistent memory requires consideration of issues such as write endurance, asymmetric read/write latency and bandwidth, and write granularity of persistent memory. In this paper, we propose Hybrid-Memcached, an optimized Memcached framework based on a hybrid combination of DRAM and persistent memory. Hybrid-Memcached includes three key components: (1) a DRAM-based data aggregation buffer to avoid multiple fine-grained writes, which extends the write endurance of persistent memory, (2) a data-object alignment mechanism to avoid write amplification, and (3) a non-temporal store instruction-based writing strategy to improve the bandwidth utilization. We have implemented Hybrid-Memcached on the Intel Optane persistent memory. Several micros-benchmarks are designed to evaluate Hybrid-Memcached by varying read/write ratios, access distributions, and key-value item sizes. Additionally, we evaluated it with the YCSB benchmark, showing a 21.2% performance improvement for fully write-intensive workloads and 11.8% for read-write balanced workloads.
Zhang Jiang, Xianduo Li, Tianxiang Peng, Haoran Li 0014, Jingxuan Hong, Jin Zhang 0003, Xiaoli Gong
IEEE Trans. Computers6
2023 Privacy-Preserving Multi-Source Domain Adaptation for Medical Data
abstract
Great progress has been made in diagnosing medical diseases based on deep learning. Large-scale medical data are expected to improve deep learning performance further. It is almost impossible for a single institution to collect so much data due to the time-consuming and costly collection and labeling of medical data. Many studies have turned attention to data sharing among multiple medical institutions. However, due to different data acquiring and processing procedures, multiple institutions' medical data is characterized by distribution heterogeneity. Besides, the protection of patient privacy in medical data sharing has also been a common concern. To simultaneously address the problems of heterogeneous data distribution and privacy protection, we propose a novel multi-source source free domain adaptation. When aligning distributed heterogeneous data, our method only require to transfer the pre-trained source models rather than the direct source domain data, thus protecting patients' privacy. In addition, it has the advantages of being efficient and less costly in network resources. The proposed method is evaluated on the multi-site fMRI database Autism Brain Imaging Data Exchange (ABIDE) and yields an average accuracy of 69.37%. We also analyzed its effectiveness on network resource-saving and conducted additional experiments on Camelyon17 to validate the generalization.
Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Yu Zhang 0009
IEEE J. Biomed. Health Informatics4
2023 Liberator: A Data Reuse Framework for Out-of-Memory Graph Computing on GPUs
abstract
Graph analytics are widely used including recommender systems, scientific computing, and data mining. Meanwhile, GPU has become the major accelerator for such applications. However, the graph size increases rapidly and often exceeds the GPU memory, incurring severe performance degradation due to frequent data transfers between the main memory and GPUs. To relieve this problem, we focus on the utilization of data in GPUs by taking advantage of the data reuse across iterations. In our studies, we deeply analyze the memory access patterns of graph applications at different granularities. We have found that the memory footprint is accessed with a roughly sequential scan without a hotspot, which infers an extremely long reuse distance. Based on our observation, we propose a novel framework, calledLiberator, to exploit the data reuse within GPU memory. InLiberator, GPU memory is reserved for the data potentially accessed across iterations to avoid excessive data transfer between the main memory and GPUs. For the data not existing in GPU memory, a Merged and Aligned memory access manner is employed to improve the transmission efficiency. We also further optimize the framework by parallel processing of data in GPU memory and data in the main memory. We have implemented a prototype of theLiberatorframework and conducted a series of experiments on performance evaluation. The experimental results show thatLiberatorcan significantly reduce the data transfer overhead, which achieves an average of 2.7x speedup over a state-of-the-art approach.
Ruiqi Tang, Xiaoli Gong, Wenwen Wang 0001, Jin Zhang 0003, Pen-Chung Yew
IEEE Trans. Parallel Distributed Syst.7
2021 Ascetic: Enhancing Cross-Iterations Data Efficiency in Out-of-Memory Graph Processing on GPUs
abstract
Graph analytics are widely used in real-world applications, and GPUs are major accelerators for such applications. However, as graph sizes become significantly larger than the capacity of GPU memory, the performance can degrade significantly due to the heavy overhead required in moving a large amount of graph data between CPU main memory and GPU memory.
Ruiqi Tang, Kailun Wang, Xiaoli Gong, Jin Zhang 0003, Wenwen Wang 0001, Pen-Chung Yew
ICPP5
2021 Serial-EMD: Fast empirical mode decomposition method for multi-dimensional signals based on serialization
abstract
Empirical mode decomposition (EMD) has developed into a prominent tool for adaptive, scale-based signal analysis in various fields like robotics, security and biomedical engineering. Since the dramatic increase in amount of data puts forward higher requirements for the capability of real-time signal analysis, it is difficult for existing EMD and its variants to trade off the growth of data dimension and the speed of signal analysis. In order to decompose multi-dimensional signals at a faster speed, we present a novel signal-serialization method (serial-EMD), which concatenates multi-variate or multi-dimensional signals into a one-dimensional signal and uses various one-dimensional EMD algorithms to decompose it. To verify the effects of the proposed method, synthetic multi-variate time series, artificial 2D images with various textures and real-world facial images are tested. Compared with existing multi-EMD algorithms, the decomposition time becomes significantly reduced. In addition, the results of facial recognition with Intrinsic Mode Functions (IMFs) extracted using our method can achieve a higher accuracy than those obtained by existing multi-EMD algorithms, which demonstrates the superior performance of our method in terms of the quality of IMFs. Furthermore, this method can provide a new perspective to optimize the existing EMD algorithms, that is, transforming the structure of the input signal rather than being constrained by developing envelope computation techniques or signal decomposition methods. In summary, the study suggests that the serial-EMD technique is a highly competitive and fast alternative for multi-dimensional signal analysis.
Jin Zhang 0003, Pere Martí-Puig, Cesar F. Caiafa, Zhe Sun 0009, Feng Duan 0006, Jordi Solé i Casals
Inf. Sci.1
2021 A Thread Level SLO-Aware I/O Framework for Embedded Virtualization
abstract
With the development of virtualization technology, it is practical and necessary to integrate virtual machine software into embedded systems. I/O scheduling is important for embedded systems, because embedded systems always face different situations and their requests have more diversity on the requirement of real-time and importance. However, the semantic information associated with the I/O data is completely lost when crossing the virtualized I/O software stack. Here, we present an I/O scheduling framework to connect the semantic gap between the application threads in virtual machines and hardware schedulers in the host machine. Therefore, the details for the I/O request can be passed through the layers of the software stack and each layer can get the specific information about the device environment. Also, various scheduling points have been provided to implement different I/O strategies. Our framework was implemented based on Linux operating system, KVM, QEMU and virtio protocol. A prototype scheduler, Orthrus, was implemented to evaluate the effectiveness of the framework. Comprehensive experiments were conducted and the results show that our framework can guarantee the real-time requirements, and reserve more system resources for critical tasks, with negligible memory consumption and throughput overhead.
Xiaoli Gong, Dingyuan Cao 0001, Yusen Li, Jin Zhang 0003, Tao Li 0022
IEEE Trans. Parallel Distributed Syst.6
2018 A Webpage Offloading Framework for Smart Devices
Jin Zhang 0003, Weilai Liu, Wenjian Zhao, Haocong Xu, Xiaoli Gong
Mob. Networks Appl.1
2017 Automatically Difficulty Grading Method Based on Knowledge Tree
Jin Zhang 0003, Haoxiang Yang, Xiaoli Gong
KSEM1
2016 WWOF: An Energy Efficient Offloading Framework for Mobile Webpage
abstract
Currently, the smart-phone has become a significant part for many people. As the major role in providing excellent surfing experiences for users, web browser can not only serve web sites visiting, but also support mobile web applications in smart-phones. In the meantime, in order to attract users, mobile web applications provide more and more diverse choices, which, however, results in the increase of CPU usage, time and energy consumption. Hence, how to improve efficiency of these applications without affecting user experiences has became a significant project. One of the solutions is to migrate the heavy computing tasks of mobile web applications from browser to cloud, which is called Offloading. This paper presents the design and implementation of a generic framework to realize the offloading from local browsers to cloud--Web Worker Offloading Framework (WWOF). The designed framework is easy to apply and is fully compatible with HTML5 API. On client side, the native Web Worker API is replaced by a library to offload seamlessly. On cloud side, corresponding interfaces are implemented to run workers on the server. To prove the effect of WWOF, evaluation of several benchmarks has been made, which shows 85% energy saving and 2-4 times execution speed-up on average for some mobile web applications.
Xiaoli Gong, Weilai Liu, Jin Zhang 0003, Haocong Xu, Wenjian Zhao
MobiQuitous3
2011 Automatic Model Building and Verification of Embedded Software with UPPAAL
abstract
Embedded systems are becoming ubiquitous and taking more and more important part in our daily life. Increasingly complex functionality leads to higher develop cost and lower software quality. Model checking has the potential of alleviating these problems. In this paper, we present an approach to construct model directly from the source code. An embedded system design language, Virgil, is selected as the target. Without losing any information, the UPPAAL model is generated based on the typed intermediate language. The timing information and stack behavior are estimated and after merging the hardware platform model, the whole system can be simulated on the model checker and some safety and aliveness properties of the program are verified.
Xiaoli Gong, Qingcheng Li, Jin Zhang 0003
TrustCom4