EDBT 2026 Demo / reviewers in the wild / expert
Zhan Zhang 0002
dblp:92/6841-2
· DBLP profile ↗
28ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-7419-4272ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 2Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MARS: Multi-model Aware Real-Time Scheduler for NPU-Coordinated DLI Tasks
Zhan Zhang 0002, Yuzhou Huo, De-Cheng Zuo, Yanjun Shu |
Euro-Par (2) | 2 |
| 2025 | VLIMNet: A Visible Light And Infrared Image Matching Network Based On Segment Anything Model And SuperPointabstractThis paper introduces a novel method for matching visible light and infrared images, termed the Visible Light and Infrared Image Matching Network (VLIMNet). In the image encoding stage, we incorporate a generative architecture-based modality transformation network after the SuperPoint encoder, enabling the local geometric features extracted from infrared images to more closely resemble those of visible light images. This reduces the impact of modality differences. Simultaneously, we utilize the image encoder of the Segment Anything Model to obtain global semantic descriptors. During the decoding stage, we fuse the global semantic descriptors with the local geometric descriptors and perform joint decoding to obtain keypoints and new feature vectors. The matching process is computed by solving a differentiable optimal transport problem based on the LightGlue network, a graph neural network with an attention mechanism. Compared to other matching models, our approach demonstrates improvements across various metrics in the domain of infrared and visible light image matching, particularly excelling in matching images with significant pose differences. Specifically, our method achieves approximately a 5% improvement in matching accuracy compared to the highest accuracy matching methods. Zhongyuan Chen, Zhan Zhang 0002, De-Cheng Zuo, Liufeng Fan |
ICASSP | 2 |
| 2025 | Data Glove-based Personalized Continuous Gesture SegmentationabstractIn recent years, gesture recognition based on data gloves has attracted increasing attention as a human-computer interaction (HCI) method that is natural, convenient, stable, robust, easy to recognize, and applicable to various usage environments. This research first proposes an advanced smart data glove that integrates cutting-edge flexible capacitive sensors on the fingertips and a 6-axis IMU on the back of the hand to recognize gestures. Secondly, this study proposes a personalized continuous gesture segmentation (PCGS) model that can adaptively calculate the most appropriate gesture segmenting threshold based on the current user and introduces the multi-sliding window theory and kinematic knowledge to perform personalized gesture segmentation. The accuracy of gesture segmentation can reach 94.3%. The result shows that our PCGS model achieves an average segmentation accuracy of 94.3% and outperforms the state-of-the-art pproaches by 11.2% to 18.5%. Liufeng Fan, Zhan Zhang 0002, De-Cheng Zuo, Yinran Wang, Zhongyuan Chen |
ICASSP | 2 |
| 2025 | SDAD: A Service Deployment Method Based on Association Rule and Reinforcement Learning for Edge Computing
Hanzhi Xu, Yanjun Shu, Wei Zhang 0098, Zhuangyu Ma, Zhan Zhang 0002, De-Cheng Zuo |
ICSOC (1) | 5 |
| 2025 | ReIDFaaS: An Energy-Efficient Serverless Person Re-Identification System Across the Edge-Cloud ContinuumabstractPerson re-identification (Re-ID) systems in edgecloud continuum face critical trade-offs between latency sensitivity and energy efficiency due to the resource-constrained edge environment. This paper proposes Re-IDFaaS, a serverless ReID system that dynamically optimizes energy consumption and computational performance across the edge-cloud continuum. Leveraging serverless architectures, our system implements three key improvements: (1) An event-driven workflow triggered by motion detection, eliminating idle GPU resource consumption during inactive periods. (2) A hardware-aware dynamic scheduler that allocates tasks based on real-time energy states and container availability, achieving balanced resource utilization across heterogeneous nodes. (3) An adaptive batching mechanism that reduces cold-start frequency through latency-constrained request grouping while maintaining the efficiency of GPU memory. Experiments demonstrate a 23.3% improvement in edge node availability and 55% reduction in memory usage compared to existing methods. The system design provides practical insights for building AI services in hybrid computing environments requiring cross-framework compatibility and adaptive resource orchestration, achieving 53% higher throughput than traditional architectures. These innovations address the challenges of dynamic workload scheduling and runtime optimization in hardware-diverse scenarios, ensuring sustainable operation under bursty surveillance workloads. Jianping Pei, Yanjun Shu, Zhuangyu Ma, De-Cheng Zuo, Zhan Zhang 0002 |
ICWS | 5 |
| 2024 | ITRMD: A Dimensionality Reduction Framework for Accurate and Efficient Multivariate KPI Anomaly Detection
Tianrun Gao, De-Cheng Zuo, Yanjun Shu, Zhan Zhang 0002, Dongxin Wen, Yutong Qu |
ADMA (4) | 4 |
| 2024 | An algorithm/hardware co-optimized method to accelerate CNNs with compressed convolutional weights on FPGAabstractSummary Convolutional neural networks (CNNs) have shown remarkable advantages in a wide range of domains at the expense of huge parameters and computations. Modern CNNs still tend to be more complex and larger to achieve better inference accuracy. However, the complex and large structures of CNNs could slow down the inference speed. Recently, Compressing the convolutional weights to be sparse by pruning the unimportant parameters has been demonstrated as an efficient way to reduce the computations of CNNs. On the other hand, field‐programmable gate arrays (FPGAs) have been a popular hardware platform to accelerate CNN inference. In this paper, we propose an algorithm/hardware co‐optimized method for accelerating CNN inference on FPGAs. For the algorithm, we take advantage of unstructured and structured parameter sparsifying methods to achieve high sparsity and keep the regularity of convolutional weights. Correspondingly, hardware‐friendly index representations of sparse convolutional weights are proposed. For the hardware architecture, we propose row‐wise input‐stationary dataflow, which is tightly coupled with the algorithm. A row‐wise computing engine (RConv Engine) is proposed, which is based on the dataflow. Inside the RConv Engine, the scalar‐vector structure is applied to implement the basic processing elements (PEs). To flexibly calculate the feature map with various sizes, the PEs are organized in a 2D structure with two work modes. The experimental results demonstrate that our co‐optimized method implements high sparsity of convolutional weights, and the computing engine achieves high computation efficiency. Compared with other accelerators, our co‐optimized method implements a 10.9 speedup on FPS at most with the highest sparsity of convolutional weights and negligible accuracy loss. Jiangwei Shang, Zhan Zhang 0002, Chuanyou Li, Hongwei Liu 0002 |
Concurr. Comput. Pract. Exp. | 2 |
| 2024 | KLNK: Expanding Page Boundaries in a Distributed Shared Memory SystemabstractSoftware-based distributed shared memory (DSM) allows multiple processes to access shared data without the need for specialized hardware. However, this flexibility comes at a significant cost due to the need for data synchronization. One approach to mitigate these costs is to relax the consistency model, which can lead to delayed updates to the shared data. This approach typically requires the use of explicit synchronization primitives to regulate access to the shared memory and determine the timing of data synchronization. To circumvent the need for explicit synchronization, an alternative approach is to manage shared memory transparently using the underlying system. While this can simplify programming, it often imposes a fixed granularity for data sharing, which can limit the expansion of the coherence domain and increase the synchronization requirements. To overcome this limitation, we propose an abstraction called the elastic coherence domain, which dynamically adjusts the scope of data synchronization and is supported by the underlying system for transparent management of shared memory. The experimental results show that this approach can improve the efficiency of memory sharing in distributed environments. Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | QoS Prediction via Multi-scale Feature Fusion Based on Convolutional Neural Network
Hanzhi Xu, Yanjun Shu, Zhan Zhang 0002, De-Cheng Zuo |
ICSOC (1) | 3 |
| 2023 | A high-performance convolution block oriented accelerator for MBConv-Based CNNs
Jiangwei Shang, Zhan Zhang 0002, Chuanyou Li, Hongwei Liu 0002 |
Integr. | 3 |
| 2023 | An adaptive non-migrating load-balanced distributed stream window join system
De-Cheng Zuo, Zhan Zhang 0002, Tianming Liu 0003 |
J. Supercomput. | 3 |
| 2022 | SepJoin: A Distributed Stream Join System with Low Latency and High ThroughputabstractIn the field of real-time analytics, stream joins are the basis for complex queries and greatly affect system performance. In order to satisfy the real-time requirements of streaming applications, the system imposes high requirements on the latency and throughput of the stream join operator. In this paper, we model the latency and throughput of distributed stream join systems based on queuing theory. Based on the analysis of this model, we demonstrate the impact of indexing-related overhead on the latency and throughput of stream join systems and propose a new distributed stream join system, SepJoin, which is oriented to the hash join problem. SepJoin reduces the number of tuples stored in each processing unit belonging to each input stream by designing a novel partitioning scheme that uses as many processing units as possible to store tuples belonging to each input stream, thereby reducing the index-related overhead of each processing unit when performing join operations and ultimately achieving performance benefits in terms of latency and throughput. We provide both theoretical analysis and extensive experimental evaluations to evaluate the processing latency and max throughput of SepJoin. De-Cheng Zuo, Zhan Zhang 0002, Tianming Liu 0003 |
ICPADS | 3 |
| 2022 | Dynamic Adaptive Checkpoint Mechanism for Streaming Applications Based on Reinforcement LearningabstractFor a stream processing system that uses checkpoints as a fault-tolerant method, selecting the appropriate checkpoint period is the key to ensuring the efficient operation of streaming applications. State-of-art stream processing systems currently only support fixed-cycle checkpoints, which is difficult to make a good trade-off between fault-tolerant processing and the cost of failure recovery in dynamically changing streaming application scenarios. Moreover, in a complex distributed streaming application environment, the dynamic environmental indicators (e.g., the values of workloads and failure rates) are not in coincidence with the model assumptions, such as the dynamics of Twitter’s hot events data changing quickly. In this paper, we consider the dynamic changes of environmental indicators and adaptively optimize the processing delay and fault recovery time. Then, we propose a dynamic adjustment method for the checkpoint interval by reinforcement learning, which is named DACM. DACM adaptively optimizes the processing delay and fault recovery time, while avoiding the overall environment modeling of streaming applications. The experiments conducted on the Flink platform show that DACM reduces the processing delay by 10% and the failure recovery time by 37% compared with the existing checkpoint interval optimization models. Zhan Zhang 0002, Tianming Liu 0003, Yanjun Shu |
ICPADS | 1 |
| 2022 | Toward optimal operator parallelism for stream processing topology with limited buffers
Zhan Zhang 0002, Yanjun Shu, Hongwei Liu 0002, Tianming Liu 0003 |
J. Supercomput. | 2 |
| 2021 | Research on Optimal Checkpointing-Interval for Flink Stream Processing Applications
Zhan Zhang 0002, Xiao Qing, Hongwei Liu 0002 |
Mob. Networks Appl. | 1 |
| 2020 | Random Priority-Based Thrashing Control for Distributed Shared MemoryabstractShared memory is widely used for inter-process communication. The shared memory abstraction allows computation to be decoupled from communication, which offers benefits, including portability and ease of programming. To enable shared memory access by processes that are on different machines, distributed shared memory (DSM) can be employed. However, DSM systems can suffer from thrashing: while different processes update certain hot data items, the largest amount of effort is spent on data synchronization, and little progress is made by each process. To avoid interference between processes during data updating while providing shared memory at page granularity, more time is reserved for a writer to hold a page in a traditional manner. In this paper, we report on complex thrashing, which can explain why extending the time of holding a page might not be sufficient to control thrashing. To increase the throughput, we propose a thrashing control mechanism that allows each process to update a set of pages during a period of time, where the pages compose a logical area. Because of the isolation of areas, updates on different areas can be performed concurrently. To allow the areas to be fairly well used, each process is assigned with a random priority for thrashing control. The thrashing control mechanism is implemented on a Linux-based DSM system. Performance results show that the execution time of the applications that are apt to cause system thrashing can be significantly reduced by our approach. Yiwei Ci, Michael R. Lyu, Zhan Zhang 0002, De-Cheng Zuo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Improving the Dependability of Self-Adaptive Cyber Physical System With Formal Compositional ContractabstractTo adapt to the uncertain environment smartly and timely, cyber physical systems (CPSs) have to interact with the physical world in a decentralized but rigorous, organized way. Guaranteeing the timing reliability is key to achieve consensus on the order of distributed events, as well as dependable cooperative decision processing. Based on our hierarchically decentralized compositional self-adaptive framework, we propose a formal compositional reliability-contract-based solution to guarantee the timing reliability of event observation and decision processing in a large-scale, geographically distributed CPS. As the prophetic decision may not fit the local situation well because of the uncertainties, we propose a gradual contract optimization solution to refine the dependability, timeliness, and energy consumption. Following the seven proposed composition schemes, we employ the nondominated sorting genetic algorithm II (NSGA-II) algorithm to optimize arrangement of decision. Moreover, a topology-aware time reserving solution is applied to improve the resilience of processing time and to tolerance timing failures. Both simulation results and real-world testing are introduced to evaluate the efficacy of our proposal. We believe that the formal compositional contract will be a competitive CPS solution to analyze requirements and optimize the self-adaptation decision at runtime. Peng Zhou 0005, De-Cheng Zuo, Kun Mean Hou, Zhan Zhang 0002, Jian Dong 0010 |
IEEE Trans. Reliab. | 4 |
| 2019 | Reducing the upfront cost of private clouds with clairvoyant virtual machine placement
Hongwei Liu 0002, Yan Wang 0002, Zhan Zhang 0002, De-Cheng Zuo |
J. Supercomput. | 4 |
| 2018 | Predicting the quality of online health expert question-answering services with temporal features in a deep learning framework
Ze Hu, Zhan Zhang 0002, Haiqin Yang, De-Cheng Zuo |
Neurocomputing | 2 |
| 2018 | Factorization machines and deep views-based co-training for improving answer quality prediction in online health expert question-answering services
Zhan Zhang 0002, Ze Hu, Haiqin Yang, De-Cheng Zuo |
J. Biomed. Informatics | 1 |
| 2017 | A deep learning approach for predicting the quality of online health expert question-answering services
Ze Hu, Zhan Zhang 0002, Haiqin Yang, De-Cheng Zuo |
J. Biomed. Informatics | 2 |
| 2015 | An imperfect software debugging model considering log-logistic distribution fault content function
Jinyong Wang, Zhibo Wu, Yanjun Shu, Zhan Zhang 0002 |
J. Syst. Softw. | 4 |
| 2012 | A multi-cycle checkpointing protocol that ensures strict 1-rollback
Yiwei Ci, Zhan Zhang 0002, De-Cheng Zuo, Zhibo Wu |
Inf. Process. Lett. | 2 |
| 2010 | Study for Performance Benchmark of Bank Intermediary Business on High-Performance Fault-Tolerant ComputersabstractThe dominant position of High-Performance Fault-Tolerant (HPFT) computers in security and economics has advanced the studies on the performance benchmarks on the HPFT computers in the specific field, such as bank finance and telecommunication etc. Although TPC (Transaction Processing Council) has proposed some benchmarks models for different OLTP (On-Line Transaction Processing) complex business, such as TPC-C and TPC-E, there is still a lack of the performance benchmark model dedicated to the bank intermediary business on HPFT computers. This paper proposes a Bank Intermediary Business performance benchmark (BIBbench), and gives a solution to test and evaluate this benchmark on HPFT computers for bank intermediary business. In this paper, we present the architecture of BIBbench, defining the structures and attributes of the business model, the database model and the transaction/frame model, and illuminating the workload generation mechanism of the intermediary business system as well. The BIBbench testing environment architecture is also discussed in the paper, as well as the testing solutions and tools. Currently, this BIBbench has been partly implemented on the Oracle 10g database system, and some performance testing experiences based on the BIBbench for HPFT computers have been made. Haiying Zhou, De-Cheng Zuo, Zhan Zhang 0002 |
ISPA | 4 |
| 2010 | Dependency mining-based causal message logging
Yiwei Ci, Zhan Zhang 0002, De-Cheng Zuo, Zhibo Wu |
Inf. Process. Lett. | 2 |
| 2009 | Communication-Based Prevention of Non-P-PatternabstractAn issue pertinent to the design of checkpointing protocols is how to improve the autonomy of checkpointing and keep computation loss under control. To address the problem, a time-based multi-cycle checkpointing protocol is proposed in this paper. In this protocol, processes are allowed to take checkpoints with desired checkpoint cycles. To enable recent checkpoints to be used to form a consistent global checkpoint, a communication-based checkpoint cycle adjustment approach is also proposed. In this approach, the checkpoint cycle adjustment of each process follows a P-pattern. Simulation results show that the rollback deviation of the proposed protocol can be well controlled under a low checkpointing overhead. Yiwei Ci, Zhan Zhang 0002, De-Cheng Zuo, Zhibo Wu |
SRDS | 2 |
| 2009 | Message fragment based causal message logging
Yiwei Ci, Zhan Zhang 0002, De-Cheng Zuo |
J. Parallel Distributed Comput. | 2 |
| 2008 | Area Difference Based Recovery Information Placement for Mobile Computing SystemsabstractIn a mobile computing system, mobile hosts may move around cells, resulting in a considerable cost for locating and retrieving the recovery information, which is necessary for fault tolerance. To speed up the recovery, traditionally, recovery information is migrated according to the location of the mobile host. In this paper, a scheme for efficiently handling the recovery information is proposed. When a mobile host moves out of a certain range, only partial recovery information of the mobile host needs to be migrated to mobile support stations. It can avoid the unnecessary migration of recovery information. Moreover, the performance of the proposed scheme is evaluated and compared with the traditional movement based scheme. Yiwei Ci, Zhan Zhang 0002, De-Cheng Zuo, Zhibo Wu |
ICPADS | 2 |