Zeqi Zhu

dblp:275/4096 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-8614-9855ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 100%
Network and information security
1 paper
Privacy and data protection · 87% Cryptographic protocols and secure computation · 13%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 50% Information retrieval · 50%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Hardware accelerators and domain-specific architectures · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
1.122025
MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks · CVPR 2025
ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration · ECCV (11) 2024
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
0.912025
MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks · CVPR 2025
Machine learning › Efficient and distributed learning
model compression
0.912025
MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks · CVPR 2025
Machine learning › Efficient and distributed learning
inference acceleration
0.812024
ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration · ECCV (11) 2024
Machine learning › Efficient and distributed learning › model compression › sparsity
sparsity exploitation
0.812024
ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration · ECCV (11) 2024
Query processing and optimization
similarity query processing
0.812024
FedSQ: A Secure System for Federated Vector Similarity Queries · Proc. VLDB Endow. 2024
Information retrieval › similarity search
vector similarity search
0.812024
FedSQ: A Secure System for Federated Vector Similarity Queries · Proc. VLDB Endow. 2024
Privacy and data protection
federated query processing
0.812024
FedSQ: A Secure System for Federated Vector Similarity Queries · Proc. VLDB Endow. 2024
Privacy and data protection
privacy-preserving query processing
0.812024
FedSQ: A Secure System for Federated Vector Similarity Queries · Proc. VLDB Endow. 2024
Cryptographic protocols and secure computation
secure multiparty computation
0.212024
FedSQ: A Secure System for Federated Vector Similarity Queries · Proc. VLDB Endow. 2024

Methods — techniques the papers use, named apart from their topics

weight compression · 1.7network architecture co-design · 1.7delta-sigma convolution · 1.7secure multiparty computation · 1.5sampling · 1.5line-based sparsity · 1.5indexing · 1.5deep neural network inference · 1.5
YearPublicationVenuePosition
2025 MEET: Towards Memory-Efficient Temporal Sparse Deep Neural Networks
abstract
Deep Neural Networks (DNNs) are accurate but compute-intensive, leading to substantial energy consumption during inference. Exploiting temporal redundancy through ∆-Σ convolution [26] in video processing has proven to greatly enhance computation efficiency. However, temporal ∆-Σ DNNs typically require substantial memory for storing neuron states to compute inter-frame differences, hindering their on-chip deployment. To mitigate this memory cost, directly compressing the states can disrupt the linearity of temporal ∆-Σ convolution, causing accumulated errors in long-term ∆-Σ processing. Thus, we propose MEET, an optimization framework for MEmory-Efficient Temporal ∆-Σ DNNs. MEET transfers the state compression challenge to a well-established weight compression problem by trading fewer activations for more weights and introduces a co-design of network architecture and suppression method to optimize for mixed spatial-temporal execution. Evaluations on three vision applications demonstrate a reduction of 5.1∼13.3 × in total memory compared to the most computation-efficient temporal DNNs, while preserving the computation efficiency and model accuracy in long-term ∆-Σ processing. MEET facilitates the deployment of temporal ∆-Σ DNNs within on-chip memory of embedded event-driven platforms, empowering low-power edge processing.
Zeqi Zhu, Ibrahim Batuhan Akkaya, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
CVPR1
2025 Neuromorphic Edge Computing: Challenges, Opportunities, and Current Solutions
abstract
Neuromorphic computing is emerging as a paradigm for high-performance, energy-efficient edge intelligence. Yet the transition from laboratory prototypes to deployable edge platforms is slowed by four intertwined obstacles: (1) complex near-sensor integration, where spiking inference must co-locate with analogue sensing to minimise latency and data-movement energy; (2) novel event-based optimisation, requiring weight compression and sparsity techniques tailored to event-driven workloads; (3) heterogeneous integration of emerging devices, such as RRAM and other non-volatile memories, into reliable, manufacturable stacks; and (4) novel security risks, including spike-pattern side channels and model-specific attacks that demand to develop neuromorphic security primitives. This paper surveys the state of the art across these four fronts, drawing on recent advances in spiking microcontrollers, mixed-precision compute-in-memory fabrics, sparsity-aware compilation, and hardware-anchored security primitives (physical unclonable functions, true random number generators, computing-in-memory-based cryptography). By distilling lessons from academic research and industrial prototyping, the paper outlines current solutions and future research directions aimed at accelerating the adoption of neuromorphic platforms in real-world edge AI systems.
Federico Corradi, Amir Zjajo, Letícia Maria Veiras Bolzani, Milos Krstic, Orlando Moreira, Zeqi Zhu, Farhad Merchant
ISLPED6
2024 ELSE: Efficient Deep Neural Network Inference Through Line-Based Sparsity Exploration
Zeqi Zhu, Alberto García Ortiz, Luc Waeijen, Egor Bondarev, Arash Pourtaherian, Orlando Moreira
ECCV (11)1
2024 CATS: Combined Activation and Temporal Suppression for Efficient Network Inference
abstract
Brain-inspired event-driven processors execute deep neural networks (DNNs) in a sparsity-aware manner, leading to superior performance compared to conventional platforms. In the pursuit of higher event sparsity, prior studies suppress non-zero events by either eliminating the intra-frame activations (spatially) or leveraging the redundancy in the inter-frame differences for a video (temporally). However, we have empirically observed that simultaneously enhancing activation and temporal sparsity can lead to a synergistic suppression outcome. To this end, we propose an end-to-end event suppression training approach CATS −− Combined Activation and Temporal Suppression for efficient network inference. It utilizes a gradient-based method to search for the optimal temporal thresholds per layer while penalizing the presence of events in both spatial and temporal domains. Our experimental results show that CATS achieves 2 ∼ 6× higher event suppression compared to the inherent ReLU suppression across a wide range of vision applications, consistently outperforming the state-of-the-art (SOTA) methods by a significant margin at all accuracy levels. Furthermore, a case study on the commercial event-driven processor GrAI-VIP highlights that the induced event sparsity in SSD on the EgoHands dataset can be efficiently translated into a performance enhancement of 2.5× in FPS, 2.1× in latency, and 3.8× in energy consumption, while maintaining the model accuracy.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Ibrahim Batuhan Akkaya, Egor Bondarev, Orlando Moreira
WACV1
2024 FedSQ: A Secure System for Federated Vector Similarity Queries
abstract
Vector databases have emerged as crucial tools for managing and retrieving representation embeddings of unstructured data. Given the explosive growth of data, vector data is often distributed and stored across multiple organizations. However, privacy concerns and regulations like GDPR present new challenges in collaborative and secure queries, also known as federated queries, over those vector data distributed across various data owners. Although existing research has attempted to enable such query services for low-dimensional data, such as relational and spatial data, these solutions can be inefficient in answering vector similarity queries involving high-dimensional data. Therefore, we are motivated to develop a new prototype system called FedSQ that (1) ensures privacy protection across data owners and (2) balances query efficiency and result accuracy when processing federated vector similarity queries. To achieve these goals, FedSQ utilizes advanced secure multi-party computation techniques to prevent information leakage during query processing and incorporates indexing and sampling based optimizations to strike a proper performance balance.
Zeqi Zhu, Zeheng Fan, Yuxiang Zeng, Yexuan Shi, Yi Xu 0013, Mengmeng Zhou, Jin Dong 0004
Proc. VLDB Endow.1
2022 ARTS: An adaptive regularization training schedule for activation sparsity exploration
abstract
Brain-inspired event-based processors have attracted considerable attention for edge deployment because of their ability to efficiently process Convolutional Neural Networks (CNNs) by exploiting sparsity. On such processors, one critical feature is that the speed and energy consumption of CNN inference are approximately proportional to the number of non-zero values in the activation maps. Thus, to achieve top performance, an efficient training algorithm is required to largely suppress the activations in CNNs. We propose a novel training method, called Adaptive-Regularization Training Schedule (ARTS), which dramatically decreases the non-zero activations in a model by adaptively altering the regularization coefficient through training. We evaluate our method across an extensive range of computer vision applications, including image classification, object recognition, depth estimation, and semantic segmentation. The results show that our technique can achieve 1.41 × to 6.00 × more activation suppression on top of ReLU activation across various networks and applications, and outperforms the state-of-the-art methods in terms of training time, activation suppression gains, and accuracy. A case study for a commercially-available event-based processor, Neuronflow, shows that the activation suppression achieved by ARTS effectively reduces CNN inference latency by up to 8.4 × and energy consumption by up to 14.1 ×.
Zeqi Zhu, Arash Pourtaherian, Luc Waeijen, Lennart Bamberg, Egor Bondarev, Orlando Moreira
DSD1
2022 Applications of graph convolutional networks in computer vision
Pingping Cao, Zeqi Zhu, Qiang Niu
Neural Comput. Appl.2