VLDB 2026 Research / reviewers in the wild / expert
Zhe Tao
dblp:199/7027
· DBLP profile ↗
13ranked-venue papers
7as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Language Guided Concept Bottleneck Models for Interpretable Continual LearningabstractContinual learning (CL) aims to enable learning systems to acquire new knowledge constantly without forgetting previously learned information. CL faces the challenge of mitigating catastrophic forgetting while maintaining interpretability across tasks. Most existing CL methods focus primarily on preserving learned knowledge to improve model performance. However, as new information is introduced, the interpretability of the learning process becomes crucial for understanding the evolving decision-making process, yet it is rarely explored. In this paper, we introduce a novel framework that integrates language-guided Concept Bottleneck Models (CBMs) to address both challenges. Our approach leverages the Concept Bottleneck Layer, aligning semantic consistency with CLIP models to learn human-understandable concepts that can generalize across tasks. By focusing on interpretable concepts, our method not only enhances the model’s ability to retain knowledge over time but also provides transparent decision-making insights. We demonstrate the effectiveness of our approach by achieving superior performance on several datasets, outperforming state-of-the-art methods with an improvement of up to 3.06% in final average accuracy on ImageNet-subset. Additionally, we offer concept visualizations for model predictions, further advancing the understanding of interpretable continual learning. Code is available at https://github.com/FisherCats/CLG-CBM. Lu Yu 0004, Zhe Tao, Hantao Yao, Changsheng Xu |
CVPR | 3 |
| 2025 | Leveraging Multiple Deep Experts for Online Class-incremental LearningabstractOnline incremental learning aims to enable learning systems to continuously accumulate new knowledge from streaming data in a single-pass manner while preserving previously acquired information. This more realistic and challenging setting has gained increasing attention in recent years. The state-of-the-art methods treat each module of a model, from shallow to deep, as a separate sub-expert network and transfer all the shallow expert knowledge into the final deep expert network. Although this yields notable improvements, we argue that directly supervising shallow layers hampers their acquisition of task-invariant knowledge. Furthermore, explicitly designating the final expert as a student network to absorb knowledge from other experts lacks adaptability, considering that different experts may not excel uniformly across all tasks. To address the aforementioned limitations, we leverage the shallow layers of the model as a shared feature extractor, while the deeper layers form a set of experts capable of learning robust and diverse features. Moreover, to facilitate knowledge transfer between multiple experts, we introduce the LEEP score to assess the feature transferability of each expert on new tasks, thereby selecting the most suitable expert as the teacher network for the new task. Extensive experiments on two evaluation benchmarks verify the effectiveness of our method (e.g, up to 1.3% on Split CIFAR-100 and 2.5% on Split Tiny-ImageNet). Code is available at https://github.com/untitledunmastered1998/MDE-OIL. Zhe Tao, Lu Yu 0004, Hantao Yao, Changsheng Xu |
ICME | 1 |
| 2025 | Provable Gradient Editing of Deep Neural NetworksabstractIn explainable AI, DNN gradients are used to interpret the prediction; in safety-critical control systems, gradients could encode safety constraints; in scientific-computing applications, gradients could encode physical invariants. While recent work on provable editing of DNNs has focused on input-output constraints, the problem of enforcing hard constraints on DNN gradients remains unaddressed. We present ProGrad, the first efficient approach for editing the parameters of a DNN to provably enforce hard constraints on the DNN gradients. Given a DNN $\mathcal{N}$ with parameters $\theta$, and a set $\mathcal{S}$ of pairs $(\mathrm{x}, \mathrm{Q})$ of input $\mathrm{x}$ and corresponding linear gradient constraints $\mathrm{Q}$, ProGrad finds new parameters $\theta'$ such that $\bigwedge_{(\mathrm{x}, \mathrm{Q}) \in \mathcal{S}} \frac{\partial}{\partial \mathrm{x}}\mathcal{N}(\mathrm{x}; \theta') \in \mathrm{Q}$ while minimizing the changes $\lVert\theta' - \theta\rVert$. The key contribution is a novel *conditional variable gradient* of DNNs, which relaxes the NP-hard provable gradient editing problem to a linear program (LP), enabling ProGrad to use an LP solver to efficiently and effectively enforce the gradient constraints. We experimentally evaluated ProGrad via enforcing (i) hard Grad-CAM constraints on ImageNet ResNet DNNs; (ii) hard Integrated Gradients constraints on Llama 3 and Qwen 3 LLMs; (iii) hard gradient constraints in training a DNN to approximate a target function as a proxy for safety constraints in control systems and physical invariants in scientific applications. The results highlight the unique capability of ProGrad in enforcing hard constraints on DNN gradients. Zhe Tao, Aditya V. Thakur |
NeurIPS | 1 |
| 2025 | Thoughts Behind Attack: Enhancing Security Against Jailbreak Attacks Using Chain-of-Thought
Zhe Tao, Muyun Yang, Hongjiao Guan, Wenpeng Lu, Hailong Cao, Conghui Zhu, Tiejun Zhao |
NLPCC (4) | 1 |
| 2024 | Provable Editing of Deep Neural Networks using Parametric Linear RelaxationabstractEnsuring that a DNN satisfies a desired property is critical when deploying DNNs in safety-critical applications. There are efficient methods that can verify whether a DNN satisfies a property, as seen in the annual DNN verification competition (VNN-COMP). However, the problem of provably editing a DNN to satisfy a property remains challenging. We present PREPARED, the first efficient technique for provable editing of DNNs. Given a DNN $\mathcal{N}$ with parameters $\theta$, input polytope $P$, and output polytope $Q$, PREPARED finds new parameters $\theta'$ such that $\forall \mathrm{x} \in P . \mathcal{N}(\mathrm{x}; \theta') \in Q$ while minimizing the changes $\lVert{\theta' - \theta}\rVert$. Given a DNN and a property it violates from the VNN-COMP benchmarks, PREPARED is able to provably edit the DNN to satisfy this property within 45 seconds. PREPARED is efficient because it relaxes the NP-hard provable editing problem to solving a linear program. The key contribution is the novel notion of Parametric Linear Relaxation, which enables PREPARED to construct tight output bounds of the DNN that are parameterized by the new parameters $\theta'$. We demonstrate that PREPARED is more efficient and effective compared to prior DNN editing approaches i) using the VNN-COMP benchmarks, ii) by editing CIFAR10 and TinyImageNet image-recognition DNNs, and BERT sentiment-classification DNNs for local robustness, and iii) by training a DNN to model a geodynamics process and satisfy physics constraints. Zhe Tao, Aditya V. Thakur |
NeurIPS | 1 |
| 2024 | Class Incremental Learning for Light-Weighted NetworksabstractDespite deep neural networks (DNNs) show impressive performance across diverse tasks, they suffer from catastrophic forgetting when dealing with continuous data streams. Incremental learning aims to alleviate this phenomenon and enable DNNs to accumulate new knowledge to cope with the ever-changing world. Recently numerous advanced methods have been developed to enhance the incremental learning capabilities of neural networks. However, these methods mainly focus on the large networks, neglecting the unique needs of edged-device applications, which is surprisingly under-investigated in previous literature. In this paper, we propose two strategies for transferring knowledge from large teacher networks to light-weighted networks in class incremental learning. Specifically, in cases where the initial task contains a large number of categories, our static teacher strategy involves transferring knowledge from the teacher to the student network on the initial task to enhance the plasticity of the student network, and applying regularization constraints on the subsequent task to improve its stability. In a more challenging scenario where each task includes an equal number of categories, the dynamic teacher strategy continuously guides the student network on each task. We evaluate the proposed methods on CIFAR100, Tiny-ImageNet and ImageNet-subset datasets with different types of light-weighted networks (MobileNet, ShuffleNet). We observed that effective knowledge transfer resulting in the student network achieving performance comparable or even outperform the teacher network. Extensive and detailed experiments conducted on three datasets demonstrated the simplicity and effectiveness of our proposed method. Comprehensive analysis are also conducted including different factors and visualization. Zhe Tao, Lu Yu 0004, Hantao Yao, Shucheng Huang, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Shallow multi-branch attention convolutional neural network for micro-expression recognition
Shucheng Huang, Zhe Tao |
Multim. Syst. | 3 |
| 2023 | Architecture-Preserving Provable Repair of Deep Neural NetworksabstractDeep neural networks (DNNs) are becoming increasingly important components of software, and are considered the state-of-the-art solution for a number of problems, such as image recognition. However, DNNs are far from infallible, and incorrect behavior of DNNs can have disastrous real-world consequences. This paper addresses the problem of architecture-preserving V-polytope provable repair of DNNs. A V-polytope defines a convex bounded polytope using its vertex representation. V-polytope provable repair guarantees that the repaired DNN satisfies the given specification on the infinite set of points in the given V-polytope. An architecture-preserving repair only modifies the parameters of the DNN, without modifying its architecture. The repair has the flexibility to modify multiple layers of the DNN, and runs in polynomial time. It supports DNNs with activation functions that have some linear pieces, as well as fully-connected, convolutional, pooling and residual layers. To the best our knowledge, this is the first provable repair approach that has all of these features. We implement our approach in a tool called APRNN. Using MNIST, ImageNet, and ACAS Xu DNNs, we show that it has better efficiency, scalability, and generalization compared to PRDNN and REASSURE, prior provable repair methods that are not architecture preserving. Zhe Tao, Stephanie Nawas, Jacqueline L. Mitchell, Aditya V. Thakur |
Proc. ACM Program. Lang. | 1 |
| 2023 | SyReNN: A tool for analyzing deep neural networks
Matthew Sotoudeh, Zhe Tao, Aditya V. Thakur |
Int. J. Softw. Tools Technol. Transf. | 2 |
| 2023 | Filtering Out High Noise Data for Distributed Deep Neural NetworksabstractArtificial intelligence-based cyber-physical systems (CPS) applications have been spread across various fields such as smart cities, medical services, and industrial controls. When CPS devices are connected to a cloud server, big data streams generated by CPS devices impose enormous bandwidth pressure and exert excessive compute loads to the cloud server. Due to unpredictable environments and uncertainty in reality, these issues are mainly attributed to a large amount of high noise data captured and uploaded by CPS devices. To overcome these issues, this paper proposes a cyber-physical-cloud based framework for distributed deep neural networks (DDNNs) to prevent high noise data from being uploaded to the cloud. The proposed framework features a lightweight data filtering module enabled by depthwise separable convolutions to identify and filter out the high noise data that the cloud cannot recognize. Extensive experimental results demonstrate that the proposed data filtering module can achieve an accuracy of up to 83.72% in identifying high noise data and the proposed framework can effectively save bandwidth of up to 63.42% as compared to benchmarking methods. Note to Practitioners—This paper is motivated by the problems of enormous bandwidth pressure and excessive cloud compute loads in cyber-physical-cloud distributed computing paradigms. These problems are mainly caused by high noise data generated by CPS devices, because CPS devices often work in disturbing and unstable environments and there are uncontrollable uncertainties in reality. Especially for the emerging artificial intelligence-driven cyber-physical-cloud distributed paradigms, there is no existing research to solve the unnecessary transmission and cloud compute loads caused by high noise data. To tackle the challenge, this paper develops a novel cyber-physical-cloud distributed framework with data filtering capabilities to prevent high noise data from being uploaded. The proposed framework supports two popular loosely coupled and closely coupled distributed computing paradigms. Extensive experiments confirm that the proposed cyber-physical-cloud distributed framework can efficiently filter out high noise data and alleviate unnecessary transmission and needless cloud compute loads introduced by high noise data. Yangguang Cui, Liying Li 0002, Zhe Tao, Mingsong Chen 0001, Tongquan Wei |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2021 | DICE*: A Formally Verified Implementation of DICE Measured Boot
Zhe Tao, Aseem Rastogi, Kapil Vaswani, Aditya V. Thakur |
USENIX Security Symposium | 1 |
| 2021 | FDA$^3$: Federated Defense Against Adversarial Attacks for Cloud-Based IIoT ApplicationsabstractAlong with the proliferation of artificial intelligence and Internet of things (IoT) techniques, various kinds of adversarial attacks are increasingly emerging to fool deep neural networks (DNNs) used by industrial IoT (IIoT) applications. Due to biased training data or vulnerable underlying models, imperceptible modifications on inputs made by adversarial attacks may result in devastating consequences. Although existing methods are promising in defending such malicious attacks, most of them can only deal with limited existing attack types, which makes the deployment of large-scale IIoT devices a great challenge. To address this problem, in this article, we present an effective federated defense approach named FDA3that can aggregate defense knowledge against adversarial examples from different sources. Inspired by federated learning, our proposed cloud-based architecture enables the sharing of defense capabilities against different attacks among IIoT devices. Comprehensive experimental results show that the generated DNNs by our approach can not only resist more malicious attacks than existing attack-specific adversarial training methods, but also prevent IIoT applications from new attacks. Yunfei Song, Tian Liu 0005, Tongquan Wei, Xiangfeng Wang 0001, Zhe Tao, Mingsong Chen 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2017 | Realtime Online Hot Topics Prediction in Sina Weibo for News Earlier ReportabstractWith the continuous growth of micro-blog services, Sina Weibo is increasingly found in the daily lives of ordinary Chinese individuals. More than one hundred million tweets are released in Sina Weibo everyday. By analyzing these mass data timely, media companies could learn how to generate buzz for new films, famous stars, or fashion shows more effectively. However, how to predict which topics will be the most popular search terms in Sina Weibo in realtime remains unknown. In this paper, we present a realtime hot topic prediction method in an online platform. Experiments are carried out on the platform to evaluate the proposed scheme. The results show that our model gets an average precision 44.32% and the median value is 45.83%. The proposed hot topic prediction method can predict the hot topics about 9.5 hours in average in advance. Sha Yuan, Zhe Tao, Tingshao Zhu, Shuotian Bai |
AINA | 2 |